Prosecution Insights
Last updated: October 01, 2026
Application No. 18/827,133

Systems and Methods for Improved Video Understanding

Non-Final OA §102§103
Filed
Sep 06, 2024
Priority
Jul 08, 2021 — divisional of 12/112,538
Examiner
TUCKER, WESLEY J
Art Unit
Tech Center
Assignee
Google LLC
OA Round
1 (Non-Final)
84%
Grant Probability
Favorable
1-2
OA Rounds
11m
Est. Remaining
89%
With Interview

Examiner Intelligence

Grants 84% — above average
84%
Career Allowance Rate
614 granted / 734 resolved
+23.7% vs TC avg
Moderate +5% lift
Without
With
+5.4%
Interview Lift
resolved cases with interview
Typical timeline
3y 0m
Avg Prosecution
19 currently pending
Career history
744
Total Applications
across all art units

Statute-Specific Performance

§101
13.5%
-26.5% vs TC avg
§103
37.8%
-2.2% vs TC avg
§102
37.0%
-3.0% vs TC avg
§112
8.4%
-31.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 734 resolved cases

Office Action

§102 §103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 102 The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. Claims 1-3 and 8-14 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by USPN 2022/0222940 to Rawat et al. With regard to claim 1, Rawat discloses a computer-implemented method for training a video understanding model for classifying video data with improved accuracy, the method comprising: obtaining, by a computing system comprising a plurality of computing devices, pretrained model data descriptive at least in part of a video understanding model, the video understanding model comprising at least one parameter of a video transformer encoder model (Fig. 2A, and paragraphs [0005]-[0006] and [0046]-[0047], Localization network shown I Fig. 2A is a pretrained model that is used to recognize video objects/actions. The localization network is used to identify and label the actions with bounding box annotations that are both spatial and temporal, and these are interpreted as parameters. The localization network uses a video encoder as shown.); training, by the computing system, the pretrained model data based at least in part on a first dataset, the first dataset comprising image data, to determine first updated model data (paragraphs [0055]-[0058], Rawat discloses two sets of data sets are used for training the system and then another training set for evaluating the system. The first datasets are used to train the model of the classifier network); and training, by the computing system, the first updated model data based at least in part on a second dataset, the second dataset comprising video data, to determine trained model data descriptive of a trained version of the video understanding model (paragraphs [0055]-[0058], The second datasets are considered to be the evaluating datasets that are used on the model once it is trained using the first datasets. All the data sets are annotated video data and are used to determine the accuracy or loss of the video action recognition for the video frame sets determined by the classification network). With regard to claim 2, Rawat discloses the computer-implemented method of claim 1, wherein the video transformer encoder comprises a factorized encoder, the factorized encoder comprising a spatial transformer encoder and a temporal transformer encoder (paragraph [0044] "To perform the localization of each video clip, the method uses an encoder-decoder structure (such as a 3D convolution-based encoder, e.g. I3D [3]) which extracts class-agnostic action features (such as spatio-temporal features that are required for activity localization) and generates segmentation masks for each clip. The decoder uses these features to segment regions from the original input which contain activities." The encoder is sued to generate spatiotemporal annotations for the video segments). With regard to claim 3, Rawat discloses the computer-implemented method of claim 2, wherein the spatial transformer encoder is configured to receive the plurality of video tokens and produce, in response to receipt of the plurality of video tokens, a plurality of temporal representations (paragraphs [0013]-[0015] and [0041], spatio-temporal action tubelets are identified using bounding boxes over a series of video frames to identify objects/actions. The video tokens are interpreted as the annotations in the form of bounding boxes indexed to specific image frames. The annotations and identified action frames are used to create tubelets that define the action in a spatiotemporal format); and wherein the temporal transformer encoder is configured to receive the plurality of temporal representations and produce, in response to receipt of the plurality of temporal representations, a spatiotemporal representation of the video data, wherein the spatiotemporal representation is classified to produce the classification output (paragraphs [0013]-[0015] and [0041], spatio-temporal action tubelets are identified using bounding boxes over a series of video frames to identify objects/actions. The annotations and identified action frames are used to create tubelets that define the action in a spatiotemporal format). With regard to claim 8, Rawat discloses the computer-implemented method of claim 1, wherein training, by the computing system, the first updated model data based at least in part on a second dataset comprises extracting a plurality of video tokens from the video data (paragraphs [0013]-[0015] and [0041], spatio-temporal action tubelets are identified using bounding boxes over a series of video frames to identify objects/actions. The annotations and identified action frames are used to create tubelets that define the action in a spatiotemporal format. The identified tubelets are used and compared with known annotated video frames in order to train the classification network as shown in Fig. 2A). With regard to claim 9, Rawat discloses the computer-implemented method of claim 8, wherein extracting the plurality of video tokens comprises: extracting, by the computing system, a plurality of video tubelets from the video data (paragraphs [0013]-[0015] and [0041], spatio-temporal action tubelets are identified using bounding boxes over a series of video frames to identify objects/actions.); projecting, by the computing system, the plurality of video tubelets to a plurality of tensor representations of the plurality of video tubelets (paragraphs [0050]-[0051], Rawat discloses a tensor representation in the vector set as described); and merging, by the computing system, the plurality of tensor representations along at least one dimension to produce the plurality of video tokens (paragraphs [0050]-[0052], As series of tubelets are recognized and categorized, they can be merged to represent as action defined by multiple sequential tubelets, thus merging the representation defining the recognized action). With regard to claim 10, Rawat discloses the computer-implemented method of claim 9, wherein each of the plurality of video tubelets spans one of the plurality of video frames (paragraphs [0013]-[0015] and [0041], spatio-temporal action tubelets are identified using bounding boxes over a series of video frames to identify objects/actions. A tubelet spanning one frame would just be a single bounding box identification for a single frame). With regard to claim 11, Rawat discloses the computer-implemented method of claim 9, wherein each of the plurality of video tubelets spans two or more of the plurality of video frames (paragraphs [0013]-[0015] and [0041], spatio-temporal action tubelets are identified using bounding boxes over a series of video frames to identify objects/actions). With regard to claim 12, Rawat discloses the computer-implemented method of claim 9, wherein the plurality of video tokens are single-dimensional (paragraphs [0013]-[0015] and [0041]-[0043], spatio-temporal action tubelets are identified using bounding boxes over a series of video frames to identify objects/actions. A tubelet spanning one frame would just be a single bounding box identification for a single frame which is interpreted as single-dimensional. Video tokens identify the frames as well as the space within the frame that an object/action has been identified. The frame number would be an example of a single-dimensional token). With regard to claim 13, Rawat discloses the computer-implemented method of claim 9, wherein the plurality of video tubelets are nonoverlapping (paragraph [0041], the plurality of tubelets can be merged. Since the tubelets occur in series in video frames, it is interpreted that they are not overlapping). With regard to claim 14, Rawat discloses the computer-implemented method of claim 8, wherein positional embeddings are added to the plurality of video tokens and input to the video understanding model (paragraphs [0041], the tubelets are passed to a classification network where the tubelets are recognized and classified in terms of their respective depicted actions. The localization network determines which pixels and their positions correspond to the recognized actions). Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim 15 is rejected under 35 U.S.C. 103 as being unpatentable over the combination of USPN 2022/0222940 to Rawat et al and publication titled “Social Fabric: Tunelet Compositions for Video Relation Detection” to Chen et al. With regard to claim 15, Rawat discloses the computer-implemented method of claim 1, but does not explicitly disclose wherein the transformer encoder model comprises at least one normalization layer. Chen discloses a tubelet composition determination similar to that of Rawat and further teaches a normalization layer (Section 3, third paragraph: “On top of the features, we apply layer normalization [3], followed by a linear layer to obtain embedded representation Ri ⊂ R ∈ RN× D. In this D-dimensional embedding space, we learn a set C ∈ RK×D consisting of K primitives. The idea behind our encoding is to describe a tubelet pair entirely as a weighted combination of these primitives.”). Therefore it would have been obvious to one of ordinary skill in the art before time of filing to use the normalization taught by Chen in order to represent the tubelet accurately with weighted factors. Claim 16 is rejected under 35 U.S.C. 103 as being unpatentable over the combination of USPN 2022/0222940 to Rawat et al and publication titled “Spatio-temporal Tubelet Feature Aggregation and Object Linkin in Videos” to Cores at al. With regard to claim 16, Rawat discloses the computer-implemented method of claim 1, but does not explicitly disclose wherein the transformer encoder model comprises at least one multi-layer perceptron layer. Cores teaches a spatio-temporal tubelet generating system similar to that of Rawat and further teaches a multi-layer perceptron with two fully connected layers (Fig. 1). Therefore it would have been obvious to one of ordinary skill in the art before time of filing to use a perceptron with multiple layers as taught by Cores in order to aid in the classification output of the classifier network of Rawat. Allowable Subject Matter Claims 4-7 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. Claims 4 and 6 contain allowable subject matter. Claims 5 and 7 depend from claims 4 and 6 respectively. Contact Information Any inquiry concerning this communication or earlier communications from the examiner should be directed to WESLEY J TUCKER whose telephone number is (571)272-7427. The examiner can normally be reached 9AM-5PM Monday-Friday. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, JOHN VILLECCO can be reached at 571-272-7319. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /WESLEY J TUCKER/Primary Examiner, Art Unit 2661
Read full office action

Prosecution Timeline

Sep 06, 2024
Application Filed
Aug 13, 2026
Non-Final Rejection mailed — §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12749338
IMAGE PROCESSING DEVICE, IMAGE PROCESSING METHOD, AND PROGRAM
3y 0m to grant Granted Sep 29, 2026
Patent 12743874
MODEL PRECONDITIONING FOR FACE RECOGNITION
3y 0m to grant Granted Sep 22, 2026
Patent 12737903
METHODS FOR DETERMINING AND REPORTING VEHICLE FOLLOWING DISTANCE
2y 0m to grant Granted Sep 15, 2026
Patent 12730868
PROCESSING SYSTEM, INFORMATION PROCESSING APPARATUS, NON-TRANSITORY COMPUTER-READABLE STORAGE MEDIUM STORING CONTROL PROGRAM, AND IMAGE PROCESSING APPARATUS
3y 7m to grant Granted Sep 08, 2026
Patent 12731280
FACE IMAGE DISPLAYING METHOD, READABLE STORAGE MEDIUM, PROGRAM PRODUCT, AND ELECTRONIC DEVICE
2y 8m to grant Granted Sep 08, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
84%
Grant Probability
89%
With Interview (+5.4%)
3y 0m (~11m remaining)
Median Time to Grant
Low
PTA Risk
Based on 734 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month