Prosecution Insights
Last updated: August 18, 2026
Application No. 18/807,734

SELF-SUPERVISED COMPOSITIONAL FEATURE REPRESENTATION FOR VIDEO UNDERSTANDING

Non-Final OA §103
Filed
Aug 16, 2024
Priority
Nov 14, 2023 — provisional 63/598,860
Examiner
KOETH, MICHELLE M
Art Unit
Tech Center
Assignee
Toyota Motor Corporation
OA Round
1 (Non-Final)
77%
Grant Probability
Favorable
1-2
OA Rounds
2m
Est. Remaining
94%
With Interview

Examiner Intelligence

Grants 77% — above average
77%
Career Allowance Rate
337 granted / 436 resolved
+17.3% vs TC avg
Strong +16% interview lift
Without
With
+16.4%
Interview Lift
resolved cases with interview
Fast prosecutor
2y 2m
Avg Prosecution
35 currently pending
Career history
473
Total Applications
across all art units

Statute-Specific Performance

§101
6.1%
-33.9% vs TC avg
§103
68.9%
+28.9% vs TC avg
§102
7.9%
-32.1% vs TC avg
§112
10.7%
-29.3% vs TC avg
Black line = Tech Center average estimate • Based on career data from 436 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Interpretation The following is a quotation of 35 U.S.C. 112(f): (f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph: An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked. As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph: (A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function; (B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and (C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function. Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function. Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function. Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitations are: “an intermediate feature selection module to pass,” “a video tubelet generation module to cluster,” “a video concept discovery module to cluster,” and “a video concept importance module to calculate,” in claim 17 and a sensor module to track in claim 20. Because these claim limitations are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, they are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof. Specifically, these limitations are interpreted to be computer-implemented means-plus-function limitations. See MPEP §2181(II)(B). If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. Claim Objections Claims 1, 9 and 17 and therefore all claims depending therefrom are objected to because of the following informalities: all of claims 1, 9 and 17 recite “to select an intermediate video feature of each of the set of videos” but should instead recite “to select intermediate video feature of each video of the set of videos,” and also all of claims 1, 9 and 17 recite “clustering the intermediate video feature of each of the set of videos” but should instead recite “clustering the intermediate video feature of the each video of the set of videos.” Appropriate correction is required. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claims 1–4, 9–12 and 17–18 are rejected under 35 U.S.C. 103 as being unpatentable over Gritsenko et al., US Patent Application Publication No. US 2024/0346824 A1 (herein “Gritsenko”) in view of Luo et al., "Trajectories as Topics: Multi-Object Tracking by Topic Discovery," in IEEE Transactions on Image Processing, vol. 28, no. 1, pp. 240-252, Jan. 2019, doi: 10.1109/TIP.2018.2866955 (herein “Luo”). Regarding claims 1, 9 and 17, with substantive differences between the claims noted in curly brackets, deficiencies of Gritsenko noted in square brackets, and claim 1 as exemplary, Gritsenko teaches {a method – claim 1 / A non-transitory computer-readable medium having program code recorded thereon – claim 9 / A system – claim 17 } for discovering human-interpretable concepts from video-based transformer models, {the program code being executed by a processor and – claim 9 / the system – claim 17} comprising (Gritsenko Abstract, method system and computer programs encoded on computer storage media for performing action localization on an input video, and specifying an action that is being performed in the video (human-interpretable concepts) and where ¶56 teaches processing including a video vision transformer): {program code to – claim 9 / an intermediate feature selection module to – claim 17} passing a set of videos through a video-based transformer model (Gritsenko fig. 1, input video (where a set can be comprised of one member video) is input to an video encoder neural network, and where ¶56 teaches that a species to the generic disclosed video encoder neural network is a Video Vision Transformer (ViVit) encoder) to select an intermediate video feature of each of the set of videos (Gritsenko fig. 1, ¶¶ 34 and 56, video transformer encoder generates (to select) a feature representation of the input video, where the feature representation 132 is further processed by being input to the decoder so it is an intermediate video feature); {program code to – claim 9 / a video tubelet generation module to – claim 17} clustering the intermediate video feature of each of the set of videos to obtain corresponding tubelets to the selected intermediate video features of each of the set of videos (Gritsenko fig. 3, ¶¶35, 57, 62–66, 82, features representation are input to a decoder and localization attention head and bounding boxes are output, the bounding boxes being grouped (clustering) according to whether they are for a background action, or for an agent action, and a single tubelet is generated for a group of bounding boxes and actions generated for a given spatial index); {program code to – claim 9 / a video concept discovery module to – claim 17} [clustering] an entire dataset of tubelets to form concepts of the set of videos (Gritsenko ¶82, classification head generates actions (concepts) for each tubelet (entire dataset)); and {program code to – claim 9 / a video concept importance module to – claim 17} calculating an importance of each of the concepts of the set of videos to an output of the video-based transformer model (Gritsenko fig. 3, ¶82, a set of scores respective to each potential action classified by the classification head is output, such as shown in fig. 3, the last tublet has a series of scores indicating the correspondence (importance) of the walk action (concept) as 0.8, the sit action (concept) as 0.3 and the watch action (concept) as 0.8). Gritsenko does not explicitly teach where Luo teaches clustering (Luo Abstract, page 240, introduction, and page 243, left column, frame by frame detections of a video are clustered to determine multi-objects as topics, where tracklets are spatiotemporal representations of an object in a video, and where the trackets are input to a dynamic clustering procedure grouping tracklets belonging to a specific object as a cluster, the clusters corresponding to topics of the video). Therefore, taking the teachings of Gritsenko and Luo together as a whole, it would have been obvious to a person having ordinary skill in the art (herein “PHOSITA”) before the effective filing date of the claimed invention to have modified the action label prediction for tubelets disclosed in Gritsenko to include the clustering specifically cited to above in Luo at least because doing so would provide more consistent tracking of objects in a video, leading to greater accuracy in video processing. See Luo page 247 right column, Part 1. Regarding claims 2 and 10, with differences between the claims noted in curly brackets, Gritsenko teaches {program code – claim 10} determining training protocols that produce models with desired concepts to provide downstream applications (Gritsenko ¶¶36, 95–97, and 111, fig. 5, training of the vision encoder neural network and decoder by a training system under any of a variety of labelling (supervised training) paradigms (training protocols), including labels from the AVA dataset (thus the training being specific to labels in the AVA dataset which are actions (desired concepts) for downstream applications such as detecting actions in video). Regarding claims 3 and 11, with differences between the claims noted in curly brackets, Gritsenko teaches {program code – claim 11} in which the training protocols comprise model pruning for improved performance and efficient action recognition, and model debugging (Gritsenko ¶108, for dense annotations in the training data, the training process (protocol) performs tubelet matching and selection of the one that minimizes the sum of the frame losses (thus is a type of pruning), with the resultant minimized frame loss being improved performance, also ¶109 teaching using an optimizer during the training (model debugging)). Regarding claims 4, 12 and 18, with differences between the claims noted in curly brackets, while Grisenko teaches to obtain the corresponding tubelets to the selected intermediate video features of each of the set of videos (Gritsenko fig. 3, ¶¶35, 57, 62–66, 82, features representation are input to a decoder and localization attention head and bounding boxes are output, the bounding boxes being grouped (clustering) according to whether they are for a background action, or for an agent action, and a single tubelet is generated for a group of bounding boxes and actions generated for a given spatial index), Grisenko does not explicitly teach where Luo teaches in which {the program code to – claim 12 } clustering the intermediate video feature of each of the set of videos comprises {program code – claim 12/ the video tubelet generation module is further to} applying simple linear iterative clustering (SLIC) to the intermediate video feature of each of the set of videos (Luo pages 243–244, section IV(A), fig. 2, input video is processed to determine bounding boxes segmented into superpixels using SLIC, each superpixel being a 5-dimensional vector (intermediate video feature) and clustered using K-means). Therefore, taking the teachings of Gritsenko and Luo together as a whole, it would have been obvious to a PHOSITA before the effective filing date of the claimed invention to have modified the video processing disclosed in Gritsenko to include the SLIC processing and clustering specifically cited to above in Luo at least because doing so would provide more consistent tracking of objects in a video, leading to greater accuracy in video processing. See Luo page 247 right column, Part 1. Claims 6–7 and 14–15 are rejected under 35 U.S.C. 103 as being unpatentable over Gritsenko in view of Luo as set forth above regarding the independent claims, and further in view of Dravid et al., “Rosetta Neurons: Mining the Common Units in a Model Zoo,” arXiv:2306.09346v2, June 16, 2023, [cs.CV], https://doi.org/10.48550/arXiv.2306.09346 (herein “Dravid”). Regarding claims 6 and 14, with differences between the claims noted in curly brackets {}, Gritsenko teaches in which {the program code to – claim 14} calculating the importance comprises: {program code to – claim 14} a set of the concepts from features of the video-based transformer model (Gritsenko ¶82, classification head that is part of the video transformer model generates actions (concepts) for each tubelet and a corresponding score (importance)) and {the program code to – claim 14}from a baseline performance of the video-based transformer model (Gritsenko ¶¶99–101, training loss is calculated (baseline performance) for the video transformer model), but does not teach the remainder of the limitations where Dravid teaches simultaneously removing a set of the concepts from features and calculating a performance degradation ((Dravid section 4.2, Fine-grained Rosetta Neurons edit section, concepts are removed from “rosetta neurons” which correspond to image features, and apply a re-optimization using an Adam optimizer (involving optimizing around a loss (performance degradation)). Therefore, taking the teachings of Gritsenko modified by Luo above and Dravid together as a whole, it would have been obvious to a PHOSITA before the effective filing date of the claimed invention to have modified the video processing disclosed in Gritsenko to include the rosetta neuron editing and re-optimization specifically cited to above in Dravid at least because doing so would allow for cross-class alignments, shifting, zooming and more without the need for specialized training. See Dravid Abstract. Regarding claims 7 and 15, with differences between the claims noted in curly brackets {}, Gritsenko teaches {program code to – claim 15} as human-interpretable concepts of the video-based transformer model (Gritsenko fig. 3, predicted labels from the video transformer model corresponding with tubelets including human-interpretable concepts of walk, sit and watch). Gritsenko does not explicitly teach where Dravid teaches displaying the set of the concepts associated with an increase of the performance degradation (Dravid fig. 2, a display of the set of all concepts emerging from the image processing model (thus including those associated with increased performance degradation as well)). Therefore, taking the teachings of Gritsenko modified by Luo above and Dravid together as a whole, it would have been obvious to a PHOSITA before the effective filing date of the claimed invention to have modified the video processing disclosed in Gritsenko to include the displaying of concepts specifically cited to above in Dravid at least because doing so would allow for cross-class alignments, shifting, zooming and more without the need for specialized training. See Dravid Abstract. Claims 8, 16 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Gritsenko in view of Luo as set forth above regarding the independent claims, and further in view of Van Hoorick et al., “Tracking through Containers and Occluders in the Wild,” arXiv:2305.03052v1 [cs.CV], May 4, 2023, https://doi.org/10.48550/arXiv.2305.03052 (herein “Hoorick”). Regarding claims 8, 16 and 20, with differences between the claims noted in curly brackets {}, Gritsenko teaches {program code to – claim 16} / a sensor module to – claim 20} tracking objects in the set of videos based on the importance of each of the concepts of the set of videos (Gritsenko fig. 3, ¶82, bounding boxes corresponding to one or more agent (tracked objects) are evaluated to determine a score (based on ) for each action (each concept) in the set of actions). Gritsenko as modified by Luo above does not, but Hoorick teaches tracking occluded objects (Hoorick page 1, introduction, tracking target objects as it becomes occluded or contained by other dynamic objects in a scene). Therefore, taking the teachings of Gritsenko modified by Luo above and Hoorick together as a whole, it would have been obvious to a PHOSITA before the effective filing date of the claimed invention to have modified the video processing disclosed in Gritsenko to include tracking occluded objects specifically cited to above in Hoorick at least because doing so would allow for a road agent to understand traffic situations more richly. See Hoorick page 1, Introduction section, left column. Allowable subject matter Claims 5, 13 and 19 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. Specifically, while the combination cited and applied to the independent claims 1, 9 and 17 of Gritsenko and Luo at least teaches claims noted in curly brackets {}, Gritsenko as modified by Luo in the rationale above for the independent claims teaches clustering the entire dataset of tubelets comprises clustering the entire dataset of tubelets to form the concepts of the set of videos, and cited but not applied reference Fel et al., “A Holistic Approach to Unifying Automatic Concept Extraction and Concept Importance Estimation,” arXiv:2306.07304v2, October 29, 2023, [cs.LG], https://doi.org/10.48550/arXiv.2306.07304 (herein “Fel”) teaches through non-negative matrix factorization (Fel page 3, section 2, page 5, first full paragraph, convex NMF (non-negative matrix factorization) used to extract meaningful concepts from high-dimensional image data), none of Gritsenko, Luo, Fel or any of the other cited art of record in any combination obvious to a PHOSITA teaches or suggests convex NMF as claimed. Therefore claims 5, 13 and 19 are allowable over the prior art. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure: Rawat et al., US Patent Application Publication No. US 2025/0029410 A1, directed towards an active sparse labeling system for video. Rawat et al., US Patent Application Publication No. US 2022/0222940 A1, directed towards categorizing actions in video by evaluating spatio-temporal action tubelets. Yu et al., US Patent No. US 10,068,138 B2, directed towards video-event classification by extracting frame-level sets of visual features from video. Any inquiry concerning this communication or earlier communications from the examiner should be directed to MICHELLE M KOETH whose telephone number is (571)272-5908. The examiner can normally be reached Monday-Thursday, 09:00-17:00, Friday 09:00-13:00, EDT/EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Vincent Rudolph can be reached at 571-272-8243. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. MICHELLE M. KOETH Primary Examiner Art Unit 2671 /MICHELLE M KOETH/Primary Examiner, Art Unit 2671
Read full office action

Prosecution Timeline

Aug 16, 2024
Application Filed
Jul 13, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12705775
METHOD AND APPARATUS FOR OBTAINING 3D INFORMATION OF VEHICLE
3y 11m to grant Granted Aug 11, 2026
Patent 12700397
SOUND OUTPUT CONTROL DEVICE, SOUND OUTPUT CONTROL METHOD, AND SOUND OUTPUT CONTROL PROGRAM
2y 11m to grant Granted Aug 04, 2026
Patent 12682672
IDENTIFYING DOCUMENT GENERATORS BY COLOR FOOTPRINTS
3y 11m to grant Granted Jul 14, 2026
Patent 12670545
CASCADED LOCAL IMPLICIT TRANSFORMER FOR ARBITRARY-SCALE SUPER-RESOLUTION
3y 2m to grant Granted Jun 30, 2026
Patent 12664808
Fake Signature Detection
3y 8m to grant Granted Jun 23, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
77%
Grant Probability
94%
With Interview (+16.4%)
2y 2m (~2m remaining)
Median Time to Grant
Low
PTA Risk
Based on 436 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month