DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Interpretation
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked.
As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph:
(A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function;
(B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and
(C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function.
Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function.
Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function.
Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action.
This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitations are: “an intermediate feature selection module to pass,” “a video tubelet generation module to cluster,” “a video concept discovery module to cluster,” and “a video concept importance module to calculate,” in claim 17 and a sensor module to track in claim 20.
Because these claim limitations are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, they are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof. Specifically, these limitations are interpreted to be computer-implemented means-plus-function limitations. See MPEP §2181(II)(B).
If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph.
Claim Objections
Claims 1, 9 and 17 and therefore all claims depending therefrom are objected to because of the following informalities: all of claims 1, 9 and 17 recite “to select an intermediate video feature of each of the set of videos” but should instead recite “to select intermediate video feature of each video of the set of videos,” and also all of claims 1, 9 and 17 recite “clustering the intermediate video feature of each of the set of videos” but should instead recite “clustering the intermediate video feature of the each video of the set of videos.” Appropriate correction is required.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 1–4, 9–12 and 17–18 are rejected under 35 U.S.C. 103 as being unpatentable over Gritsenko et al., US Patent Application Publication No. US 2024/0346824 A1 (herein “Gritsenko”) in view of Luo et al., "Trajectories as Topics: Multi-Object Tracking by Topic Discovery," in IEEE Transactions on Image Processing, vol. 28, no. 1, pp. 240-252, Jan. 2019, doi: 10.1109/TIP.2018.2866955 (herein “Luo”).
Regarding claims 1, 9 and 17, with substantive differences between the claims noted in curly brackets, deficiencies of Gritsenko noted in square brackets, and claim 1 as exemplary, Gritsenko teaches {a method – claim 1 / A non-transitory computer-readable medium having program code recorded thereon – claim 9 / A system – claim 17 } for discovering human-interpretable concepts from video-based transformer models, {the program code being executed by a processor and – claim 9 / the system – claim 17} comprising (Gritsenko Abstract, method system and computer programs encoded on computer storage media for performing action localization on an input video, and specifying an action that is being performed in the video (human-interpretable concepts) and where ¶56 teaches processing including a video vision transformer):
{program code to – claim 9 / an intermediate feature selection module to – claim 17} passing a set of videos through a video-based transformer model (Gritsenko fig. 1, input video (where a set can be comprised of one member video) is input to an video encoder neural network, and where ¶56 teaches that a species to the generic disclosed video encoder neural network is a Video Vision Transformer (ViVit) encoder) to select an intermediate video feature of each of the set of videos (Gritsenko fig. 1, ¶¶ 34 and 56, video transformer encoder generates (to select) a feature representation of the input video, where the feature representation 132 is further processed by being input to the decoder so it is an intermediate video feature);
{program code to – claim 9 / a video tubelet generation module to – claim 17} clustering the intermediate video feature of each of the set of videos to obtain corresponding tubelets to the selected intermediate video features of each of the set of videos (Gritsenko fig. 3, ¶¶35, 57, 62–66, 82, features representation are input to a decoder and localization attention head and bounding boxes are output, the bounding boxes being grouped (clustering) according to whether they are for a background action, or for an agent action, and a single tubelet is generated for a group of bounding boxes and actions generated for a given spatial index);
{program code to – claim 9 / a video concept discovery module to – claim 17} [clustering] an entire dataset of tubelets to form concepts of the set of videos (Gritsenko ¶82, classification head generates actions (concepts) for each tubelet (entire dataset)); and
{program code to – claim 9 / a video concept importance module to – claim 17} calculating an importance of each of the concepts of the set of videos to an output of the video-based transformer model (Gritsenko fig. 3, ¶82, a set of scores respective to each potential action classified by the classification head is output, such as shown in fig. 3, the last tublet has a series of scores indicating the correspondence (importance) of the walk action (concept) as 0.8, the sit action (concept) as 0.3 and the watch action (concept) as 0.8).
Gritsenko does not explicitly teach where Luo teaches clustering (Luo Abstract, page 240, introduction, and page 243, left column, frame by frame detections of a video are clustered to determine multi-objects as topics, where tracklets are spatiotemporal representations of an object in a video, and where the trackets are input to a dynamic clustering procedure grouping tracklets belonging to a specific object as a cluster, the clusters corresponding to topics of the video).
Therefore, taking the teachings of Gritsenko and Luo together as a whole, it would have been obvious to a person having ordinary skill in the art (herein “PHOSITA”) before the effective filing date of the claimed invention to have modified the action label prediction for tubelets disclosed in Gritsenko to include the clustering specifically cited to above in Luo at least because doing so would provide more consistent tracking of objects in a video, leading to greater accuracy in video processing. See Luo page 247 right column, Part 1.
Regarding claims 2 and 10, with differences between the claims noted in curly brackets, Gritsenko teaches {program code – claim 10} determining training protocols that produce models with desired concepts to provide downstream applications (Gritsenko ¶¶36, 95–97, and 111, fig. 5, training of the vision encoder neural network and decoder by a training system under any of a variety of labelling (supervised training) paradigms (training protocols), including labels from the AVA dataset (thus the training being specific to labels in the AVA dataset which are actions (desired concepts) for downstream applications such as detecting actions in video).
Regarding claims 3 and 11, with differences between the claims noted in curly brackets, Gritsenko teaches {program code – claim 11} in which the training protocols comprise model pruning for improved performance and efficient action recognition, and model debugging (Gritsenko ¶108, for dense annotations in the training data, the training process (protocol) performs tubelet matching and selection of the one that minimizes the sum of the frame losses (thus is a type of pruning), with the resultant minimized frame loss being improved performance, also ¶109 teaching using an optimizer during the training (model debugging)).
Regarding claims 4, 12 and 18, with differences between the claims noted in curly brackets, while Grisenko teaches to obtain the corresponding tubelets to the selected intermediate video features of each of the set of videos (Gritsenko fig. 3, ¶¶35, 57, 62–66, 82, features representation are input to a decoder and localization attention head and bounding boxes are output, the bounding boxes being grouped (clustering) according to whether they are for a background action, or for an agent action, and a single tubelet is generated for a group of bounding boxes and actions generated for a given spatial index), Grisenko does not explicitly teach where Luo teaches in which {the program code to – claim 12 } clustering the intermediate video feature of each of the set of videos comprises {program code – claim 12/ the video tubelet generation module is further to} applying simple linear iterative clustering (SLIC) to the intermediate video feature of each of the set of videos (Luo pages 243–244, section IV(A), fig. 2, input video is processed to determine bounding boxes segmented into superpixels using SLIC, each superpixel being a 5-dimensional vector (intermediate video feature) and clustered using K-means).
Therefore, taking the teachings of Gritsenko and Luo together as a whole, it would have been obvious to a PHOSITA before the effective filing date of the claimed invention to have modified the video processing disclosed in Gritsenko to include the SLIC processing and clustering specifically cited to above in Luo at least because doing so would provide more consistent tracking of objects in a video, leading to greater accuracy in video processing. See Luo page 247 right column, Part 1.
Claims 6–7 and 14–15 are rejected under 35 U.S.C. 103 as being unpatentable over Gritsenko in view of Luo as set forth above regarding the independent claims, and further in view of Dravid et al., “Rosetta Neurons: Mining the Common Units in a Model Zoo,” arXiv:2306.09346v2, June 16, 2023, [cs.CV], https://doi.org/10.48550/arXiv.2306.09346 (herein “Dravid”).
Regarding claims 6 and 14, with differences between the claims noted in curly brackets {}, Gritsenko teaches in which {the program code to – claim 14} calculating the importance comprises: {program code to – claim 14} a set of the concepts from features of the video-based transformer model (Gritsenko ¶82, classification head that is part of the video transformer model generates actions (concepts) for each tubelet and a corresponding score (importance)) and {the program code to – claim 14}from a baseline performance of the video-based transformer model (Gritsenko ¶¶99–101, training loss is calculated (baseline performance) for the video transformer model), but does not teach the remainder of the limitations where Dravid teaches simultaneously removing a set of the concepts from features and calculating a performance degradation ((Dravid section 4.2, Fine-grained Rosetta Neurons edit section, concepts are removed from “rosetta neurons” which correspond to image features, and apply a re-optimization using an Adam optimizer (involving optimizing around a loss (performance degradation)).
Therefore, taking the teachings of Gritsenko modified by Luo above and Dravid together as a whole, it would have been obvious to a PHOSITA before the effective filing date of the claimed invention to have modified the video processing disclosed in Gritsenko to include the rosetta neuron editing and re-optimization specifically cited to above in Dravid at least because doing so would allow for cross-class alignments, shifting, zooming and more without the need for specialized training. See Dravid Abstract.
Regarding claims 7 and 15, with differences between the claims noted in curly brackets {}, Gritsenko teaches {program code to – claim 15} as human-interpretable concepts of the video-based transformer model (Gritsenko fig. 3, predicted labels from the video transformer model corresponding with tubelets including human-interpretable concepts of walk, sit and watch). Gritsenko does not explicitly teach where Dravid teaches displaying the set of the concepts associated with an increase of the performance degradation (Dravid fig. 2, a display of the set of all concepts emerging from the image processing model (thus including those associated with increased performance degradation as well)).
Therefore, taking the teachings of Gritsenko modified by Luo above and Dravid together as a whole, it would have been obvious to a PHOSITA before the effective filing date of the claimed invention to have modified the video processing disclosed in Gritsenko to include the displaying of concepts specifically cited to above in Dravid at least because doing so would allow for cross-class alignments, shifting, zooming and more without the need for specialized training. See Dravid Abstract.
Claims 8, 16 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Gritsenko in view of Luo as set forth above regarding the independent claims, and further in view of Van Hoorick et al., “Tracking through Containers and Occluders in the Wild,” arXiv:2305.03052v1 [cs.CV], May 4, 2023, https://doi.org/10.48550/arXiv.2305.03052 (herein “Hoorick”).
Regarding claims 8, 16 and 20, with differences between the claims noted in curly brackets {}, Gritsenko teaches {program code to – claim 16} / a sensor module to – claim 20} tracking objects in the set of videos based on the importance of each of the concepts of the set of videos (Gritsenko fig. 3, ¶82, bounding boxes corresponding to one or more agent (tracked objects) are evaluated to determine a score (based on ) for each action (each concept) in the set of actions). Gritsenko as modified by Luo above does not, but Hoorick teaches tracking occluded objects (Hoorick page 1, introduction, tracking target objects as it becomes occluded or contained by other dynamic objects in a scene).
Therefore, taking the teachings of Gritsenko modified by Luo above and Hoorick together as a whole, it would have been obvious to a PHOSITA before the effective filing date of the claimed invention to have modified the video processing disclosed in Gritsenko to include tracking occluded objects specifically cited to above in Hoorick at least because doing so would allow for a road agent to understand traffic situations more richly. See Hoorick page 1, Introduction section, left column.
Allowable subject matter
Claims 5, 13 and 19 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. Specifically, while the combination cited and applied to the independent claims 1, 9 and 17 of Gritsenko and Luo at least teaches claims noted in curly brackets {}, Gritsenko as modified by Luo in the rationale above for the independent claims teaches clustering the entire dataset of tubelets comprises clustering the entire dataset of tubelets to form the concepts of the set of videos, and cited but not applied reference Fel et al., “A Holistic Approach to Unifying Automatic Concept Extraction and Concept Importance Estimation,” arXiv:2306.07304v2, October 29, 2023, [cs.LG], https://doi.org/10.48550/arXiv.2306.07304 (herein “Fel”) teaches through non-negative matrix factorization (Fel page 3, section 2, page 5, first full paragraph, convex NMF (non-negative matrix factorization) used to extract meaningful concepts from high-dimensional image data), none of Gritsenko, Luo, Fel or any of the other cited art of record in any combination obvious to a PHOSITA teaches or suggests convex NMF as claimed. Therefore claims 5, 13 and 19 are allowable over the prior art.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure:
Rawat et al., US Patent Application Publication No. US 2025/0029410 A1, directed towards an active sparse labeling system for video.
Rawat et al., US Patent Application Publication No. US 2022/0222940 A1, directed towards categorizing actions in video by evaluating spatio-temporal action tubelets.
Yu et al., US Patent No. US 10,068,138 B2, directed towards video-event classification by extracting frame-level sets of visual features from video.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to MICHELLE M KOETH whose telephone number is (571)272-5908. The examiner can normally be reached Monday-Thursday, 09:00-17:00, Friday 09:00-13:00, EDT/EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Vincent Rudolph can be reached at 571-272-8243. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
MICHELLE M. KOETH
Primary Examiner
Art Unit 2671
/MICHELLE M KOETH/Primary Examiner, Art Unit 2671