Prosecution Insights
Last updated: August 17, 2026
Application No. 18/988,381

WEAKLY SUPERVISED ACTION SELECTION LEARNING IN VIDEO

Non-Final OA §101§DP
Filed
Dec 19, 2024
Priority
Apr 19, 2021 — provisional 63/176,858 +1 more
Examiner
BHATNAGAR, ANAND P
Art Unit
Tech Center
Assignee
The Toronto-dominion Bank
OA Round
1 (Non-Final)
91%
Grant Probability
Favorable
1-2
OA Rounds
11m
Est. Remaining
94%
With Interview

Examiner Intelligence

Grants 91% — above average
91%
Career Allowance Rate
662 granted / 724 resolved
+31.4% vs TC avg
Minimal +2% lift
Without
With
+2.3%
Interview Lift
resolved cases with interview
Typical timeline
2y 7m
Avg Prosecution
18 currently pending
Career history
740
Total Applications
across all art units

Statute-Specific Performance

§101
21.1%
-18.9% vs TC avg
§103
29.0%
-11.0% vs TC avg
§102
32.6%
-7.4% vs TC avg
§112
6.9%
-33.1% vs TC avg
Black line = Tech Center average estimate • Based on career data from 724 resolved cases

Office Action

§101 §DP
DETAILED ACTION Notice of Pre-AIA or AIA Status 1. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 101 2. 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. 3. Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The claims recite a mental process. This judicial exception is not integrated into a practical application. The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. The following reasons are provided to evaluate subject matter eligibility. (1) Are the claims directed to a process, machine, manufacture or composition of matter; (2A) Prong One: Are the claims directed to a judicially recognized exception, i.e., a law of nature, a natural phenomenon, or an abstract idea; Prong Two: If the claims are directed to a judicial exception under Prong One, then is the judicial exception integrated into a practical application; (2B) If the claims are directed to a judicial exception and do not integrate the judicial exception, do the claims provide an inventive concept. With regard to (1), the analysis is a ‘yes’, claim 1 recites a Machine, claim 8 recites a process and claim 15 recites a non-transitory computer readable medium. With regard to (2A) Prong One, the analysis is a “yes”. Claim 1 recites “applying a classification model to video data for a video to generate action class predictions for video segments of the video, wherein an action class prediction represents a likelihood that an action of an action class is depicted in a video segment, wherein an action class indicates a type of action.” When viewed under the broadest most reasonable interpretation the claim recites an abstract idea of mental processes. The step of “applying” is generically recited because there is no description of how this is accomplished. It can be interpreted as merely looking at the data, and evaluating the data in the mind. The concepts, as claimed, are observations and/or evaluations (“applying....”), judgements (“predictions..” and “compare...”), and opinions (“generating one or more predictions….”). There is nothing in the claim that requires more than an operation that a human, armed with the appropriate apparatus, pen/paper, can perform. One can perform the process using pen and paper, and the recitation of modules (such as judgers, gathering unit, an inference model) in the system/device claim is a mere use of generic computer components. See MPEP 2106.04 and the 2019 PEG. With regard to (2A) Prong Two: the analysis is a “No”. Claim 1 recites the additional elements of “applying an actionness model to the video data to generate actionness predictions for the video segments of the video, wherein an actionness prediction represents a likelihood that an action is depicted in a video segment; and generating one or more video class predictions based on the action class predictions generated by the classification model and the actionness predictions generated by the actionness model, wherein a video class prediction represents a likelihood that an action of an action class is depicted in the video.” These additional elements represents mere data gathering and indexing the data all together that is necessary for use of the recited abstract idea. Therefore, the limitation(s) is/are insignificant extra-solution activity, resulting in a generic operation. See MPEP 2106.05(1). The claim as a whole, looking at the additional elements individually and in combination, does not integrate the abstract idea into a practical application. With regard to (2B): the pending claims do not show what is more than a routine in the art presented in the claims, i.e., the additional elements are nothing more than routine and well-known steps. The additional elements do not reflect an improvement to a technology or technical field, including the use of a particular machine or particular transformation. It has not been shown that the mental process allows the “technology” to do something that it previously was not able to do. Claims 8 and 15 are similarly rejected for the same reasons as claim 1. Dependent claims 2-7, 9-14, and 16-20 do not include additional elements that are sufficient to amount to significantly more than the judicial exception. The claims are rejected for the same reasons and not repeated herewith. Double Patenting 4. The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969). A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on nonstatutory double patenting provided the reference application or patent either is shown to be commonly owned with the examined application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. See MPEP § 717.02 for applications subject to examination under the first inventor to file provisions of the AIA as explained in MPEP § 2159. See MPEP § 2146 et seq. for applications not subject to examination under the first inventor to file provisions of the AIA . A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b). The filing of a terminal disclaimer by itself is not a complete reply to a nonstatutory double patenting (NSDP) rejection. A complete reply requires that the terminal disclaimer be accompanied by a reply requesting reconsideration of the prior Office action. Even where the NSDP rejection is provisional the reply must be complete. See MPEP § 804, subsection I.B.1. For a reply to a non-final Office action, see 37 CFR 1.111(a). For a reply to final Office action, see 37 CFR 1.113(c). A request for reconsideration while not provided for in 37 CFR 1.113(c) may be filed after final for consideration. See MPEP §§ 706.07(e) and 714.13. The USPTO Internet website contains terminal disclaimer forms which may be used. Please visit www.uspto.gov/patent/patents-forms. The actual filing date of the application in which the form is filed determines what form (e.g., PTO/SB/25, PTO/SB/26, PTO/AIA /25, or PTO/AIA /26) should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto-processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to www.uspto.gov/patents/apply/applying-online/eterminal-disclaimer. U.S. application 18/988,381 U.S. patent 12,211,274 B2 1. A system comprising: a processor; and a non-transitory computer-readable medium storing instructions executable by the processor for: applying a classification model to video data for a video to generate action class predictions for video segments of the video, wherein an action class prediction represents a likelihood that an action of an action class is depicted in a video segment, wherein an action class indicates a type of action; applying an actionness model to the video data to generate actionness predictions for the video segments of the video, wherein an actionness prediction represents a likelihood that an action is depicted in a video segment; and generating one or more video class predictions based on the action class predictions generated by the classification model and the actionness predictions generated by the actionness model, wherein a video class prediction represents a likelihood that an action of an action class is depicted in the video. 3. The system of claim 1, wherein the instructions are further executable for: updating the classification model by comparing the one or more video class predictions with one or more action class labels associated with the video, wherein an action class label is a label of whether the video depicts an action of an action class. 1. A system comprising: a processor; and a non-transitory computer-readable medium storing a classification model that is trained to predict action classes for video segments of a video, wherein the classification model is trained on each training example of a plurality of training examples by performing steps comprising: applying the classification model to video data for a video of the training example to generate action class predictions for video segments of the video, wherein an action class prediction represents a likelihood that an action of an action class is depicted in a video segment, wherein an action class indicates a type of action; applying an actionness model to the video data of the training example to generate actionness predictions for the video segments of the video, wherein an actionness prediction represents a likelihood that an action is depicted in a video segment; generating one or more video class predictions based on the action class predictions generated by the classification model and the actionness predictions generated by the actionness model, wherein a video class prediction represents a likelihood that an action of an action class is depicted in the video of the training example; and updating the classification model by comparing the one or more video class predictions with one or more action class labels associated with the training example, wherein an action class label is a label of whether the video depicts an action of an action class. 2. The system of claim 1, wherein the instructions are further executable for: identifying one or more video segments of the video that depict an action for an action class based on the one or more video class predictions. 4. The system of claim 1, wherein the computer-readable medium further stores the actionness model, and wherein the actionness model is trained based on the video by: identifying a set of video segments as being likely to depict an action based on the action class predictions generated by the classification model and the actionness predictions generated by the actionness model; and updating the actionness model to predict whether an action is depicted in a video segment based on the identified set of video segments. 2. The system of claim 1, wherein the computer-readable medium further stores the actionness model, and wherein the actionness model is trained based on the training example by: identifying a set of video segments as being likely to depict an action based on the action class predictions generated by the classification model and the actionness predictions generated by the actionness model; and updating the actionness model to predict whether an action is depicted in a video segment based on the identified set of video segments. 5. The system of claim 1, wherein generating one or more video class predictions comprises: identifying, for each of one or more action classes, a set of video segments as being likely to depict an action of the action class, wherein the set of video segments are identified based on the action class predictions generated by the classification model and the actionness predictions generated by the actionness model; and generating a video class prediction for each of the one or more action classes based on action class predictions for the set of video segments associated with the action class. 3. The system of claim 1, wherein generating one or more video class predictions comprises: identifying, for each of one or more action classes, a set of video segments as being likely to depict an action of the action class, wherein the set of video segments are identified based on the action class predictions generated by the classification model and the actionness predictions generated by the actionness model; and generating a video class prediction for each of the one or more action classes based on action class predictions for the set of video segments associated with the action class. 6. The system of claim 1, wherein identifying a set of video segments for an action class comprises identifying a pre-determined number of video segments that are most likely to depict an action of the action class based on the action class predictions generated by the classification model and the actionness predictions generated by the actionness model. 4. The system of claim 1, wherein identifying a set of video segments for an action class comprises identifying a pre-determined number of video segments that are most likely to depict an action of the action class based on the action class predictions generated by the classification model and the actionness predictions generated by the actionness model. 7. The system of claim 1, wherein a video segment is a frame in the video or a time interval in the video. 5. The system of claim 1, wherein a video segment is one of a frame in the video or a time interval in the video. 8. A method comprising: applying a classification model to video data for a video to generate action class predictions for video segments of the video, wherein an action class prediction represents a likelihood that an action of an action class is depicted in a video segment, wherein an action class indicates a type of action; applying an actionness model to the video data to generate actionness predictions for the video segments of the video, wherein an actionness prediction represents a likelihood that an action is depicted in a video segment; and generating one or more video class predictions based on the action class predictions generated by the classification model and the actionness predictions generated by the actionness model, wherein a video class prediction represents a likelihood that an action of an action class is depicted in the video. 10. The method of claim 8, further comprising wherein the instructions are further executable for: updating the classification model by comparing the one or more video class predictions with one or more action class labels associated with the video, wherein an action class label is a label of whether the video depicts an action of an action class. 6. A method comprising: training a classification model to predict action classes for video segments of a video, wherein the classification model is trained on each training example of a plurality of training examples by performing steps comprising: applying the classification model to video data for a video of the training example to generate action class predictions for video segments of the video, wherein an action class prediction represents a likelihood that an action of an action class is depicted in a video segment, wherein an action class indicates a type of action; applying an actionness model to the video data of the training example to generate actionness predictions for the video segments of the video, wherein an actionness prediction represents a likelihood that an action is depicted in a video segment; generating one or more video class predictions based on the action class predictions generated by the classification model and the actionness predictions generated by the actionness model, wherein a video class prediction represents a likelihood that an action of an action class is depicted in the video of the training example; and updating the classification model by comparing the one or more video class predictions with one or more action class labels associated with the training example, wherein an action class label is a label of whether the video depicts an action of an action class. 9. The method of claim 8, further comprising: identifying one or more video segments of the video that depict an action for an action class based on the one or more video class predictions. 11. The method of claim 8, further comprising training the actionness model by: identifying a set of video segments as being likely to depict an action based on the action class predictions generated by the classification model and the actionness predictions generated by the actionness model; and updating the actionness model to predict whether an action is depicted in a video segment based on the identified set of video segments. 7. The method of claim 6, further comprising training the actionness model based on each training example of the plurality of training examples by performing steps comprising: identifying a set of video segments as being likely to depict an action based on the action class predictions generated by the classification model and the actionness predictions generated by the actionness model; and updating the actionness model to predict whether an action is depicted in a video segment based on the identified set of video segments. 12. The method of claim 8, wherein generating one or more video class predictions comprises: identifying, for each of one or more action classes, a set of video segments as being likely to depict an action of the action class, wherein the set of video segments are identified based on the action class predictions generated by the classification model and the actionness predictions generated by the actionness model; and generating a video class prediction for each of the one or more action classes based on action class predictions for the set of video segments associated with the action class. 8. The method of claim 6, wherein generating one or more video class predictions comprises: identifying, for each of one or more action classes, a set of video segments as being likely to depict an action of the action class, wherein the set of video segments are identified based on the action class predictions generated by the classification model and the actionness predictions generated by the actionness model; and generating a video class prediction for each of the one or more action classes based on action class predictions for the set of video segments associated with the action class. 13. The method of claim 8, wherein identifying a set of video segments for an action class comprises identifying a pre-determined number of video segments that are most likely to depict an action of the action class based on the action class predictions generated by the classification model and the actionness predictions generated by the actionness model. 9. The method of claim 6, wherein identifying a set of video segments for an action class comprises identifying a pre-determined number of video segments that are most likely to depict an action of the action class based on the action class predictions generated by the classification model and the actionness predictions generated by the actionness model. 14. The method of claim 8, wherein a video segment is a frame in the video or a time interval in the video. 10. The method of claim 6, wherein a video segment is one of a frame in the video or a time interval in the video. 15. A non-transitory computer-readable medium comprising instructions executable by a processor for: applying a classification model to video data for a video to generate action class predictions for video segments of the video, wherein an action class prediction represents a likelihood that an action of an action class is depicted in a video segment, wherein an action class indicates a type of action; applying an actionness model to the video data to generate actionness predictions for the video segments of the video, wherein an actionness prediction represents a likelihood that an action is depicted in a video segment; and generating one or more video class predictions based on the action class predictions generated by the classification model and the actionness predictions generated by the actionness model, wherein a video class prediction represents a likelihood that an action of an action class is depicted in the video. 16. The non-transitory computer-readable medium of claim 15, wherein the instructions are further executable for: identifying one or more video segments of the video that depict an action for an action class based on the one or more video class predictions. 17. The non-transitory computer-readable medium of claim 15, wherein the instructions are further executable for: updating the classification model by comparing the one or more video class predictions with one or more action class labels associated with the video, wherein an action class label is a label of whether the video depicts an action of an action class. 18. The non-transitory computer-readable medium of claim 15, wherein the computer-readable medium further stores the actionness model, and wherein the actionness model is trained based on the video by: identifying a set of video segments as being likely to depict an action based on the action class predictions generated by the classification model and the actionness predictions generated by the actionness model; and updating the actionness model to predict whether an action is depicted in a video segment based on the identified set of video segments. 11. A non-transitory computer-readable medium storing a set of weights for a video localization system, wherein the set of weights are generated by a process comprising: accessing the set of weights, wherein the set of weights comprise: a set of weights for a classification model for predicting action classes within video segments; and a set of weights for an actionness model for predicting whether a video segment depicts an action; storing a set of training examples, wherein each training example comprises: video data for a video, wherein the video data comprises a plurality of video segments of the video; and an action class label indicating an action class for an action performed in the video; for each training example in the set of training examples: generating a set of action class predictions for each video segment in the plurality of video segments of the training example by applying the classification model to the video data of the training example, where each action class prediction is associated with an action class of a set of action classes and represents a likelihood that the video segment depicts an action of the associated action class; generating an actionness prediction for each video segment in the plurality of video segments by applying the actionness model to the video data, where each actionness prediction represents a likelihood that an action is depicted by the video segment; identifying, for each action class of the set of action classes, a subset of the plurality of video segments as being likely to depict an action of the action class based on the set of action class predictions and the actionness prediction for each video segment of the plurality of video segments; updating the set of weights for the classification model based on an identified subset of the plurality of video segments associated with the action class indicated by the action class label for the training example; and updating the set of weights for the actionness model based on each of the identified subsets of the plurality of video segments; and storing the updated set of weights for the classification model and the updated set of weights for the actionness model on the computer-readable medium. 19. The non-transitory computer-readable medium of claim 15, wherein generating one or more video class predictions comprises: identifying, for each of one or more action classes, a set of video segments as being likely to depict an action of the action class, wherein the set of video segments are identified based on the action class predictions generated by the classification model and the actionness predictions generated by the actionness model; and generating a video class prediction for each of the one or more action classes based on action class predictions for the set of video segments associated with the action class. 16. The computer-readable medium of claim 11, wherein updating the set of weights for the classification model comprises: generating a video class prediction for the action class of the action class label based on action class predictions associated with the identified subset of video segments associated with the action class of the action class label, wherein the video class prediction represents a likelihood that an action of the action class is depicted in the video; and updating the set of weights of the classification model based on the video class prediction and the action class label. 20. The non-transitory computer-readable medium of claim 15, wherein identifying a set of video segments for an action class comprises identifying a pre- determined number of video segments that are most likely to depict an action of the action class based on the action class predictions generated by the classification model and the actionness predictions generated by the actionness model. 13. The computer-readable medium of claim 11, wherein identifying a subset of video segments for an action class of the set of action classes comprises identifying, based on actionness predictions and action class predictions for the action class associated with each video segment, a pre-determined number of video segments that are most likely to depict an action of the action class. Claims 1 and 3 are rejected on the ground of nonstatutory double patenting as being unpatentable over claim 1 of U.S. Patent No. 12,211,274 B2. Claims 2 and 4 are rejected on the ground of nonstatutory double patenting as being unpatentable over claim 2 of U.S. Patent No. 12,211,274 B2. Claim 5 is rejected on the ground of nonstatutory double patenting as being unpatentable over claim 3 of U.S. Patent No. 12,211,274 B2. Claim 6 is rejected on the ground of nonstatutory double patenting as being unpatentable over claim 4 of U.S. Patent No. 12,211,274 B2. Claim 7 is rejected on the ground of nonstatutory double patenting as being unpatentable over claim 5 of U.S. Patent No. 12,211,274 B2. Claims 8 and 10 are rejected on the ground of nonstatutory double patenting as being unpatentable over claim 6 of U.S. Patent No. 12,211,274 B2. Claim 9 is rejected on the ground of nonstatutory double patenting as being unpatentable over claim 7 of U.S. Patent No. 12,211,274 B2. Claim 11 is rejected on the ground of nonstatutory double patenting as being unpatentable over claim 7 of U.S. Patent No. 12,211,274 B2. Claim 12 is rejected on the ground of nonstatutory double patenting as being unpatentable over claim 8 of U.S. Patent No. 12,211,274 B2. Claim 13 is rejected on the ground of nonstatutory double patenting as being unpatentable over claim 9 of U.S. Patent No. 12,211,274 B2. Claim 14 is rejected on the ground of nonstatutory double patenting as being unpatentable over claim 10 of U.S. Patent No. 12,211,274 B2. Claims 15-18 are rejected on the ground of nonstatutory double patenting as being unpatentable over claim 11 of U.S. Patent No. 12,211,274 B2. Claim 19 is rejected on the ground of nonstatutory double patenting as being unpatentable over claim 16 of U.S. Patent No. 12,211,274 B2. Claim 20 is rejected on the ground of nonstatutory double patenting as being unpatentable over claim 13 of U.S. Patent No. 12,211,274 B2. Although the claims at issue are not identical, they are not patentably distinct from each other because the scope of claims of this instant invention are encompassed by the patented claims. Regarding claims 1-20: Prior art rejection is not made on the claimed subject matter since prior art was not found on the claimed subject matter. Contact Information 5. Any inquiry concerning this communication or earlier communications from the examiner should be directed to ANAND BHATNAGAR whose telephone number is (571)272-7416. The examiner can normally be reached on M-F 7:30am-4:00pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Vu Le can be reached on 571-272-4650. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /ANAND P BHATNAGAR/ Primary Examiner, Art Unit 2668 July 25, 2026
Read full office action

Prosecution Timeline

Dec 19, 2024
Application Filed
Jul 29, 2026
Non-Final Rejection mailed — §101, §DP (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12694473
ELECTRONIC DEVICE AND METHOD WITH IMAGE PROCESSING.
2y 9m to grant Granted Jul 28, 2026
Patent 12676973
IMAGE DECODING DEVICE, IMAGE DECODING METHOD, AND PROGRAM
2y 4m to grant Granted Jul 07, 2026
Patent 12664661
IMAGE PROCESSING APPARATUS AND IMAGE PROCESSING METHOD
2y 0m to grant Granted Jun 23, 2026
Patent 12657704
TIME PHASE DETERMINATION APPARATUS AND TIME PHASE DETERMINATION METHOD
3y 2m to grant Granted Jun 16, 2026
Patent 12657779
DATA PROCESSING METHOD AND APPARATUS, ELECTRONIC DEVICE, AND STORAGE MEDIUM
2y 6m to grant Granted Jun 16, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
91%
Grant Probability
94%
With Interview (+2.3%)
2y 7m (~11m remaining)
Median Time to Grant
Low
PTA Risk
Based on 724 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month