Prosecution Insights
Last updated: October 02, 2026
Application No. 18/498,153

METHODS AND SYSTEMS FOR FACILITATING ANNOTATION OF VIDEOS

Final Rejection §103
Filed
Oct 31, 2023
Examiner
HAUK, EMILY ROSE
Art Unit
2669
Tech Center
2600 — Communications
Assignee
Toyota Motor Corporation
OA Round
2 (Final)
100%
Grant Probability
Favorable
3-4
OA Rounds
0m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 100% — above average
100%
Career Allowance Rate
7 granted / 7 resolved
+38.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
2y 5m
Avg Prosecution
12 currently pending
Career history
14
Total Applications
across all art units

Statute-Specific Performance

§101
11.4%
-28.6% vs TC avg
§103
51.9%
+11.9% vs TC avg
§102
13.9%
-26.1% vs TC avg
§112
21.5%
-18.5% vs TC avg
Black line = Tech Center average estimate • Based on career data from 7 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Amendment The present application, amendments filed on 05/21/2026 have been entered in full. Claims 1-20 remain pending in this application. The amendments to the Specification have overcome each and every objection to the Specification. Response to Arguments Applicant’s arguments, see Page 8, filed 05/21/2026, with respect to U.S.C. 101 rejections of pending claims 1-20 have been fully considered and are persuasive. The U.S.C. 101 rejections of claims 1-20 has been withdrawn. Applicant’s arguments, see 8-10, filed 05/21/2026, with respect to the rejection(s) of claim(s) 1 and 11 under U.S.C. 103 have been fully considered and are persuasive. Therefore, the rejection has been withdrawn. However, upon further consideration, a new ground(s) of rejection is made in view of the incorporation of Williams US12321385 (hereinafter “Williams”). Sahni nor Sharma teach the extraction of frames of a video. However, Williams, teaches the use of a video data modeling engine to extract video frames of the video’s content. Regarding dependent claims 2-10 and 12-20, as the independent base claims are rejected with the incorporation of Williams, the arguments the dependent claims patentable are not persuasive. Claim Rejections - 35 USC § 103 The text of those sections of Title 35, U.S. Code not included in this action can be found in a prior Office action. Claims 1, 3-11, 13-15, and 17-20 are rejected under 35 U.S.C. 103 as being unpatentable over US 12406480 (hereinafter “Sahni”) in view of Sharma US10977518 (hereinafter “Sharma”) in view of Williams US12321385 (hereinafter “Williams”). Regarding claim 1, Sahni teaches a method for facilitating human annotation of videos comprising (see col 4 lines 4-7 and 10-17, an annotation system for videos that send the content [videos] to participating annotators): “selecting” (see col 7 lines 43-51, the data selector selects a set of content items from the content database [col 6 lines 1-8, the database stores content and the content includes images or videos] for annotating that are associated with a job [a job is created and managed based on a job description provided by an administrator [user], col 4 line 51 through col 5 line 5]. The selection criteria for selecting content may be associated with a prediction metric made by the machine learning engine [col 7 lines 36-38] which may use neural networks [col 6 lines 25-27]. While Sahni teaching selecting content items, it does not explicitly state the step of extracting frames in video); assigning, using the one or more trained neural networks, the annotation task to one or more annotators based on matchability scores (see col 5 lines 55-62, a job description may identify a set of annotators for receiving the content items for the annotation job based on specific criteria including field of expertise and level of experience [matchability score]. See figure 1, the annotation system 100 includes machine learning engine 126 which may use neural networks [col 6 lines 25-27]), (see col 9 lines 35-41, a user [annotator] can flag content, the content may be flagged for low quality); and sending notifications of the flagged annotations to the user (see col 9 lines 35-37, the content flagged by the annotator is flagged for review by the administrator [user], interpreted as the content is flagged and sent to the user); and wherein the one or more neural networks are pre-trained based on sample input videos, sample annotation tasks, and sample annotations (Col 8 lines 35-43, the training and updating of the machine learning model associated with a particular job [task] based on the labels [annotations] and associated content items [input image or videos, col 4 line 62-63]). Sahni does not explicitly teach extracting frames in a video, assigning the annotation task to one or more annotators based on historical annotation performance of the annotators; determining, using the one or more trained neural networks, whether one or more confidence scores of one or more annotations in annotated frames received from the one or more annotators are below a threshold; in response to a determination that the one or more confidence scores are below the threshold, flagging the one or more annotations associated with the one or more confidence scores below the threshold; and pre-training based on sample annotation associated with sample annotators and continuously training based on the assigned annotators. Sharma teaches assigning the annotation task to one or more annotators based on historical annotation performance of the annotators (see col 6 lines 28, the Annotation Job Controller may select a set of annotators for the job via the annotator information in the annotation job repository. The annotator information may include the quality score [see figure 1 and col 7 lines 21-23], the quality score [historical annotation performance indicating the level of correctness of the annotations provides by the annotator over time [col 7, lines 45-49]); determining, using the one or more trained neural networks, whether one or more confidence scores of one or more annotations in annotated frames received from the one or more annotators are below a threshold (see col 8 lines 21-28 and line 63 through col 9 line 5, the use of a confidence score [indicating the level of confidence that the annotation is proper, col 7 lines 3-12] to compare to a threshold as part of the active learning model [which may comprise of a machine learning model such as a convolutional neural network trained with input data, col 8 lines 51-54); in response to a determination that the one or more confidence scores are below the threshold, flagging the one or more annotations associated with the one or more confidence scores below the threshold (see col 8 lines 63-65, when the confidence score is below a threshold the data element may be marked as “bad”); and pre-training based on sample annotation associated with sample annotators and continuously training based on the assigned annotators (See col 8 lines 46-62, training a ML model with input data using the annotations of the involved annotators and the scores of the involved annotators. The training may iterative). Sahni and Sharma are analogous art because they are from the same field of endeavor of a video annotation system, which includes machine learning using neural networks, that selects a set of annotators for the annotation jobs and marks the content based on its quality. Before the effective filling date of the invention, it would have been obvious to one of ordinary skill in the art to modify Sahni to incorporate the assigning tasks using historical annotation performance, determining whether the confidence score of an annotation is below a threshold, and flagging annotation below a threshold as taught by Sharma. The motivation for doing so would have been to give higher weight to the annotations of the annotators with higher historical performance and compare the annotation of a task across annotators (Sharma, col 7 lines 13-44). Sharma nor Sahni explicitly teach extracting frames in a video. Williams teaches extracting frames in a video (see col 3 lines 45-49, the video data modeling engine to extract video frames of the video’s content); determining whether one or more confidence scores of frames are below a threshold (see col 3 line 8-22, determining the confidence score of the video and whether the confidence score is below a threshold; flagging the one or more annotations (see col 5, line1-5, flagging the video-label pair [annotation]). Williams, Sahni and Sharma are analogous art because they are from the same field of endeavor of the automating of video annotation or labelling by sending videos to annotators for consideration. Before the effective filling date of the invention, it would have been obvious to one of ordinary skill in the art to modify Sahni and Sharma to extract frames of a video as taught by Williams. The motivation for doing so would have been to properly summarize the video (Williams, col 3 lines 45-49). Regarding claim 3, Sharma, Williams, and Sahni teach the method of claim 1. Sharma teaches the historical annotation performance of the annotators is determined based on (see col 7 lines 50-60, the quality score [historic annotation performance] is may be based on the difference between the annotators annotation and the consolidated annotation [a combination of multiple annotations from multiple annotators, col 7 lines 6-10]). Sharma does not teach annotating completion rates. Sahni teaches annotating completion rates (see col 8 lines 44-54, the cycle manager tracks the progress of the annotation job and determines when the job is complete using predefined completion criteria which includes obtaining labels for at least a predefined number of content items [completion rate]). Sahni and Sharma are analogous art because they are from the same field of endeavor of a video annotation system, which includes machine learning using neural networks, that selects a set of annotators for the annotation jobs and marks the content based on its quality. Before the effective filling date of the invention, it would have been obvious to one of ordinary skill in the art to modify Sharma to incorporate annotating completion rates as taught by Sahni. The motivation for doing so would have been to track progress of each job and determine when the job is completed by the standards of the cycle manager (Sahni, col 8 lines 44-54). Regarding claim 4, Sharma, Williams, and Sahni teach the method of claim 1. Sharma teaches the annotators are one or more servers or persons on an annotation platform (see col 4 lines 58-63, annotators are humans or machine learning models that provide annotations of some sort). Regarding claim 5, Sharma, Williams, and Sahni teach the method of claim 1. Sharma teaches after assigning, the method further comprises monitoring annotation progress, the annotation progress comprising (see col 7 lines 15-40, the use of the annotation consolidation module to combine a collection of annotations and determine how similar or different they are, and using annotators with a higher quality score to weight scores and determine what is the correct annotation [accurate annotation]). Sharma does not teach completion rates. Sahni teaches after assigning, the method further comprises monitoring annotation progress, the annotation progress comprising completion rates and accuracy (see col 8 lines 44-54, the cycle manager tracks the progress of the annotation job and determines when the job is complete using predefined completion criteria which includes obtaining labels for at least a predefined number of content items [completion rate] and at least a predefined average prediction metric for prediction made [accuracy]). Regarding claim 6, Sharma, Williams, and Sahni teach the method of claim 1. Williams teaches the method further comprises sending the notifications of the flagged annotations to the one or more annotators (see col 4 lines 60-67, if the verification of the successful labelling fails the verification task may be re-issued to a different annotator). Regarding claim 7, Sharma, Williams, and Sahni teach the method of claim 1. Sharma teaches the annotation task is selected from object recognition and, segmentation, pose estimation, action recognition, attribute recognition of people, or a combination thereof detection (see col 5 lines 38-45, the user may select between annotation tasks including image classification, object detection with additional embodiments allowing for different annotation tasks). Regarding claim 8, Sharma, Williams, and Sahni teach the method of claim 1. Sahni teaches the one or more trained neural networks are trained based on historical manual selections of frames in historical videos associated with historical annotation tasks and feedback of historical frame selections from the user (see col 4 lines 45-51, training the machine learning model [which may include neural networks] which includes identifying training content items [interpreted as identified frames selected] and the associated annotations [interpreted as feedback]). Regarding claim 9, Sharma, Williams, and Sahni teach the method of claim 1. Sharma teaches the one or more trained neural networks are trained based on historical annotator selections by the user in association with the annotation task and feedback of generated assignment (see col 7, the weighted selection of annotator based on the annotator quality score and the updating [training] of quality scores based on the contribution to the consolidated annotations [interpreted as feedback of the annotation job assignment]). Regarding claim 10, Sharma, Williams, and Sahni teach the method of claim 1. Sharma teaches the one or more trained neural networks are trained based on feedback of the flagged annotations from the user and manual flags marked by the user (see col 8 lines 46-60 the machine learning model such as a convolutional neural network is trained using input data with consolidated annotations, as well as metadata including confidence score of annotations [used to flag annotations], individual annotations, and quality scores of annotators. The job submitters [user] may select representation of “bad” examples to aid as part of the instructions [col 2, lines 43-50], interpreted as manually marked by the user). Claims 11 and 13-20 are analogous system to method claims 1 and 3-10, respectively, thus are analyzed and rejected similar to claims 1 and 3-10. Claims 2 and 12 are rejected under 35 U.S.C. 103 as being unpatentable over Sahni in view of Sharma in view of Williams in view of Jannink US 20210362363 (hereinafter “Jannink”). Regarding claim 2, Sharma, Williams, and Sahni teach the method of claim 1. Sahni teaches the matchability scores of the annotators are determined based on (see col 5 lines 55-67, the use of the job description including specific criteria of the annotator [annotator task] and the annotator having specific criteria including field of expertise [annotation task performed by the annotator] to determine which annotators or group is assigned the job. Sahni does not explicitly teach the distance between the annotation task and the historical annotation task; an additional reference will provide obviousness). Sahni, Williams, and Sharma do not explicitly teach the distances between the annotation task and historical annotation tasks performed by the annotators. Jannink teaches the distances between the annotation task and historical annotation tasks performed by the annotators (see paragraph 0070, the classifier may compute a similarity between representation of user request [annotation task] and a representation of previously received user requests [historical annotation task]. The user request comprising a task to be performed [paragraph 0040]). Sharma, Williams, Jannink, and Sahni are analogous art because they are from the same field of endeavor of using machine learning models with annotation based on user inputs. Before the effective filling data of the inventions, it would have been obvious to one of ordinary skill in the art to modify Sahni, Williams, and Sharma to incorporate the distance between the annotation task and the historical annotation tasks preformed as taught by Jannink. The motivation for doing so would have been to make selections based on the smallest distance when comparing possible tasks (Jannink, paragraph 0070). Claim 12 is analogous system to the method of claim 2, thus claim 12 is analyzed and rejected similar to claim 2. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to EMILY R. HAUK whose telephone number is (571)272-5966. The examiner can normally be reached M-F 8:00-5:00. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Chan Park can be reached at 571-272-7409. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /EMILY ROSE HAUK/Examiner, Art Unit 2669 /CHAN S PARK/Supervisory Patent Examiner, Art Unit 2669
Read full office action

Prosecution Timeline

Show 2 earlier events
Apr 26, 2026
Interview Requested
May 04, 2026
Examiner Interview Summary
May 04, 2026
Applicant Interview (Telephonic)
May 21, 2026
Response Filed
Aug 20, 2026
Final Rejection mailed — §103
Sep 21, 2026
Interview Requested
Sep 30, 2026
Examiner Interview Summary
Sep 30, 2026
Applicant Interview (Telephonic)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12700113
LANGUAGE-BASED LEARNING FOR MONOCULAR DEPTH ESTIMATION
2y 6m to grant Granted Aug 04, 2026
Patent 12670718
VEHICLE VIOLATION DETECTION METHOD AND VEHICLE VIOLATION DETECTION SYSTEM
2y 7m to grant Granted Jun 30, 2026
Patent 12646158
DEVICE AND METHOD FOR DETECTING WILDFIRE
2y 9m to grant Granted Jun 02, 2026
Patent 12620216
LATENT DIFFUSION MODEL AUTODECODERS
2y 4m to grant Granted May 05, 2026
Study what changed to get past this examiner. Based on 4 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
100%
Grant Probability
99%
With Interview (+0.0%)
2y 5m (~0m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 7 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month