Prosecution Insights
Last updated: October 01, 2026
Application No. 19/181,064

VIDEO GENERATION METHOD, APPARATUS, DEVICE, MEDIUM AND PROGRAM PRODUCT

Non-Final OA §102
Filed
Apr 16, 2025
Priority
Apr 16, 2024 — CN 202410458518.3
Examiner
PARRA, OMAR S
Art Unit
2421
Tech Center
2400 — Computer Networks
Assignee
Beijing Youzhuju Network Technology Co., Ltd.
OA Round
1 (Non-Final)
74%
Grant Probability
Favorable
1-2
OA Rounds
1y 4m
Est. Remaining
84%
With Interview

Examiner Intelligence

Grants 74% — above average
74%
Career Allowance Rate
518 granted / 696 resolved
+16.4% vs TC avg
Moderate +9% lift
Without
With
+9.2%
Interview Lift
resolved cases with interview
Typical timeline
2y 10m
Avg Prosecution
20 currently pending
Career history
721
Total Applications
across all art units

Statute-Specific Performance

§101
7.2%
-32.8% vs TC avg
§103
52.1%
+12.1% vs TC avg
§102
23.6%
-16.4% vs TC avg
§112
4.2%
-35.8% vs TC avg
Black line = Tech Center average estimate • Based on career data from 696 resolved cases

Office Action

§102
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 102 The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. Claim(s) 1-3, 7-10, 14-17 and 20 is/are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Gupta et al. (“Towards Generating Ultara-High Resolution Talking-Face Videos with Lip synchronization”, 2023 IEEE/CVF; hereinafter “Gupta-IEEE”, presented on IDS 07/28/2025). Regarding claims 1, 8 and 15, Gupta-IEEE teaches a computer device (computing devices on the deep networks) (with corresponding method and non-transitory computer readable medium), comprising: a memory and a processor, the memory and the processor communicating with each other, the memory having computer instructions stored therein, and the processor (these elements are inherent in all the computing devices of the neural networks) being configured to execute the computer instructions to: acquire target audio data and first video data of a target object (input speech and input face, Fig. 3: Overall Inference pipeline column); acquire second video data, the second video data being obtained by performing mask processing on a lip area in video data of the target object (masked face, Fig. 3: Overall Inference pipeline column); perform feature processing on the target audio data based on a target multimodal model to obtain a target audio feature, the target multimodal model being obtained based on performing synchronization alignment training of a sample audio feature and a sample video feature on paired sample audio and sample video; perform feature extraction on the first video data and the second video data to obtain a feature to be processed; and predict a lip area in the second video data based on the target audio feature and the feature to be processed, to determine a target video corresponding to the target audio data (Stage 2, Fig. 3: Overall Inference pipeline: post-processing). Regarding claims 2, 9 and 16, Gupta-IEEE teaches wherein the target multimodal model is determined by: acquiring positive sample data and negative sample data, the positive sample data comprising synchronous first sample audio and first sample video, and the negative sample data comprising asynchronous second sample audio and second sample video; and performing synchronization alignment training of the sample audio feature and the sample video feature on a preset multimodal model based on the positive sample data and the negative sample data, to obtain the target multimodal model (Section 3.1” For training the lip-sync expert, we sample video-speech pairs from the same time step (in-sync, i.e., positive pairs) and random pairs from different time-steps (out-of-sync, i.e., negative pairs)”). Regarding claim 3, 10 and 17, Gupta-IEEE teaches wherein predicting the lip area in the second video data based on the target audio feature and the feature to be processed, to determine the target video corresponding to the target audio data comprises: inputting the target audio feature and the feature to be processed into a target image generation model, and predicting the lip area in the second video data to obtain target feature data, the target image generation model being obtained based on performing parameter update of a sample audio feature output by the target multimodal model and a video feature of a sample video of a sample object; and decoding the target feature data to obtain the target video (target video generation and lip synchronization for model audio updates in Fig. 3: Overall Inference pipeline). Regarding claims 7, 14 and 20, Gupta-IEEE teaches wherein acquiring the target audio data comprises: acquiring a target text and a target timbre; and converting the target text into the target audio data based on the target timbre (“Our model works for any in-the-wild unseen identities, languages”, page 5200). (a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. Claim(s) 1, 2, 8, 9, 15 and 16 is/are rejected under 35 U.S.C. 102(a)(2) as being anticipated by Gupta et al. (hereinafter ‘Gupta’, Patent No. 12,505,863). Regarding claims 1, 8 and 15, Gupta teaches computer device (210, Fig. 2) (with corresponding method and non-transitory computer readable medium), comprising: a memory and a processor, the memory and the processor communicating with each other (215 and memory, Fig. 2; col. 4 lines 9-32), the memory having computer instructions stored therein, and the processor being configured to execute the computer instructions to: acquire target audio data and first video data of a target object (102a and 102b, Fig. 1; col. 1 line 63 to col. 2 line 8); acquire second video data, the second video data being obtained by performing mask processing on a lip area in video data of the target object (viseme clips of the lips are obtained based on the target, dubbed audio, col. 2 line 40) ; perform feature processing on the target audio data based on a target multimodal model to obtain a target audio feature, the target multimodal model being obtained based on performing synchronization alignment training of a sample audio feature and a sample video feature on paired sample audio and sample video (based on the target audio, language features, phoneme; multiple corresponding visemes are identified, including viseme features such as matrix of coordinates for facial landmarks, lines and vertices, col. 2 lines 21-40); perform feature extraction on the first video data and the second video data to obtain a feature to be processed (source video and audio are analyzed to determine lip poses of the speaker, and based on phonemes of the source audio, lip pose that share indices with the source visemes are compared, col. 2 line 41-67) ; and predict a lip area in the second video data based on the target audio feature and the feature to be processed, to determine a target video corresponding to the target audio data (col. 2 line 57 to col. 3 line 16, where the best matching visemes are predicted) . Regarding claims 2, 9 and 16, Gupta teaches wherein the target multimodal model is determined by: acquiring positive sample data and negative sample data, the positive sample data comprising synchronous first sample audio and first sample video (machine learning is used to determine how lip poses and visemes positively sync through error), col. 2 line 52 to col. 3 line 16; col. 6 line 25 to col. 7 line 4), and the negative sample data comprising asynchronous second sample audio and second sample video (when negative synchronization or de-synchronization is detected based on lip pose-viseme distance is over a threshold value, further modification is performed for improve alignment, col. 7 lines 5-28); and performing synchronization alignment training of the sample audio feature and the sample video feature on a preset multimodal model based on the positive sample data and the negative sample data, to obtain the target multimodal model (col. 2 line 52 to col. 3 line 16; col. 6 line 25 to col. 7 line 4; col. 7 lines 5-28). Allowable Subject Matter Claims 4-6, 11-13, 18 and 19 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to OMAR S PARRA whose telephone number is (571)270-1449. The examiner can normally be reached M-F: Mostly 10-6PM. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Nathan Flynn can be reached at 571-2721915. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /OMAR S PARRA/Primary Examiner, Art Unit 2421
Read full office action

Prosecution Timeline

Apr 16, 2025
Application Filed
Aug 11, 2026
Non-Final Rejection mailed — §102 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12744961
SET-TOP BOX WITH ENHANCED BEHAVIORAL CONTROLS AND SYSTEM AND METHOD FOR USE OF SAME
1y 8m to grant Granted Sep 22, 2026
Patent 12732674
GENERATING A CUSTOMIZED HIGHLIGHT SEQUENCE DEPICTING MULTIPLE EVENTS
1y 9m to grant Granted Sep 08, 2026
Patent 12726678
MODEL-BASED DATA PROCESSING METHOD AND APPARATUS
3y 3m to grant Granted Sep 01, 2026
Patent 12720129
SYSTEMS AND METHODS FOR ALTERING A PROGRESS BAR TO PREVENT SPOILERS IN A MEDIA ASSET
2y 10m to grant Granted Aug 25, 2026
Patent 12708857
PROP FOR AN ATTRACTION SYSTEM
2y 12m to grant Granted Aug 18, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
74%
Grant Probability
84%
With Interview (+9.2%)
2y 10m (~1y 4m remaining)
Median Time to Grant
Low
PTA Risk
Based on 696 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month