Prosecution Insights
Last updated: October 01, 2026
Application No. 19/084,573

VIDEO GENERATION METHOD, READABLE MEDIUM, AND ELECTRONIC DEVICE

Non-Final OA §103§112
Filed
Mar 19, 2025
Priority
Mar 19, 2024 — CN 202410316872.2
Examiner
AHMAD, NAUMAN UDDIN
Art Unit
Tech Center
Assignee
Beijing Youzhuju Network Technology Co., Ltd.
OA Round
1 (Non-Final)
74%
Grant Probability
Favorable
1-2
OA Rounds
12m
Est. Remaining
91%
With Interview

Examiner Intelligence

Grants 74% — above average
74%
Career Allowance Rate
36 granted / 49 resolved
+13.5% vs TC avg
Strong +17% interview lift
Without
With
+17.2%
Interview Lift
resolved cases with interview
Typical timeline
2y 6m
Avg Prosecution
33 currently pending
Career history
75
Total Applications
across all art units

Statute-Specific Performance

§101
3.6%
-36.4% vs TC avg
§103
76.6%
+36.6% vs TC avg
§102
3.3%
-36.7% vs TC avg
§112
14.2%
-25.8% vs TC avg
Black line = Tech Center average estimate • Based on career data from 49 resolved cases

Office Action

§103 §112
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 3-5, 8, 11-13, 16 and 19-20 rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Claims 3, 11 and 19 recites the limitation “performing feature fusion on the image” in line 5, line 5 and line 6 (respectively). There is insufficient antecedent basis for this limitation in the claim. This is because it is unclear whether this “image” is of the initial image sequence or of the target image sequence. Claims 4, 12 and 20 further add to this lack of clarity by referring to the image by stating “feature corresponding to the image” in line 2 of each claim which is still unclear if this “image” is of the initial image sequence or of the target image sequence. Note: this lack of clarity can be corrected by referring to “the image” as “the image of the initial image sequence” or “the image of the target image sequence”. Claims 3, 11 and 19 recites the limitation "and generating the video frame sequence" in line 7, line 7 and line 8 (respectively). There is insufficient antecedent basis for this limitation in the claim. This is because it is unclear if this is the same “generating” as before (aforementioned in each parent claim) or a newer instance of generating the video frame sequence. Claims 5 and 13 recites the limitation "extracting…down-sampling" in line 5 of each claim. There is insufficient antecedent basis for this limitation in the claim. This is because it is unclear if this “extracting” and “down-sampling” is the same as that of each parent claim or a newer instance of “extracting” and “down-sampling”. Claims 8 and 16 recites the limitation "for video generation" in line 1 of each claim. There is insufficient antecedent basis for this limitation in the claim. This is because it is unclear whether these are new instances or old instances (first introduced in parent claim(s)) of “video generation”, Note. Most likely these claims depend on some dependent claim or are missing elements. In order to fix this issue, dependency should be reviewed and any first instance of an element should be made clear that it’s a first instance and should be referred to as “a” or “an” instead of “the”, and if multiple instances exist, further instances should be further distinguished for example by saying “first”, “second”, and/or “third” etc. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1-4, 9-12, and 17-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Prajwal et al. (A Lip Sync Expert Is All You Need for Speech to Lip Generation In the Wild), hereinafter referenced as Prajwal in view of Li et al. (U.S. Patent Application Publication No. 2023/0177646), hereinafter referenced as Li. Regarding claim 1, Prajwal teaches A video generation method, comprising: obtaining a talking video of a target object and a target text for video generation (fig. 1 and description thereof teaches “Our novelWav2Lip model produces significantly more accurate lip-synchronization in dynamic, unconstrained talk-ing face videos” and page 7, left col. (LHC), first paragraph teaches “The third and final type of videos, “TTS", has been specifically chosen for benchmarking the lip-syncing performance on synthetic speech obtained from a text-to-speech system”); this shows the video generation model obtaining talking face video and the TTS shows target text for video generation; and generating, by using the talking video, the target text, and a video generation model, a target video of a digital human corresponding to the target object talking according to the target text, (figs. 1-2 show Wav2lip model producing target video of human (in digital format) corresponding to object talking/lip syncing according to target text and page 1, right col. (RHC) teaches “our challenging benchmarks show that the lip-sync accuracy of the videos generated by our Wav2Lip model is almost as good as real synced videos”); this shows aforementioned target text and talking video used in model to generate target video of digital human for target object talking according to target text (associated with the lip sync); wherein the video generation model is configured to generate the target video by following operations: extracting an initial image sequence from the talking video, wherein each of the images in the initial image sequence comprises a face of the target object (page 4, RHC, first paragraph teaches “Speech Encoder is also a stack of 2Dconvolutions to encode the input speech segment 𝑆 which is then concatenated with the face representation” and fig. 1 shows images in initial sequence comprising face of target object); face representations concatenated to encode input speech shows extracting initial image sequence from talking video and each of these are shown as images in an initial image sequence with face of target object; generating a video frame sequence corresponding to a video of the digital human corresponding to the target object talking according to the target text, according to the target image sequence and an audio sequence corresponding to the target text (page 4, RHC 2nd paragraph teaches “encoder-decoder network that generates each frame independently” and page 4, LHC, first paragraph teaches “for each sample that indicates the probability that the input audio-video pair is in sync:”); frame sequence generated here is the decoded version therefore corresponding to video of digital human corresponding to target object talking according to target text and this would be in accordance to audio (and corresponding target text thereof from TTS) due to speech and lip-sync as well as in accordance to target image sequence (by down sampled frames from the below combination) because entire video is first downsampled (into target image sequence) to save storage in the reference below. However, Prajwal fails to explicitly teach and down-sampling images in the initial image sequence to obtain a target image sequence; and up-sampling video frames in the video frame sequence to obtain the target video. However, Li explicitly teaches and down-sampling images in the initial image sequence to obtain a target image sequence (Li, paragraph 98 teaches “To reduce a storage amount occupied by the video, compression or downsampling processing may be performed on the video, to obtain video data that occupies a smaller storage amount”); this shows down-sampling video (which shows frames/initial image sequence) to obtain compressed video (target image sequence); and up-sampling video frames in the video frame sequence to obtain the target video (Li, paragraph 98 teaches “When the user plays the video by using the terminal, super-resolution processing may be performed on the stored video data by using the image processing method provided in this disclosure, to obtain video data with higher resolution, thereby improving viewing experience of the user.” and paragraph 113 teaches “the first image may be decomposed in a manner of downsampling and upsampling,”); this shows up-sampling of video frames in the video frame sequence would occur to obtain target video. Li is considered to be analogous art because it is reasonably pertinent to the problem faced by the inventor of down-sampling then up-sampling to decrease processing costs and reduce storage usage. Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify Prajwal's invention with the down-sampling then up-sampling techniques of Li to reduce a storage amount (Li paragraph 98) and ensure improving viewing experience of the user (Li, paragraph 113). This would be done by the improved efficiency due to down-sampling then up-sampling. Regarding claim 2, the combination of Prajwal and Li teaches wherein the images in the initial image sequence have a first resolution, and the down-sampling the images in the initial image sequence to obtain the target image sequence comprises: down-sampling the images in the initial image sequence to obtain the target image sequence having a second resolution, wherein the first resolution is higher than the second resolution (Li, paragraph 103 teaches “If the video data is a high-resolution video, to reduce a bandwidth occupied for transmitting the video data, the server may perform downsampling on the video data to obtain a low-resolution video”); video data/images being high-resolution shows images in initial image sequence having a first resolution, then after the down-sampling, the lower-resolution indicates a second resolution which means the first resolution is higher than the second; and the up-sampling the video frames in the video frame sequence to obtain the target video comprises: up-sampling the video frames in the video frame sequence to obtain the target video having the first resolution (Li, paragraph 103 teaches “Therefore, after receiving the low-resolution video, as shown in FIG. 5B, the local device may perform super-resolution processing on the low-resolution video to obtain a high-resolution video”); super-resolution/up-sampling the video frames in the video frame sequence to obtain high-resolution shows obtaining target video as having the first resolution. The same motivations used in claim 1 apply here in claim 2. Regarding claim 3, the combination of Prajwal and Li teaches wherein the generating the video frame sequence corresponding to the video of the digital human corresponding to the target object talking according to the target text, according to the target image sequence and the audio sequence corresponding to the target text comprises: for each image in the target image sequence, determining a previous image and a next image of the image in the target image sequence, (Li, paragraph 12 teaches “performing fusion on a second structure sub-image obtained in previous iteration and a second detail sub-image obtained in the previous iteration, to obtain a first fused image of current iteration” and paragraph 16 teaches ”where the first hidden state information is used to process a next frame of image that is in the video data and that is arranged in the first image”); this shows that the aforementioned generating step (from claim 1) would have sub-image from previous iteration determined (previous image) and next frame/image determined due to the first hidden state information and since this is for a iterative process, it is done for each image in the target image sequence; and performing feature fusion on the image, the previous image, and the next image to obtain an image fusion feature (Li, paragraph 26 teaches “perform iterative fusion on the second structure sub-image and the second detail sub-image for at least one time to obtain an updated second structure sub-image and an updated second detail sub-image; and extract a feature from the updated second structure sub-image”, paragraph 119 teaches “hidden state information may be understood as a feature map generated by a network, including a feature extracted from a past frame, and is stored historical information. In a super-resolution processing process, a hidden state information provides historical information, and performs time-space fusion with a feature of a current input frame, so that more feature expressions can be obtained”, and paragraph 177 teaches “where the first hidden state information includes a feature extracted from a second image, and the second image includes at least one frame of image in the video data adjacent to the first image”); iterative fusion would mean it would also be for next frame when next iteration occurs and this describes feature fusion for previous/past frame/image and current frame/image; and generating the video frame sequence corresponding to the video of the digital human corresponding to the target object talking according to the target text, according to image fusion features of all images in the target image sequence and the audio sequence corresponding to the target text (Li, paragraph 6 teaches “fusing first hidden state information and the first structure sub-image to obtain a second structure sub-image, and splicing the first hidden state information and the first detail sub-image to obtain a second detail sub-image, where the first hidden state information includes a feature extracted from a second image, and the second image includes at least one frame of image in the video data adjacent to the first image”); this shows that the aforementioned generating step (from claim 1) would also be according to image fusion features of all images in target image sequence due to the aforementioned iteration and first hidden state information including features from multiple adjacent frames in the video data. The same motivations used in claim 1 apply here in claim 3. Regarding claim 4, the combination of Prajwal and Li teaches wherein the image fusion feature comprises a first fusion feature corresponding to the image, (Li, paragraph 14 teaches “fusing the structure feature and the detail feature to obtain a second fused image”); second fused image with fusion feature shows first fusion feature corresponding to the image; a second fusion feature corresponding to the previous image, (Li, paragraph 13 teaches “in each iterative fusion process, the second structure sub-image obtained in the previous iteration and the second detail sub-image obtained in the previous iteration may be fused, and the second structure sub-image and the second detail sub-image are fused separately by using the first fused image obtained through fusion,”); since this is for previous iteration image, it shows second fusion corresponding to previous sub-image, and this fuses details/features (therefore second fusion feature); and a third fusion feature corresponding to the next image (Li, paragraph 14 teaches “performing amplification processing on the second fused image to obtain the output image, where the resolution of the output image is higher than resolution of the second fused image”); output image of second fused feature would be considered next image (which is also of a fusion feature therefore considered third fusion feature) since it is after the second fused image; and the generating the video frame sequence corresponding to the video of the digital human corresponding to the target object talking according to the target text, according to the image fusion features of all the images in the target image sequence and the audio sequence corresponding to the target text comprises: for the image fusion feature of each image in the target image sequence, according to the image fusion feature and the audio sequence, generating a first video frame corresponding to the first fusion feature, (Li, paragraph 6 teaches “the second image includes at least one frame of image in the video data adjacent to the first image” and paragraph 10 teaches “performing iterative fusion”); this shows that the aforementioned generating step according to image fusion features and audio sequence comprises first video data/frame generated (which corresponds to first fusion feature since second image above is equated to first fusion feature corresponding to the image); a second video frame corresponding to the second fusion feature, (Li, paragraph 6 teaches “first structure sub-image and a first detail sub-image, where the first image is any frame of image in video data”); this shows that the aforementioned generating step according to image fusion features and audio sequence comprises second video data/frame generated (which corresponds to second fusion feature since sub-image above is equated to second fusion feature corresponding to the previous image); and a third video frame corresponding to the third fusion feature (Li, paragraph 7 teaches “processing of the video data… a finally obtained output image are more enriched”); this shows that the aforementioned generating step according to image fusion features and audio sequence comprises third video data/frame generated (which corresponds to third fusion feature since output image above is equated to third fusion feature corresponding to the next image); and obtaining the video frame sequence according to first video frames of all the images in the target image sequence (Li, paragraph 110 teaches “video data may include a plurality of frames of images, and the first image is any one of the frames of images.”); when done iteratively as aforementioned, this shows obtainment of video frame sequence would be according to first video frames of all images in the target image sequence since first image is any one of the frames of the images. The same motivations used in claim 1 apply here in claim 4. Regarding claim 9, the device claim 9 recites similar limitations as method claim 1, and thus is rejected under similar rationale. In addition, Li fig. 19 shows device with processor 1901 (processing apparatus) and memory 1902 (storage apparatus) and paragraph 32 teaches “processor invokes program code in the memory to perform a processing-related function in the image processing method”. Regarding claim 10, the device claim 10 recites similar limitations as method claim 2, and thus is rejected under similar rationale. Regarding claim 11, the device claim 11 recites similar limitations as method claim 3, and thus is rejected under similar rationale. Regarding claim 12, the device claim 12 recites similar limitations as method claim 4, and thus is rejected under similar rationale. Regarding claim 17, the transitory computer-readable medium claim 17 recites similar limitations as method claim 1, and thus is rejected under similar rationale. In addition, Li claim 15 teaches “non-transitory machine-readable storage medium having instructions stored therein, which when executed by a processor, cause the processor to perform operations” and paragraph 32 teaches “processor invokes program code in the memory to perform a processing-related function in the image processing method”. Regarding claim 18, the transitory computer-readable medium claim 18 recites similar limitations as method claim 2, and thus is rejected under similar rationale. Regarding claim 19, the transitory computer-readable medium claim 19 recites similar limitations as method claim 3, and thus is rejected under similar rationale. Regarding claim 20, the transitory computer-readable medium claim 20 recites similar limitations as method claim 4, and thus is rejected under similar rationale Claim(s) 5 and 13 is/are rejected under 35 U.S.C. 103 as being unpatentable over thec combination of Prajwal and Li as applied to claims 1 and 9 above, and further in view of Kalva et al. (U.S. Patent Application Publication No. 2026/0222616), hereinafter referenced as Kalva. Regarding claim 5, the combination of Prajwal and Li teaches wherein the video generation model comprises an audio encoding module, an image encoding module, and a decoding module, (Prajwal, page 3, RHC, second to last paragraph teaches “a face encoder and an audio encoder,”, and page 4, RHC, first paragraph teaches “a (iii) Face Decoder.”); this shows audio encoding module and face/image encoding module as well as decoding module; the image encoding module comprises a down-sampling layer and an image encoding unit, (Li, paragraph 84 teaches “including a convolutional layer and a sub-sampling layer.”); this shows that sub-sampling/down-sampling layer would exist in encoder of Prajwal and one of ordinary skill in the art would understand that the face encoder itself would be image encoding unit; the decoding module comprises a decoding unit and an up- sampling layer, (Li, paragraph 205 teaches “vector calculation unit 2007 is mainly configured to perform network calculation at a non-convolutional/fully connected layer in a neural network, for example, batch normalization, pixel-level summation, and upsampling on a feature plane”); this shows layer for upsampling would exist in decoder of Prajwal and one of ordinary skill in the art would understand that the decoder itself would be a decoding unit; and performing feature encoding on the audio sequence by using the audio encoding module to obtain an audio feature (Prajwal, page 3, RHC, second to last paragraph teaches “a face encoder and an audio encoder,”); this shows feature of audio would be obtained using audio encoding. However, the combination of Prajwal and Li fails to explicitly teach and the video generation model is configured to generate the target video by following operations: extracting the initial image sequence from the talking video, and down-sampling the images in the initial image sequence by using the down-sampling layer to obtain the target image sequence; performing feature encoding on the target image sequence by using the image encoding unit to obtain an image feature; inputting the image feature and the audio feature into the decoding unit to generate the video frame sequence corresponding to the video of the digital human corresponding to the target object talking according to the target text; and up-sampling the video frames in the video frame sequence by using the up-sampling layer to obtain the target video. However, Kalva teaches and the video generation model is configured to generate the target video by following operations: extracting the initial image sequence from the talking video, and down-sampling the images in the initial image sequence by using the down-sampling layer to obtain the target image sequence (Kalva, paragraph 45 teaches “Encoder 208 may implement downsampling in the region transform module 260. For each picture and region of picture that is downsampled”); this shows the aforementioned extracting of image sequence (since for each picture) would be done by using down-sampling layer of encoder; performing feature encoding on the target image sequence by using the image encoding unit to obtain an image feature (Kalva, paragraph 7 teaches “video encoder receives the packed frame and coordinates of the regions of interest and encodes the packed frame”); this shows encoding which would be done by image encoding unit to obtain region of interest/image feature; inputting the image feature and the audio feature into the decoding unit to generate the video frame sequence corresponding to the video of the digital human corresponding to the target object talking according to the target text (Kalva, paragraph 4 teaches “decoder includes a video decoder which receives the encoded bitstream and decompreses the bitstream to extract and identify regions of interest and region parameters therefrom”); this shows input image feature to decoder unit to generate video frame sequence (and audio feature of video of digital human corresponding to target object talking according to target text from the above combination when viewed in combination); and up-sampling the video frames in the video frame sequence by using the up-sampling layer to obtain the target video (Kalva, paragraph 32 teaches “in order to maintain accuracy of the task algorithm, the spatial and temporal resolutions of the pictures need to be recovered. This is typically done at the decoder side by the technique of upsampling”); this shows decoder side performing up-sampling meaning the aforementioned up-sampling layer from decoder module would be used to recover pictures and obtain target video by upsampling frames. Kalva is considered to be analogous art because it is reasonably pertinent to the problem faced by the inventor of downsampling and upsampling using a encoder and decoder respectively. Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify the combination of Prajwal and Li with the superresolution techniques of Kalva to reduce the bits in the 232 compressed bitstream while also ensuring that 252 machine task system performance is improved or maintained (Kalva, paragraph 39). This would be done due to compression from encoding and then later decoding to keep consistent quality. Regarding claim 13, the device claim 13 recites similar limitations as method claim 5, and thus is rejected under similar rationale. Claim(s) 6-7 and 14-15 is/are rejected under 35 U.S.C. 103 as being unpatentable over the combination of Prajwal and Li as applied to claims 1 and 9 above, and further in view of Wang et al. (U.S. Patent Application Publication No. 20250045869), hereinafter referenced as Wang. Regarding claim 6, the combination of Prajwal and Li fails to teach wherein the video generation model is trained by following operations: obtaining a sample talking video of a target sample object, and extracting a sample audio sequence and a sample image sequence from the sample talking video, wherein images in the sample image sequence have a second resolution; processing a resolution of the images in the sample image sequence to a first resolution to obtain a first sample image sequence, and processing the first sample image sequence by using a super-resolution model to obtain a second sample image sequence, wherein the first resolution is higher than the second resolution; and performing model training on the video generation model according to the sample audio sequence and the second sample image sequence to obtain a trained video generation model. However, Wang teaches wherein the video generation model is trained by following operations: obtaining a sample talking video of a target sample object, and extracting a sample audio sequence and a sample image sequence from the sample talking video, wherein images in the sample image sequence have a second resolution (Wang, paragraph 99 teaches “Cache an i.sup.th frame of sample image and an (i+1).sup.th frame of sample image from a sample video” and paragraph 100 teaches “a training set is given for training a generative network, and the training set includes a sample video”); this shows talking video from above combination which has audio sequence would be sample talking video and it has sample image sequence extracted for training the video generation model, also the current resolution the images in sample image are at is considered the second resolution;; processing a resolution of the images in the sample image sequence to a first resolution to obtain a first sample image sequence, and processing the first sample image sequence by using a super-resolution model to obtain a second sample image sequence, wherein the first resolution is higher than the second resolution (Wang, paragraph 113 teaches “according to the method provided in this embodiment, the i.sup.th frame of sample image and the (i+1).sup.th frame of sample image in the sample video are cached, the super-resolution images are predicted by using the generative network, discrimination is performed on the super-resolution result by using the discriminative network,” and paragraph 126 teaches “a discriminative network 103 performs result discrimination based on a sample image and a super-resolution image of the sample image, and outputs a true (1) result or a false (0) result”); super-resolution shows processing resolution and it is of the images in sample image sequence (therefore this gives first sample image sequence), the resolution after super-resolution becomes first resolution which is higher than the previous second resolution and then that first sample image sequence is fed into the discriminative network which acts as super-resolution model processing the first sample image sequence to obtain the second sample image sequence; and performing model training on the video generation model according to the sample audio sequence and the second sample image sequence to obtain a trained video generation model (fig. 9 step 450 teaches to train the network based on results); this shows model training on video generation model using second image sequence obtained by discriminative result and this would be according to the sample audio sequence associated thereof as well when viewed in combination. Wang is considered to be analogous art because it is reasonably pertinent to the problem faced by the inventor of model training and super-resolution for sample images and videos. Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify the combination of Prajwal and Li with the super-resolution and training techniques of Wang to ensure more accurate generation results (Wang, paragraph 113). This would be due to the training. Regarding claim 7, the combination of Prajwal, Li and Wang teaches wherein the performing model training on the video generation model according to the sample audio sequence and the second sample image sequence to obtain the trained video generation model comprises: for each sample image in the second sample image sequence, determining a previous sample image of the sample image and a next sample image of the sample image in the second sample image sequence (Wang, claim 14 teaches “constraining super-resolution stability between adjacent frames of images” and claim 16 teaches “to obtain a discrimination result”); this would be for each sample image in second sample image sequence since it’s associated with discrimination result (which is aforementioned above to result in second image sequence) and includes next as well as previous sample image from such due to mention of adjacent frames; obtaining, by using the sample audio sequence, the sample image, the previous sample image, the next sample image, and the video generation model, a target sample image corresponding to the sample image, a previous target sample image corresponding to the previous sample image, and a next target sample image corresponding to the next sample image (Wang, paragraph 19 teaches “an inter-frame stability loss between: a stability loss of a change between adjacent frames of sample images and a stability loss of a change between super-resolution images of the adjacent frames of sample images, are calculated.” and paragraph 121 teaches “optical flows are used for measuring changes between adjacent image frames.”); this shows (by using/after aforementioned audio, previous, next and current sample images and video generation model from the combination above) target sample image corresponding to the sample image, a previous target sample image corresponding to the previous sample image, and a next target sample image corresponding to the next sample image being obtained due to the tracking of change between adjacent frames and having a stability loss calculated which means target must exist; determining an image jitter loss according to the target sample image, the previous target sample image, and the next target sample image (Wang, paragraph 107 teaches “in addition to several loss functions commonly used in a generative adversarial network, for example, an adversarial loss function, the present disclosure further provides an inter-frame stability loss function, which is configured for constraining a change between adjacent image frames.”); this stability/jitter loss is determined according to adjacent (meaning previous, next and current) target (since a target must be present for a constraining to occur as well as have the stability loss calculation) sample images; and adjusting a model parameter of the video generation model at least based on the image jitter loss to obtain the trained video generation model (Wang, paragraph 109 teaches “The corresponding error loss is calculated based on at least one loss function” and paragraph 110 teaches “Train the generative network and the discriminative network alternately based on the error loss”); this shows model being trained on (and obtained based on) the aforementioned image jitter/stability loss, and one of ordinary skill in the art would understand that training would adjust model parameter of the video generation model. Regarding claim 14, the device claim 14 recites similar limitations as method claim 6, and thus is rejected under similar rationale. Regarding claim 15, the device claim 15 recites similar limitations as method claim 7, and thus is rejected under similar rationale. Claim(s) 8 and 16 is/are rejected under 35 U.S.C. 103 as being unpatentable over the combination of Prajwal and Li as applied to claims 1 and 9 above, and further in view of Lin et al. (U.S. Patent Application Publication No. 2025/0150647), hereinafter referenced as Lin. Regarding claim 8, the combination of Prajwal and Li fails to teach wherein the obtaining the target text for video generation comprises: obtaining a live streaming content text for live streaming video generation; and the generating, by using the talking video, the target text, and the video generation model, the target video of the digital human corresponding to the target object talking according to the target text comprises: generating, by using the talking video, the live streaming content text, and the video generation model, a target live streaming video of the digital human corresponding to the target object talking according to the live streaming content text. However, Lin teaches wherein the obtaining the target text for video generation comprises: obtaining a live streaming content text for live streaming video generation (Lin, abstract teaches “obtaining a speech list from a speech server, the speech list including speech texts corresponding to business links of a same live streaming business process; generating a portrait oral video”); this shows speech text as live streaming content text for live streaming process/video generation; and the generating, by using the talking video, the target text, and the video generation model, the target video of the digital human corresponding to the target object talking according to the target text comprises: generating, by using the talking video, the live streaming content text, and the video generation model, a target live streaming video of the digital human corresponding to the target object talking according to the live streaming content text (Lin, abstract teaches “generating a portrait oral video corresponding to each of the speech texts based on a template video, the template video having a preset required duration, and each image frame of the template video comprising a facial image collected based on a same person; implementing a virtual live streaming activity by pushing the portrait oral video corresponding to each speech text to the e-commerce live streaming room,”); this shows the aforementioned generation would be by using template/talking video, the speech text (live stream content text) and video generation model from the above combination to generate target live streaming video with facial image of same person (digital human corresponding to target object talking) which would be according to the speech text (live streaming content text). Lin is considered to be analogous art because it is reasonably pertinent to the problem faced by the inventor of live streaming video generation using text and humans. Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify the combination of Prajwal and Li with the live streaming techniques of Lin to ensure whole process is relatively smooth, making the interaction between people presented in the virtual live streaming activity in the e-commerce live streaming room more realistic and natural, and may significantly improve the user experience of audience users participating in the e-commerce live streaming room (Lin, paragraph 71). This would be due to the speech text used to generate the live streaming video of the digital human which enhances the interaction. Regarding claim 16, the device claim 16 recites similar limitations as method claim 8, and thus is rejected under similar rationale. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. SIDDHARTHA et al. (U.S. Patent Application Publication No. 2025/0008193) abstract teaches “generating a target audio track and a target video track based on a source audio-video track are described …an audio generation model is used to generate a target audio for replacing specific portion of a source audio track… a video generation model is used to generate a target video for replacing specific portion of a source video track to generate a seamless target video track”; this shows source audio and video to generate target video. Any inquiry concerning this communication or earlier communications from the examiner should be directed to NAUMAN U AHMAD whose telephone number is (703)756-5306. The examiner can normally be reached Monday - Friday 9:00am - 5:00pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kee Tung can be reached at (571) 272-7794. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /KEE M TUNG/Supervisory Patent Examiner, Art Unit 2611 /N.U.A./Examiner, Art Unit 2611
Read full office action

Prosecution Timeline

Mar 19, 2025
Application Filed
Sep 23, 2026
Non-Final Rejection mailed — §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12722043
SYSTEMS AND METHODS FOR CAPTURING AND VISUALIZING THE SPINNING OF SMALL BALLS IN SPORTS
3y 2m to grant Granted Sep 01, 2026
Patent 12718620
EXTRACTION OF HUMAN POSES FROM VIDEO DATA FOR ANIMATION OF COMPUTER MODELS
3y 1m to grant Granted Aug 25, 2026
Patent 12718486
METHOD AND APPARATUS FOR GRAPHICS RENDERING USING A NEURAL PROCESSING UNIT
2y 7m to grant Granted Aug 25, 2026
Patent 12705803
PSEUDO VASCULAR PATTERN GENERATION DEVICE AND METHOD OF GENERATING PSEUDO VASCULAR PATTERN
2y 3m to grant Granted Aug 11, 2026
Patent 12700184
SYSTEM AND METHODS FOR REFINING ROOM SEGMENTS TO IMPROVE AESTHETIC QUALITY FOR END-USER APPLICATIONS
2y 2m to grant Granted Aug 04, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
74%
Grant Probability
91%
With Interview (+17.2%)
2y 6m (~12m remaining)
Median Time to Grant
Low
PTA Risk
Based on 49 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month