Prosecution Insights
Last updated: October 02, 2026
Application No. 18/870,926

DIGITAL HUMAN DRIVING METHOD, DIGITAL HUMAN DRIVING DEVICE AND STORAGE MEDIUM

Non-Final OA §102§103§112
Filed
Dec 02, 2024
Priority
May 30, 2022 — CN 202210599184.2 +1 more
Examiner
DU, HAIXIA
Art Unit
Tech Center
Assignee
ZTE Corporation
OA Round
1 (Non-Final)
87%
Grant Probability
Favorable
1-2
OA Rounds
5m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 87% — above average
87%
Career Allowance Rate
497 granted / 574 resolved
+26.6% vs TC avg
Strong +18% interview lift
Without
With
+17.8%
Interview Lift
resolved cases with interview
Typical timeline
2y 3m
Avg Prosecution
15 currently pending
Career history
587
Total Applications
across all art units

Statute-Specific Performance

§101
10.5%
-29.5% vs TC avg
§103
51.7%
+11.7% vs TC avg
§102
6.8%
-33.2% vs TC avg
§112
20.8%
-19.2% vs TC avg
Black line = Tech Center average estimate • Based on career data from 574 resolved cases

Office Action

§102 §103 §112
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . This is in response to the preliminary amendment filed on 1/13/20255. Claims 9, 10, and 11 have been amended. Claims 12-20 have been added. Claims 1-20 are present for examination. Claim Interpretation The term “computer storage medium”: Specification, p. 17, 2nd para., discloses computer-readable medium may include a computer storage medium (or non-transitory medium) and a communication medium (or transitory medium). Therefore, the term “computer storage medium” has been interpreted as non-transitory. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 8, 9, and 17-20 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Regarding claim 8, it depends from claim 1 and further recites in response to not receiving the image information and the audio information of the target object or in response to the determination result indicating that the image information and the audio information are invalid, acquiring a preset action sequence and performing feature extraction on the preset action sequence to obtain a third motion feature; and inputting the third motion feature and the digital human base image to the character generator, performing driving processing on the digital human base image through the character generator, and outputting the first digital human driving image. (Emphasis added.) However, claim 1 recites “performing feature extraction processing on the image information and/or the audio information according to the determination result to obtain a first motion feature and/or a second motion feature; inputting the first motion feature and/or the second motion feature and a digital human based image into a character generator.” (Emphasis added.) According to claim 1, it has to obtain a first motion feature and/or a second motion feature and input the first motion feature and/or the second motion feature into a character generator. However, claim 8 recites the image information and the audio information are not received or determined to be invalid. Therefore, no first motion feature or second motion feature will be obtained. Because claim 8 incorporates all the limitations of claim 1, it is not clear where the first motion feature and/or second motion feature recited in claim 1 comes from. Therefore, the first motion feature and the second motion feature are indefinite in claim 8. Claim 20 depends from claim 8 but fail to cure the deficiencies of claim 8. Regarding claim 9, it recites “performing interpolation processing according to a motion feature of the first digital human driving image and a motion feature of the second digital human driving image to obtain a transitional digital human driving image.” However, claim 1, from which claim 9 depends, recites “inputting the first motion feature and/or the second motion feature and a digital human based image into a character generator, and performing driving processing on the digital human base image through the character generator, and outputting a first digital human driving image.” It is not clear whether the motion feature of the first digital human driving image recited in claim 9 are the same as or different than the first motion feature or the second motion feature recited in claim 1, which are used as input by the character generator to output the first digital human driving image. Claims 17-20 respectively recite similar limitations discussed above with respect to claim 9. Therefore, claims 8, 9, and 17-20 are rejected under 355 USC 112(b) as being indefinite. For examination purposes, claim 8 has been interpreted as the image information and the audio information of the target object have been received but invalid, and the first and second motion features are obtained from the invalid image and audio information. Claims 9 and 17-20 have been interpreted as the motion feature of the first digital human driving imagen as the first motion feature. Claim Rejections - 35 USC § 102 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. Claim(s) 1, 10, and 11 is/are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Chinese Patent Publication No. CN113886641A to Wang et al. Regarding claim 1, Wang discloses A digital human driving method (Wang, Translation, para. [n0001], disclosing a method, apparatus, device and medium for generating digital humans), comprising: acquiring image information and audio information of a target object (Wang, Translation, para. [n0007], disclosing the target audio is input into the pre-trained first generator, para. [n0008], disclosing extract the target 3D face reconstruction parameters of the person from the target image, process the target image to obtain the first intermediate image that does not include the mouth area of the person, para. [n0089], disclosing images or videos of people can be acquired using image acquisition devices, para. [n0118], disclosing extracting several image frames and corresponding audio frames from the video stream, para. [n0119], disclosing the video stream can be actual video stream to be processed, such as a video stream recorded by the user, indicating the video stream can include the image information and audio information of a target object (the person) acquired by user recording); performing recognition and determination on the image information and the audio information to obtain a determination result (Wang, Translation, para. [n0113], disclosing using a piece of audio to finally generate a digital human video, para. [n0114], disclosing each audio frame in the audio segment can be identified as a target audio, and performing steps for each target audio to obtain a digital human image corresponding to each target audio, according to the time sequence of the target audio, the digital human images corresponding to the target audio are combined to generate a digital human video, para. [n0115], disclosing inputting the target audio into a pre-trained first generator to obtain the target facial expression parameters of the person, extracting the target 3D face reconstruction parameters of the person from the target image, para. [n0118], disclosing extracting several image frames and corresponding audio frames from the video stream, indicating the extracting image and audio frames can correspond performing recognition and determination on the image information and the audio information in the video stream to obtain a determination result of extracted image and audio frames); performing feature extraction processing on the image information and/or the audio information according to the determination result to obtain a first motion feature and/or a second motion feature (Wang, Translation, para. [n0084], disclosing input the target audio into the pre-trained first generator to obtain the target facial expression parameters of the character, para. [n0086], disclosing the facial expression parameters of a person rep[resent the opening and closing state of the person’s mouth, para. [n0088], disclosing extract the target 3D face reconstruction parameters from the target image, para. [n0090], disclosing 3D face reconstruction parameters include face shape information, reflection information, texture information, and lighting information, para. [n0114], disclosing each audio frame in the audio segment can be identified as a target audio, and performing steps for each target audio to obtain a digital human image corresponding to each target audio, according to the time sequence of the target audio, the digital human images corresponding to the target audio are combined to generate a digital human video, para. [n0115], disclosing inputting the target audio into a pre-trained first generator to obtain the target facial expression parameters of the person, extracting the target 3D face reconstruction parameters of the person from the target image, para. [n0118], disclosing extracting several image frames and corresponding audio frames from the video stream, indicating the target facial expression parameters can correspond to the first motion feature, and 3D face reconstruction parameters can correspond to the second motion feature, which are obtained by performing feature extraction on the image information and/or the audio information according to the determination result of extracted image and audio frames); inputting the first motion feature and/or the second motion feature and a digital human base image into a character generator (Wang, Translation, para. [n0092], disclosing process the target image to obtain a first intermediate image that does not include the mouth area of the person, para. [n0094], disclosing determine the target mouth region information based on the target expression parameters and the target 3D face reconstruction parameters, para. [n0098], disclosing inputting the target expression parameters and the target 3D face reconstruction parameters into a preset face 3D deformation statistical model (such as 3DMM) to obtain several target mouth region key points, and determining the several target mouth region key points as target mouth region information, para. [n0100], disclosing input the target mouth region information and the first intermediate image into the pre-trained second generator to obtain the digital human image, indicating the preset face 3D deformation statistical model (such as 3DMM) and the pre-trained second generator can correspond to the character generator, the first intermediate image can correspond to the digital human base image); and performing driving processing on the digital human base image through the character generator, and outputting a first digital human driving image (Wang, Translation, para. [n0092], disclosing process the target image to obtain a first intermediate image that does not include the mouth area of the person, para. [n0094], disclosing determine the target mouth region information based on the target expression parameters and the target 3D face reconstruction parameters, para. [n0098], disclosing inputting the target expression parameters and the target 3D face reconstruction parameters into a preset face 3D deformation statistical model (such as 3DMM) to obtain several target mouth region key points, and determining the several target mouth region key points as target mouth region information, para. [n0100], disclosing input the target mouth region information and the first intermediate image into the pre-trained second generator to obtain the digital human image, indicating the preset face 3D deformation statistical model (such as 3DMM) and the pre-trained second generator can correspond to the character generator, the first intermediate image can correspond to the digital human base image, and the character generator (the preset face 3D deformation statistical model (such as 3DMM) and the pre-trained second generator) performs driving processing on the digital human base image (the first intermediate image) and outputs a first digital human driving image (the digital human image)). Regarding claim 10, it recites similar limitations of claim 1 but in a device form. The rationale of claim 1 rejection is applied to reject claim 10. In addition, Wang discloses a memory, a processor, and a computer program stored in the memory and executable by the processor, wherein the computer program, when executed by the processor, causes the processor to perform a digital human driving method (Wang, Translation, para. [n0070], disclosing an electronic device including a processor, a memory, the memory is used to store computer programs, and the processor, when executing the program stored in the memory, implements the steps of the digital human generation method). Regarding claim 11, it recites similar limitations of claim 1 but in a computer storage medium form. The rationale of claim 1 rejection is applied to reject claim 11. In addition, Wang discloses computer storage medium, storing computer-executable instructions which, when executed by a processor, cause the processor to perform a digital human driving method (Wang, Translation, para. [n0070], disclosing an electronic device including a processor, a memory, the memory is used to store computer programs, and the processor, when executing the program stored in the memory, implements the steps of the digital human generation method). Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 2, 4-7, 12, and 14-16 is/are rejected under 35 U.S.C. 103 as being unpatentable over Wang in view of US Patent Publication No. 20080320158 A1 to Simonds. Regarding claim 2, Wang discloses the digital human driving method of claim 1, wherein performing feature extraction processing on the image information and/or the audio information according to the determination result to obtain a first motion feature and/or a second motion feature comprises: respectively performing feature extraction processing on the image information and the audio information to obtain the first motion feature and the second motion feature located in the same feature space as the first motion feature (Wang, Translation, para. [n0115], disclosing inputting the target audio into a pre-trained first generator to obtain the target facial expression parameters of the person, extracting the target 3D face reconstruction parameters of the person from the target image, para. [n0118], disclosing extracting several image frames and corresponding audio frames from the video stream, para. [n0143], disclosing the mouth region information determination module can determine the target mouth region information based on the target expression parameters and the target 3D face reconstruction parameters, para. [n0145], disclosing the target expression parameters and the target 3D face reconstruction parameters are input into a preset 3D face deformation statistical model to obtain the target 3D face mesh, indicating the target expression parameters and the target 3D face reconstruction parameters as the first motion feature and the second motion feature can be respectively extracted on the audio information and the image information and they are located in the same feature space for the preset 3D face deformation statistical model). However, Wang does not expressly disclose the feature extraction processing is in response to the determination result indicating that the image information and the audio information are valid. On the other hand, Simonds discloses providing the audio frame and the video frame in response to the determination result indicating that the image information and the audio information are valid (Simonds, para. [0036], disclosing determining whether the audio frame is valid or not, and if the audio stream is valid, providing the audio stream to the server, para. [0037], disclosing determining whether the video frame is valid or not, and if the video frame is valid, para. [0038], disclosing providing valid video frame to the server). Combining Wang and Simonds would allow checking whether the audio information and the image information corresponding to the video frame are valid or not, and in response to the determinization result indicating the image information and the audio information are valid, performing feature extraction processing on the image information and/or the audio information according to the determination result. Before the invention was effectively filed, it would have been obvious for a person skilled in the art to combine Wang and Simonds. The suggestion/motivation would have been to selectively streaming multimedia content, as suggested by Simonds (see Simonds, para. [0001]). Regarding claim 4, Wang discloses the digital human driving method of claim 1, wherein performing feature extraction processing on the image information and/or the audio information according to the determination result to obtain a first motion feature and/or a second motion feature comprises: performing feature extraction processing on the image information to obtain the first motion feature (Wang, Translation, para. [n0115], disclosing inputting the target audio into a pre-trained first generator to obtain the target facial expression parameters of the person, extracting the target 3D face reconstruction parameters of the person from the target image, para. [n0118], disclosing extracting several image frames and corresponding audio frames from the video stream, para. [n0143], disclosing the mouth region information determination module can determine the target mouth region information based on the target expression parameters and the target 3D face reconstruction parameters, para. [n0145], disclosing the target expression parameters and the target 3D face reconstruction parameters are input into a preset 3D face deformation statistical model to obtain the target 3D face mesh, indicating the target 3D face reconstruction parameters as the first motion feature can be respectively extracted on the image information). However, Wang does not expressly disclose the feature extraction processing is in response to the determination result indicating that the image information is valid and the audio information is invalid. On the other hand, Simonds discloses providing the video frame in response to the determination result indicating that the image information is valid and the audio information is invalid (Simonds, para. [0037], disclosing determining whether the video frame is valid or not, and if the video frame is valid, para. [0038], disclosing providing valid video frame to the server, para. [0039], disclosing it may only providing a video frame sequence). Combining Wang and Simonds would allow checking whether the audio information and the image information corresponding to the video frame are valid or not, and in response to the determinization result indicating that the image information is valid and the audio information is invalid, performing feature extraction processing on the image information to obtain the first motion feature. Before the invention was effectively filed, it would have been obvious for a person skilled in the art to combine Wang and Simonds. The suggestion/motivation would have been to selectively streaming multimedia content, as suggested by Simonds (see Simonds, para. [0001]). Regarding claim 5, Wang in view of Simonds discloses the digital human driving method of claim 4, wherein inputting the first motion feature and/or the second motion feature and a digital human base image into a character generator further comprises: inputting the first motion feature and the digital human base image to the character generator (Wang, Translation, para. [n0092], disclosing process the target image to obtain a first intermediate image that does not include the mouth area of the person, para. [n0094], disclosing determine the target mouth region information based on the target expression parameters and the target 3D face reconstruction parameters, para. [n0098], disclosing inputting the target expression parameters and the target 3D face reconstruction parameters into a preset face 3D deformation statistical model (such as 3DMM) to obtain several target mouth region key points, and determining the several target mouth region key points as target mouth region information, para. [n0100], disclosing input the target mouth region information and the first intermediate image into the pre-trained second generator to obtain the digital human image, indicating the preset face 3D deformation statistical model (such as 3DMM) and the pre-trained second generator can correspond to the character generator, the first intermediate image can correspond to the digital human base image, and the target 3D face reconstruction parameters can correspond to the first motion feature). Regarding claim 6, Wang discloses the digital human driving method of claim 1, wherein performing feature extraction processing on the image information and/or the audio information according to the determination result to obtain a first motion feature and/or a second motion feature comprises: performing feature extraction processing on the audio information to obtain the second motion feature (Wang, Translation, para. [n0115], disclosing inputting the target audio into a pre-trained first generator to obtain the target facial expression parameters of the person, extracting the target 3D face reconstruction parameters of the person from the target image, para. [n0118], disclosing extracting several image frames and corresponding audio frames from the video stream, para. [n0143], disclosing the mouth region information determination module can determine the target mouth region information based on the target expression parameters and the target 3D face reconstruction parameters, para. [n0145], disclosing the target expression parameters and the target 3D face reconstruction parameters are input into a preset 3D face deformation statistical model to obtain the target 3D face mesh, indicating the target facial expression parameters as the second motion feature can be respectively extracted on the audio information). However, Wang does not expressly disclose the feature extraction processing is in response to the determination result indicating that the image information is invalid and the audio information is valid. On the other hand, Simonds discloses providing the audio frame in response to the determination result indicating that the image information is invalid and the audio information is valid (Simonds, para. [0036], disclosing determining whether the audio frame is valid or not, and if the audio stream is valid, providing the audio stream to the server, para. [0037], disclosing determining whether the video frame is valid or not, and if the video frame is valid, para. [0038], disclosing providing valid video frame to the server, para. [0039], disclosing it may only providing an audio frame sequence). Combining Wang and Simonds would allow checking whether the audio information and the image information corresponding to the video frame are valid or not, and in response to the determination result indicating that the image information is invalid and the audio information is valid, performing feature extraction processing on the audio information to obtain the second motion feature. Before the invention was effectively filed, it would have been obvious for a person skilled in the art to combine Wang and Simonds. The suggestion/motivation would have been to selectively streaming multimedia content, as suggested by Simonds (see Simonds, para. [0001]). Regarding claim 7, Wang in view of Simonds discloses the digital human driving method of claim 6, wherein inputting the first motion feature and/or the second motion feature and a digital human base image into a character generator further comprises: inputting the second motion feature and the digital human base image to the character generator (Wang, Translation, para. [n0092], disclosing process the target image to obtain a first intermediate image that does not include the mouth area of the person, para. [n0094], disclosing determine the target mouth region information based on the target expression parameters and the target 3D face reconstruction parameters, para. [n0098], disclosing inputting the target expression parameters and the target 3D face reconstruction parameters into a preset face 3D deformation statistical model (such as 3DMM) to obtain several target mouth region key points, and determining the several target mouth region key points as target mouth region information, para. [n0100], disclosing input the target mouth region information and the first intermediate image into the pre-trained second generator to obtain the digital human image, indicating the preset face 3D deformation statistical model (such as 3DMM) and the pre-trained second generator can correspond to the character generator, the first intermediate image can correspond to the digital human base image, and the target expression parameters can correspond to the second motion feature). Regarding claim 12, it recites similar limitations of claim 2 but in a device form. The rationale of claim 2 rejection is applied to reject claim 12. In addition, Wang discloses a memory, a processor, and a computer program stored in the memory and executable by the processor, wherein the computer program, when executed by the processor, causes the processor to perform a digital human driving method (Wang, Translation, para. [n0070], disclosing an electronic device including a processor, a memory, the memory is used to store computer programs, and the processor, when executing the program stored in the memory, implements the steps of the digital human generation method). Regarding claim 14, it recites similar limitations of claim 4 but in a device form. The rationale of claim 4 rejection is applied to reject claim 14. In addition, Wang discloses a memory, a processor, and a computer program stored in the memory and executable by the processor, wherein the computer program, when executed by the processor, causes the processor to perform a digital human driving method (Wang, Translation, para. [n0070], disclosing an electronic device including a processor, a memory, the memory is used to store computer programs, and the processor, when executing the program stored in the memory, implements the steps of the digital human generation method). Regarding claim 15, it recites similar limitations of claim 5 but in a device form. The rationale of claim 5 rejection is applied to reject claim 15. In addition, Wang discloses a memory, a processor, and a computer program stored in the memory and executable by the processor, wherein the computer program, when executed by the processor, causes the processor to perform a digital human driving method (Wang, Translation, para. [n0070], disclosing an electronic device including a processor, a memory, the memory is used to store computer programs, and the processor, when executing the program stored in the memory, implements the steps of the digital human generation method). Regarding claim 16, it recites similar limitations of claim 6 but in a device form. The rationale of claim 6 rejection is applied to reject claim 16. In addition, Wang discloses a memory, a processor, and a computer program stored in the memory and executable by the processor, wherein the computer program, when executed by the processor, causes the processor to perform a digital human driving method (Wang, Translation, para. [n0070], disclosing an electronic device including a processor, a memory, the memory is used to store computer programs, and the processor, when executing the program stored in the memory, implements the steps of the digital human generation method). Claim(s) 3 and 13 is/are rejected under 35 U.S.C. 103 as being unpatentable over Wang in view of Simonds as applied to claims 2 above, and further in view of Chinese Patent Publication No. CN 112465935 A to Li et al. Regarding claim 3, Wang in view of Simonds discloses the digital human driving method of claim 2. However, Wang or Simonds does not expressly disclose wherein inputting the first motion feature and/or the second motion feature and a digital human base image into a character generator further comprises: performing feature fusion processing according to a preset weighted fusion coefficient, the first motion feature, and the second motion feature to obtain a fused motion feature; and inputting the fused motion feature and the digital human base image to the character generator. On the other hand Li discloses performing feature fusion processing according to a preset weighted fusion coefficient, the first motion feature, and the second motion feature to obtain a fused motion feature (Li, Translation, para. [n0046], disclosing speech features are obtained by extracting speech features from speech data, para. [n0048], disclosing the facial expression features can be extracted from video data containing the speaker’s face, para. [n0052], disclosing fusion of the speech features and facial expression features and generating virtual avatar videos based on the fused features through a pre-trained avatar synthesis model, para. [n0055], disclosing the speech features and the facial expressly features are weighted and fused based on the fusion weights, indicating the fusion weights can be a preset weighted fusion coefficient, the facial expression features and the speech features can correspond to the first motion feature and the second motion feature, a feature fusion processing is performed according to the fusion weights as the preset weighted fusion coefficient, the facial expression features and the speech features as the first and second motion features to obtain the fused motion feature); and inputting the fused motion feature and the digital human base image to the character generator (Li, Translation, para. [n0052], disclosing fusion of the speech features and facial expression features and generating virtual avatar videos based on the fused features through a pre-trained avatar synthesis model, para. [n0078], disclosing the image synthesis model can achieve weighted fusion of speech features and facial expression features, and on this basis, combine virtual image mask images to synthesize virtual image videos, indicating the fusion of the speech features and facial expression features as the fused motion feature and the virtual image mask images can correspond to the digital human base image, and they are input to the pre-trained avatar synthesis model). Before the invention was effectively filed, it would have been obvious for a person skilled in the art to combine Wang in view of Simonds with Li. The suggestion/motivation would have been ensure the expressions of the virtual avatar can naturally match the video data and optimize the user experience of virtual avatar synthesis, as suggested by Li (see Li, Translation, para. [n0032]). Regarding claim 13, it recites similar limitations of claim 3 but in a device form. The rationale of claim 3 rejection is applied to reject claim 13. In addition, Wang discloses a memory, a processor, and a computer program stored in the memory and executable by the processor, wherein the computer program, when executed by the processor, causes the processor to perform a digital human driving method (Wang, Translation, para. [n0070], disclosing an electronic device including a processor, a memory, the memory is used to store computer programs, and the processor, when executing the program stored in the memory, implements the steps of the digital human generation method). Allowable Subject Matter Claims 8, 9, and 17-20 would be allowable if rewritten to overcome the rejection(s) under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), 2nd paragraph, set forth in this Office action and to include all of the limitations of the base claim and any intervening claims. Regarding claim 8, none of the prior art references on the record, alone or in combination, discloses in response to not receiving the image information and the audio information of the target object or in response to the determination result indicating that the image information and the audio information are invalid, acquiring a preset action sequence and performing feature extraction on the preset action sequence to obtain a third motion feature; and inputting the third motion feature and the digital human base image to the character generator, performing driving processing on the digital human base image through the character generator, and outputting the first digital human driving image. Regarding claim 9, none of the prior art references on the record, alone or in combination, discloses determining first driving modality information according to the first digital human driving image; determining second driving modality information according to a second digital human driving image, wherein the second digital human driving image is a frame image previous to the first digital human driving image; and in response to the first driving modality information being different from the second driving modality information, performing interpolation processing according to a motion feature of the first digital human driving image and a motion feature of the second digital human driving image to obtain a transitional digital human driving image. Claims 17-20 respectively recite similar limitations discussed above with respect to claim 9. Claim 20 also depends from claim 8. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. US Patent Publication No. 20220374625 A1 to Chu et al., which discloses image transformation and retrieval by interpolating the features between the first image and second image. Any inquiry concerning this communication or earlier communications from the examiner should be directed to HAIXIA DU whose telephone number is (571)270-5646. The examiner can normally be reached Monday - Friday 8:00 am-4:00 pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kee Tung can be reached at 571-272-7794. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /HAIXIA DU/Primary Examiner, Art Unit 2611
Read full office action

Prosecution Timeline

Dec 02, 2024
Application Filed
Aug 11, 2026
Non-Final Rejection mailed — §102, §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12743745
GENERATING DIGITAL IMAGES UTILIZING A DIFFUSION-BASED NETWORK CONDITIONED ON LIGHTING-AWARE FEATURE REPRESENTATIONS
2y 5m to grant Granted Sep 22, 2026
Patent 12737938
GRAPHICS TEXTURE PROCESSING
2y 11m to grant Granted Sep 15, 2026
Patent 12737953
ANIMATION EFFECT DISPLAY METHOD AND ELECTRONIC DEVICE
2y 5m to grant Granted Sep 15, 2026
Patent 12718339
MEDICAL IMAGING
2y 4m to grant Granted Aug 25, 2026
Patent 12718471
UNDER-DISPLAY ARRAY CAMERA PROCESSING FOR THREE-DIMENSIONAL (3D) SCENES
1y 10m to grant Granted Aug 25, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
87%
Grant Probability
99%
With Interview (+17.8%)
2y 3m (~5m remaining)
Median Time to Grant
Low
PTA Risk
Based on 574 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month