DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Arguments
The reply filed on 7/06/2026 has been entered. Applicant’s arguments with respect to claims 1, 2, 3, 4, 5,6, 7, 8, 9, 10, 11, 12, 15, 16, 19, 20, 21, 22, 23, and 24 have been considered but are moot in view of new ground(s) of rejection caused by the amendments.
Claims 1, 2, 3, 4, 5,6, 7, 8, 9, 10, 11, 12, 15, 16, 19, 20, 21, 22, 23, and 24 are pending in this application and have been considered below.
Priority
Receipt is acknowledged that application is a National Stage application of PCT/SG2023/050004. Priority to CN202210001031.3 with a priority date of 1/04/2022 is acknowledged under 35 USC 119(e) and 37 CFR 1.78.
Information Disclosure Statement
The IDS dated 7/03/2024 has been considered and placed in the application file.
1st Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 1, 6, 15, and 16 are rejected under 35 U.S.C. 103 as obvious over US Patent Publication 2022 0245961 A1, (Yan et al.) in view of US Patent Publication 2008 0187174 A1, (Metaxas et al.).
Claim 1
Regarding claim 1, Yan et al. disclose an expression driving method, comprising: acquiring a first video; ("The terminal device uploads the acquired facial image or video of the virtual object to the server," par. 51) and inputting the first video into a pre-trained expression driving model to obtain a second video; ("inputs the facial image of the virtual object and the video of the real person into the expression transfer model … Output, by the expression transfer model, a synthesized facial image or a synthesized facial video," par. 53-54) wherein the expression driving model is trained based on a third sample image and a plurality of first sample images, ("during the training process, a set of samples includes the source domain facial image, the target domain facial image and the facial feature image belonging to the same real person," par. 69) wherein a facial image in the second video is generated ("the synthesized facial images are outputted based on the expression transfer model," par. 192) based on the third sample image, ("a set of samples includes the source domain facial image, the target domain facial image," par. 69) wherein a gesture expression feature of the facial image in the second video is the same as a gesture expression feature of a facial image in the first video ("the face in the synthesized facial image to have the same size and position as the face in the target domain facial image," par. 136) and wherein the gesture expression feature comprises an expression ("It can be seen that, images having similar expressions with real human faces can be obtained by using the expression transfer model provided by the present disclosure, which not only maintains the real facial features, but also fully integrates styles of the virtual objects, thereby obtaining more vivid synthesized images," par. 191).
Yan et al. do not explicitly teach all of wherein the gesture expression feature comprises a gesture angle, and the gesture angle comprises at least one of a roll angle or a yaw angle.
However, Metaxas et al. teach wherein the gesture expression feature comprises a gesture angle and an expression, and the gesture angle comprises at least one of a roll angle or a yaw angle ("The present invention could be extended to estimate pose angles (e.g., pitch and yaw) of a subject, so as to improve tracking accuracy where there are large head rotations or discontinuities in shape spaces," par. 47).
Therefore, taking the teachings of Yan et al. and Metaxas et al. as a whole, it would have been obvious to a person having ordinary skill in the art before the time of the effective filing date of the claimed invention of the instant application to modify the expression transfer model as taught by Yan et al. to use pose angle distribution fitting as taught by Metaxas et al. The suggestion/motivation for doing so would have been that, “The present invention could be extended to estimate pose angles (e.g., pitch and yaw) of a subject, so as to improve tracking accuracy where there are large head rotations or discontinuities in shape spaces” as noted by the Metaxas et al. disclosure in paragraph [0047], which also motivates combination because the combination would predictably have a higher accuracy as there is a reasonable expectation that incorporating pose estimation into the expression transfer model would account for variations in head orientation, thereby reducing errors associated with tracking subject features during significant rotational movement; and/or because doing so merely combines prior art elements according to known methods to yield predictable results.
Claim 6
Regarding claim 6, Yan et al. and Metaxas et al. teach the method of claim 1 as noted above.
Yan et al. also teach wherein the plurality of first sample images are initial sample images ("frames of second images," par. 191).
Yan et al. do not explicitly teach all of in which a number of sample images for each gesture angle conforms to a predetermined distribution.
However, Metaxas et al. teach in which a number of sample images for each gesture angle conforms to a predetermined distribution ("The pose angle predictions are done by the function approximators F.sub.i(X) fitted locally to each cluster and are represented using Gaussian distribution," par. 47).
Yan et al. and Metaxas et al. are combined as per claim 1.
Claim 15
Regarding claim 15, Yan et al. disclose an electronic device, comprising: a processor and a memory communicatively connected with the processor, wherein the memory stores computer-executable instructions; ("a memory, a processor, and a bus system, the bus system connecting the memory to the processor; the memory being configured to store a plurality of computer programs," par. 12-13) and wherein the computer-executable instructions, upon execution of the processor, cause the processor to: ("the processor being configured to execute the plurality of computer programs," par. 14) acquire a first video; ("The terminal device uploads the acquired facial image or video of the virtual object to the server," par. 51) and input the first video into a pre-trained expression driving model to obtain a second video; ("inputs the facial image of the virtual object and the video of the real person into the expression transfer model … Output, by the expression transfer model, a synthesized facial image or a synthesized facial video," par. 53-54) wherein the expression driving model is trained based on a third sample image and a plurality of first sample images, ("during the training process, a set of samples includes the source domain facial image, the target domain facial image and the facial feature image belonging to the same real person," par. 69) wherein a facial image in the second video is generated ("the synthesized facial images are outputted based on the expression transfer model," par. 192) based on the third sample image, ("a set of samples includes the source domain facial image, the target domain facial image," par. 69) wherein a gesture expression feature of the facial image in the second video is the same as a gesture expression feature of a facial image in the first video ("the face in the synthesized facial image to have the same size and position as the face in the target domain facial image," par. 136) and wherein the gesture expression feature comprises an expression ("It can be seen that, images having similar expressions with real human faces can be obtained by using the expression transfer model provided by the present disclosure, which not only maintains the real facial features, but also fully integrates styles of the virtual objects, thereby obtaining more vivid synthesized images," par. 191).
Yan et al. do not explicitly teach all of wherein the gesture expression feature comprises a gesture angle, and the gesture angle comprises at least one of a roll angle or a yaw angle.
However, Metaxas et al. teach wherein the gesture expression feature comprises a gesture angle, and the gesture angle comprises at least one of a roll angle or a yaw angle ("The present invention could be extended to estimate pose angles (e.g., pitch and yaw) of a subject, so as to improve tracking accuracy where there are large head rotations or discontinuities in shape spaces," par. 47).
Yan et al. and Metaxas et al. are combined as per claim 1.
Claim 16
Regarding claim 16, Yan et al. and Metaxas et al. teach the steps of the expression driving method according to claim 1 as noted above.
Yan et al. also disclose a non-transitory computer-readable storage medium with computer-executable instructions stored thereon, wherein the computer-executable instructions, when executed by a processor, cause the processor to implement ("the embodiments of the present disclosure provide a non-transitory computer readable storage medium, storing a plurality of computer programs. The computer programs are configured for performing the aforementioned method for training an expression transfer model," par. 15).
Yan et al. and Metaxas et al. are combined as per claim 1.
2nd Claim Rejections - 35 USC § 103
Claims 2 and 19 are rejected under 35 U.S.C. 103 as obvious over US Patent Publication 2022 0245961 A1, (Yan et al.) and US Patent Publication 2008 0187174 A1, (Metaxas et al.) in view of US Patent Publication 2023 0126806 A1, (Kim et al.).
Claim 2
Regarding claim 2, Yan et al. and Metaxas et al. teach the method of claim 1, wherein the expression driving model is trained based on a plurality of sample image pairs determined based on the plurality of first sample images and corresponding second sample images; ("the facial feature image A and the first image are stitched and then inputted into the trained expression transfer model … The facial feature image B and the first image are stitched and then inputted into the trained expression transfer model," par. 178) a second sample image is derived based on a plurality of second facial keypoints in the third sample image ("the keypoints are extracted from the target domain facial image, to obtain a corresponding facial feature image," par. 69) and a plurality of first facial keypoints in a corresponding first sample image; ("The facial feature image set includes P facial feature images, and the facial feature images are in a one-to-one correspondence with the second images," par. 170) and a similarity between a gesture expression feature of a facial image in the second sample image and a gesture expression feature of a facial image in the corresponding first sample image ("Obtain, based on the synthesized facial image and the target domain facial image and by a discriminative network model, a first discrimination result," par. 75).
Yan et al. do not explicitly teach all of a similarity is greater than a value.
However, Kim et al. teach a similarity is greater than a value ("similarity scores being greater than the first threshold value," par. 26).
Therefore, taking the teachings of Yan et al., Metaxas et al., and Kim et al. as a whole, it would have been obvious to a person having ordinary skill in the art before the time of the effective filing date of the claimed invention of the instant application to modify the expression transfer model as taught by Yan et al. and pose angle distribution fitting as taught by Metaxas et al. to use facial authentication techniques as taught by Kim et al. The suggestion/motivation for doing so would have been that, “in response to an average value of the plurality of similarity scores being greater than the first threshold value, determining that the face authentication may be successful” as noted by the Kim et al. disclosure in paragraph [0025], which also motivates combination because the combination would predictably have a higher accuracy as there is a reasonable expectation that utilizing a threshold-based similarity score provides a reliable, objective metric for verifying identity, thereby ensuring that the combined system meets a predetermined, expected level of authentication precision; and/or because doing so merely combines prior art elements according to known methods to yield predictable results.
Claim 19
Regarding claim 19, Yan et al. and Metaxas et al. teach the electronic device of claim 15 as noted above.
Yan et al. teach wherein the expression driving model is trained based on a plurality of sample image pairs determined based on the plurality of first sample images and corresponding second sample images; ("the facial feature image A and the first image are stitched and then inputted into the trained expression transfer model … The facial feature image B and the first image are stitched and then inputted into the trained expression transfer model," par. 178) a second sample image is derived based on a plurality of second facial keypoints in the third sample image ("the keypoints are extracted from the target domain facial image, to obtain a corresponding facial feature image," par. 69) and a plurality of first facial keypoints in a corresponding first sample image; ("The facial feature image set includes P facial feature images, and the facial feature images are in a one-to-one correspondence with the second images," par. 170) and a similarity between a gesture expression feature of a facial image in the second sample image and a gesture expression feature of a facial image in the corresponding first sample image ("Obtain, based on the synthesized facial image and the target domain facial image and by a discriminative network model, a first discrimination result," par. 75).
Yan et al. do not explicitly teach all of a similarity is greater than a value.
However, Kim et al. teach a similarity is greater than a value ("similarity scores being greater than the first threshold value," par. 26).
Yan et al., Metaxas et al., and Kim et al. are combined as per claim 2.
3rd Claim Rejections - 35 USC § 103
Claims 3, 4, 5, 7, 8, 9, 10, 11, 12, 20, 21, 22, 23, and 24 are rejected under 35 U.S.C. 103 as obvious over US Patent Publication 2022 0245961 A1, (Yan et al.), US Patent Publication 2008 0187174 A1, (Metaxas et al.), and US Patent Publication 2023 0126806 A1, (Kim et al.) in view of US Patent Publication 2023 0281833 A1, (Zhu et al.).
Claim 3
Regarding claim 3, Yan et al., Metaxas et al., and Kim et al. teach the method of claim 2 as noted above.
Yan et al. do not explicitly teach all of wherein the second sample image is obtained based on displacement information between the plurality of second facial keypoints and the plurality of first facial keypoints and a corresponding facial feature map of the third sample image; for each second facial keypoint, the displacement information is displacement information between the second facial keypoint and a corresponding first facial keypoint; and the facial feature map is obtained by encoding facial information of the third sample image.
However, Zhu et al. teach wherein the second sample image is obtained based on displacement information between the plurality of second facial keypoints and the plurality of first facial keypoints and a corresponding facial feature map of the third sample image; ("The computer device generates, based on the first facial sample image and the second optical flow information, a predicted image of the first facial sample image through a generator of the image processing model," par. 53) for each second facial keypoint, the displacement information is displacement information between the second facial keypoint ("the second optical flow information representing offsets between a plurality of pixels in the first facial sample image and a plurality of corresponding pixels in the second facial sample image," par. 5) and a corresponding first facial keypoint; ("the first optical flow information representing offsets of a plurality of key points of a target object in a first facial sample image and a second facial sample image," par. 5) and the facial feature map is obtained by encoding facial information of the third sample image ("the generator extracts features of the first facial sample image to obtain a first intermediate feature map," par. 62).
Therefore, taking the teachings of Yan et al., Metaxas et al., Kim et al., and Zhu et al. as a whole, it would have been obvious to a person having ordinary skill in the art before the time of the effective filing date of the claimed invention of the instant application to modify the expression transfer model as taught by Yan et al., pose angle distribution fitting as taught by Metaxas et al., and facial authentication techniques as taught by Kim et al. to use the feature processing methods as taught by Zhu et al. The suggestion/motivation for doing so would have been that, “The facial image is driven through the image processing model, so that the facial image may reflect the same and dynamic expression changes as a driving image, thereby achieving the purpose of diversifying image processing effects, and the effects of facial expressions are more realistic” as noted by the Zhu et al. disclosure in paragraph [0142], which also motivates combination because the combination would predictably have a more realistic output as there is a reasonable expectation that combining fine-grained feature processing with facial authentication techniques improves the robustness and fidelity of expression transfer between subjects; and/or because doing so merely combines prior art elements according to known methods to yield predictable results.
Claim 4
Regarding claim 4, Yan et al., Metaxas et al., Kim et al., and Zhu et al. teach the method of claim 3 as noted above.
Yan et al. do not explicitly teach all of wherein the displacement information is determined according to difference information between the plurality of second facial keypoints and corresponding first facial keypoints, and a pre-trained network model.
However, Zhu et al. teach wherein the displacement information is determined according to difference information ("determining a difference between the plurality of key points in the second optical flow information and the first optical flow information," par. 100) between the plurality of second facial keypoints ("the first optical flow information representing offsets of a plurality of key points of a target object in a first facial sample image and a second facial sample image," par. 5) and corresponding first facial keypoints, ("the second optical flow information representing offsets between a plurality of pixels in the first facial sample image and a plurality of corresponding pixels in the second facial sample image," par. 5) and a pre-trained network model ("the optical flow information prediction sub-model 203 may be published as a trained image processing model," par. 42).
Yan et al., Metaxas et al., Kim et al., and Zhu et al. are combined as per claim 3.
Claim 5
Regarding claim 5, Yan et al., Metaxas et al., Kim et al., and Zhu et al. teach the method of claim 4 as noted above.
Yan et al. do not explicitly teach all of wherein the difference information is determined according to coordinate information of the second facial keypoint and coordinate information of the corresponding first facial keypoint under a same coordinate system.
However, Zhu et al. teach wherein the difference information is determined according to coordinate information of the second facial keypoint and coordinate information of the corresponding first facial keypoint ("For each key point in the first facial sample image, the second position of the key point in the second facial sample image is determined, and the second position of the key point is subtracted from the first position of the key point to obtain offset of the key point," par. 71) under a same coordinate system ("The first position and the second position are represented by coordinates in a same coordinate system, and the offset is represented in a vector form," par. 71).
Yan et al., Metaxas et al., Kim et al., and Zhu et al. are combined as per claim 3.
Claim 7
Regarding claim 7, Yan et al. teach a training method of an expression driving model, comprising: a plurality of first facial keypoints in each of a plurality of first sample images, respectively ("The facial feature image set includes P facial feature images, and the facial feature images are in a one-to-one correspondence with the second images," par. 170); wherein a similarity between a gesture expression feature of a facial image in the second sample image and a gesture expression feature of a facial image in the third sample image ("Obtain, based on the synthesized facial image and the target domain facial image and by a discriminative network model, a first discrimination result," par. 75); wherein the gesture expression feature comprises an expression; ("It can be seen that, images having similar expressions with real human faces can be obtained by using the expression transfer model provided by the present disclosure, which not only maintains the real facial features, but also fully integrates styles of the virtual objects, thereby obtaining more vivid synthesized images," par. 191) determining a plurality of sample image pairs according to the plurality of first sample images and corresponding second sample images; ("the facial feature image A and the first image are stitched and then inputted into the trained expression transfer model … The facial feature image B and the first image are stitched and then inputted into the trained expression transfer model," par. 178) and updating model parameters of an initial expression driving model according to the plurality of sample image pairs to obtain the expression driving model ("optimizes the model parameter of the expression transfer model to be trained based on a backpropagation algorithm. The expression transfer model is obtained when a model convergence condition is reached," par. 82).
Metaxas et al. teach wherein the gesture expression feature comprises a gesture angle, and the gesture angle comprises at least one of a roll angle or a yaw angle ("The present invention could be extended to estimate pose angles (e.g., pitch and yaw) of a subject, so as to improve tracking accuracy where there are large head rotations or discontinuities in shape spaces," par. 47).
Kim et al. teach a similarity is greater than a value ("similarity scores being greater than the first threshold value," par. 26).
Yan et al. do not explicitly teach all of extracting a plurality of second facial keypoints in a third sample image, determining, for each first sample image and each second facial keypoint, displacement information between a second facial keypoint and a first facial keypoint in the first sample image corresponding to the second facial keypoint; and generating a second sample image according to the displacement information and the third sample image.
However, Zhu et al. teach extracting a plurality of second facial keypoints in a third sample image, ("obtaining first optical flow information, the first optical flow information representing offsets of a plurality of key points of a target object in a first facial sample image," par. 5) determining, for each first sample image and each second facial keypoint, displacement information ("determining a difference between the plurality of key points in the second optical flow information and the first optical flow information," par. 100) between a second facial keypoint ("the first optical flow information representing offsets of a plurality of key points of a target object in a first facial sample image and a second facial sample image," par. 5) and a first facial keypoint in the first sample image corresponding to the second facial keypoint ("the second optical flow information representing offsets between a plurality of pixels in the first facial sample image and a plurality of corresponding pixels in the second facial sample image," par. 5); and generating a second sample image according to the displacement information and the third sample image ("The computer device generates, based on the first facial sample image and the second optical flow information, a predicted image of the first facial sample image through a generator of the image processing model," par. 53).
Yan et al., Metaxas et al., Kim et al., and Zhu et al. are combined as per claim 3.
Claim 8
Regarding claim 8, Yan et al., Metaxas et al., Kim et al., and Zhu et al. teach the method of claim 7, wherein the generating the second sample image according to the displacement information and the third sample image comprises, as noted above.
Yan et al. do not explicitly teach all of encoding facial information in the third sample image to obtain a facial feature map; and determining the second sample image according to the displacement information and the facial feature map.
However, Zhu et al. teach encoding facial information in the third sample image to obtain a facial feature map ("the generator extracts features of the first facial sample image to obtain a first intermediate feature map," par. 62); and determining the second sample image according to the displacement information ("The computer device generates, based on the first facial sample image and the second optical flow information, a predicted image of the first facial sample image through a generator of the image processing model," par. 53) and the facial feature map ("the generator extracts features of the first facial sample image to obtain a first intermediate feature map," par. 62).
Yan et al., Metaxas et al., Kim et al., and Zhu et al. are combined as per claim 3.
Claim 9
Regarding claim 9, Yan et al., Metaxas et al., Kim et al., and Zhu et al. teach the method of claim 8, wherein the determining the second sample image according to the displacement information and the facial feature map comprises, as noted above.
Yan et al. do not explicitly teach all of performing, according to the displacement information, bending transition processing and/or displacement processing on the facial feature map to obtain a processed facial feature map; and decoding the processed facial feature map to obtain the second sample image.
However, Zhu et al. teach performing, according to the displacement information, bending transition processing and/or displacement processing on the facial feature map to obtain a processed facial feature map ("offsets pixels in the first intermediate feature map through the generator based on the third optical flow information to obtain a second intermediate feature map," par. 82); and decoding the processed facial feature map to obtain the second sample image ("and upsamples the second intermediate feature map to obtain a predicted image of the first facial sample image," par. 82).
Yan et al., Metaxas et al., Kim et al., and Zhu et al. are combined as per claim 3.
Claim 10
Regarding claim 10, Yan et al., Metaxas et al., Kim et al., and Zhu et al. teach the method of claim 7, wherein the determining displacement information between the second facial keypoint and the first facial keypoint in the first sample image corresponding to the second facial keypoint comprises, as noted above.
Yan et al. do not explicitly teach all of determining difference information between the second facial keypoints and first facial keypoints in the first sample image corresponding to the second facial keypoints; and determining the displacement information according to the difference information and a pre- trained network model.
However, Zhu et al. teach determining difference information between ("determining a difference between the plurality of key points in the second optical flow information and the first optical flow information," par. 100) the second facial keypoints and ("the first optical flow information representing offsets of a plurality of key points of a target object in a first facial sample image and a second facial sample image," par. 5) first facial keypoints in the first sample image corresponding to the second facial keypoints ("the second optical flow information representing offsets between a plurality of pixels in the first facial sample image and a plurality of corresponding pixels in the second facial sample image," par. 5); and determining the displacement information according to the difference information ("The optical flow information prediction sub-model predicts optical flow information of the plurality of pixels of the first facial sample image," par. 100) and a pre- trained network model ("the optical flow information prediction sub-model 203 may be published as a trained image processing model," par. 42).
Yan et al., Metaxas et al., Kim et al., and Zhu et al. are combined as per claim 3.
Claim 11
Regarding claim 11, Yan et al., Metaxas et al., Kim et al., and Zhu et al. teach the method of claim 10, wherein the determining the difference information between the second facial keypoints and first facial keypoints in the first sample image corresponding to the second facial keypoints comprises, as noted above.
Yan et al. do not explicitly teach all of transforming the plurality of second facial keypoints and the plurality of first facial keypoints into a same coordinate system; and determining the difference information between a respective second facial keypoints and a corresponding first facial keypoints according to coordinate information of the respective second facial keypoints and the coordinate information of the corresponding first facial keypoints under the same coordinate system.
However, Zhu et al. teach transforming the plurality of second facial keypoints and the plurality of first facial keypoints into a same coordinate system ("The first position and the second position are represented by coordinates in a same coordinate system, and the offset is represented in a vector form," par. 71); and determining the difference information between a respective second facial keypoints and a corresponding first facial keypoints according to coordinate information of the respective second facial keypoints and the coordinate information of the corresponding first facial keypoints ("For each key point in the first facial sample image, the second position of the key point in the second facial sample image is determined, and the second position of the key point is subtracted from the first position of the key point to obtain offset of the key point," par. 71) under the same coordinate system ("The first position and the second position are represented by coordinates in a same coordinate system, and the offset is represented in a vector form," par. 71).
Yan et al., Metaxas et al., Kim et al., and Zhu et al. are combined as per claim 3.
Claim 12
Regarding claim 12, Yan et al., Metaxas et al., Kim et al., and Zhu et al. teach the method of claim 7 as noted above.
Yan et al. also teach acquiring a plurality of initial sample images ("Obtain a first image corresponding to a virtual object and a video material corresponding to a real person. The video material includes P frames of second images," par. 168) and determining initial sample images ("frames of second images," par. 191).
Yan et al. do not explicitly teach all of determining gesture angles of the plurality of initial sample images; and in which a number of sample images for each gesture angle conforms to a predetermined distribution as the plurality of first sample images.
However, Metaxas et al. teach determining gesture angles of the plurality of initial sample images ("provide estimates of pose angles of a subjects in images," par. 23); and in which a number of sample images for each gesture angle conforms to a predetermined distribution as the plurality of first sample images ("The pose angle predictions are done by the function approximators F.sub.i(X) fitted locally to each cluster and are represented using Gaussian distribution," par. 47).
Yan et al., Metaxas et al., Kim et al., and Zhu et al. are combined as per claim 3.
Claim 20
Regarding claim 20, Yan et al., Metaxas et al., Kim et al., and Zhu et al. teach the method of claim 19 as noted above.
Yan et al. do not explicitly teach all of wherein the second sample image is obtained based on displacement information between the plurality of second facial keypoints and the plurality of first facial keypoints and a corresponding facial feature map of the third sample image; for each second facial keypoint, the displacement information is displacement information between the second facial keypoint and a corresponding first facial keypoint; and the facial feature map is obtained by encoding facial information of the third sample image.
However, Zhu et al. teach wherein the second sample image is obtained based on displacement information between the plurality of second facial keypoints and the plurality of first facial keypoints and a corresponding facial feature map of the third sample image ("The computer device generates, based on the first facial sample image and the second optical flow information, a predicted image of the first facial sample image through a generator of the image processing model," par. 53); for each second facial keypoint, the displacement information is displacement information between the second facial keypoint ("the first optical flow information representing offsets of a plurality of key points of a target object in a first facial sample image and a second facial sample image," par. 5) and a corresponding first facial keypoint ("the second optical flow information representing offsets between a plurality of pixels in the first facial sample image and a plurality of corresponding pixels in the second facial sample image," par. 5); and the facial feature map is obtained by encoding facial information of the third sample image ("the generator extracts features of the first facial sample image to obtain a first intermediate feature map," par. 62).
Yan et al., Metaxas et al., Kim et al., and Zhu et al. are combined as per claim 3.
Claim 21
Regarding claim 21, Yan et al., Metaxas et al., Kim et al., and Zhu et al. teach the method of claim 20 as noted above.
Yan et al. do not explicitly teach all of wherein the displacement information is determined according to difference information between the plurality of second facial keypoints and corresponding first facial keypoints, and a pre-trained network model.
However, Zhu et al. teach wherein the displacement information is determined according to difference information between ("determining a difference between the plurality of key points in the second optical flow information and the first optical flow information," par. 100) the plurality of second facial keypoints ("the first optical flow information representing offsets of a plurality of key points of a target object in a first facial sample image and a second facial sample image," par. 5) and corresponding first facial keypoints, ("the second optical flow information representing offsets between a plurality of pixels in the first facial sample image and a plurality of corresponding pixels in the second facial sample image," par. 5) and a pre-trained network model ("the optical flow information prediction sub-model 203 may be published as a trained image processing model," par. 42).
Yan et al., Metaxas et al., Kim et al., and Zhu et al. are combined as per claim 3.
Claim 22
Regarding claim 22, Yan et al., Metaxas et al., Kim et al., and Zhu et al. teach the method of claim 21 as noted above.
Yan et al. do not explicitly teach all of wherein the difference information is determined according to coordinate information of the second facial keypoint and coordinate information of the corresponding first facial keypoint under a same coordinate system.
However, Zhu et al. teach wherein the difference information is determined according to coordinate information of the second facial keypoint and coordinate information of the corresponding first facial keypoint ("For each key point in the first facial sample image, the second position of the key point in the second facial sample image is determined, and the second position of the key point is subtracted from the first position of the key point to obtain offset of the key point," par. 71) under a same coordinate system ("The first position and the second position are represented by coordinates in a same coordinate system, and the offset is represented in a vector form," par. 71).
Yan et al., Metaxas et al., Kim et al., and Zhu et al. are combined as per claim 3.
Claim 23
Regarding claim 23, Yan et al., Metaxas et al., Kim et al., and Zhu et al. teach the steps of the training method of an expression driving model according to claim 7 as noted above.
Yan et al. also teach an electronic device, comprising: a processor and a memory communicatively connected with the processor, wherein the memory stores computer-executable instructions ("a memory, a processor, and a bus system, the bus system connecting the memory to the processor; the memory being configured to store a plurality of computer programs," par. 12-13); and wherein the computer-executable instructions, upon execution of the processor, cause the processor to implement the steps ("the processor being configured to execute the plurality of computer programs," par. 14).
Yan et al., Metaxas et al., Kim et al., and Zhu et al. are combined as per claim 3.
Claim 24
Regarding claim 24, Yan et al., Metaxas et al., Kim et al., and Zhu et al. teach the steps of the training method of an expression driving model according to claim 7 as noted above.
Yan et al. also teach a non-transitory computer-readable storage medium with computer-executable instructions stored thereon, wherein the computer-executable instructions, when executed by a processor, cause the processor to implement ("the embodiments of the present disclosure provide a non-transitory computer readable storage medium, storing a plurality of computer programs. The computer programs are configured for performing the aforementioned method for training an expression transfer model," par. 15).
Yan et al., Metaxas et al., Kim et al., and Zhu et al. are combined as per claim 3.
Conclusion
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Karsten F. Lantz whose telephone number is (571)272-4564. The examiner can normally be reached Monday-Friday 8:00-4:00.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Ms. Jennifer Mehmood can be reached on 571-272-2976. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/Karsten F. Lantz/Examiner, Art Unit 2664
Date: 8/21/2026
/JENNIFER MEHMOOD/ Supervisory Patent Examiner, Art Unit 2664