DETAILED ACTION
Claims 1-22 filed September 25th 2026 are pending in the current action.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Arguments
Applicant’s arguments with respect to claim(s) 1-22 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument. The Examiner apologizes for the error in the application of Kim et al. (US11,562,597). The previous rejection is withdrawn and a new rejection has been issued below.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-5, 8-10, 12, 14, 16-22 is/are rejected under 35 U.S.C. 103 as being unpatentable over Chandran et al. (US2021/0279956) in view of Li et al. (US2015/0213604)
Consider claim 1, where Chandran teaches a decoder comprising: circuitry; and memory coupled to the circuitry, wherein in operation, (See Chandran Fig. 1 where there is a decoder 156 located in memory 144 coupled to a processor 142 located on a processor 142) the circuitry: decodes expression data indicating information expressed by a person; generates a person equivalent image through a neural network according to the expression data and at least one face model, the person equivalent image corresponding to the person; and outputs the person equivalent image. (See Chandran Fig. 2 and ¶48-49, 35 where As shown, the facial identity code 206 and the facial expression code 208 are concatenated together into a concatenated code 210 that is fed to the decoder 156. The concatenated code 210 can be in the form of a vector having dimension nid+nexp. The decoder 156 is trained to reconstruct a given identity in a desired expression. In some embodiments, each of the identity encoder 152 and the expression encoder 154 may include a deep neural network, such as an encoder from a variational autoencoder (VAE). Similarly, the decoder 156 may also include a deep neural network in some embodiments.)
Chandran teaches a synthetic face model; however Chandran does not explicitly teach a profile image of the person. However, in an analogous field of endeavor Li teaches a profile image of the person. (See Li ¶59-61 where a 3D morphable face model is created from 2D face image inputs.) Therefore, it would have been obvious for one of ordinary skill in the art to recognize that the synthetic face model of Chandran would be created from images of the person as taught by Li. One of ordinary skill in the art would have been motivated to perform the modification for the advantage of/ benefit of using known methods of creating a face model to yield the desired results.
Consider claim 2, where Chandran in view of Li teaches the decoder according to claim 1, wherein the expression data includes data originated from a video of the person. (See Chandran Fig. 2 and ¶43 where he application 146 could perform blendweight retargeting in which the face model 150 is used to transfer facial expression(s) from an image or video to a new facial identity by determining blendweights associated with the facial expression(s) in the image or video, inputting the blendweights into the expression encoder 154.)
Consider claim 3, where Chandran in view of Li teaches the decoder according to claim 1, wherein the expression data includes audio data of the person. (See Li ¶17 where during the initial video recording, the facial characteristics of the user are detected and changes therein are tracked, as may result, for example, from changing of the user's facial expression, movement of the user's head, etc. Thereafter, those changes are mapped to the selected avatar on a frame-by-frame basis, and the resultant collection of avatar frames can be encoded with the original audio (if any)) Therefore, it would have been obvious for one of ordinary skill in the art to recognize that the video of Chandran may comprise sound taught by Li. One of ordinary skill in the art would have been motivated to perform the modification for the advantage of/ benefit of using existing data in a video.
Consider claim 4, where Chandran in view of Li teaches the decoder according to claim 1, wherein the at least one profile image is composed of a plurality of profile images, and the circuitry: selects one profile image from among the plurality of profile images according to the expression data; and generates the person equivalent image through the neural network according to the one profile image. (See Chandran ¶2 where such models typically generate a tensor of different dimensions that a user is permitted to control, such as the identity and expressions of faces that are being generated. User control over the identity and expressions of faces is oftentimes referred to as having “semantic control” of those facial dimensions. Also See Li ¶42 where avatar control module 210 may include custom, proprietary, known, and/or after-developed graphics processing code (or instruction sets) that are generally well-defined and operable to generate parameters for animating the avatar selected by avatar selection module 208 based on the face/head position and/or facial characteristics 206 detected by face detection module 204.)
Consider claim 5, where Chandran in view of Li teaches the decoder according to claim 4, wherein the expression data includes an index indicating a facial expression of the person, and the plurality of profile images correspond to a plurality of facial expressions of the person. (See Chandran Fig. 4B and ¶57-58 where FIG. 4B illustrates exemplar facial expressions 402 along an expression dimension generated using the face model 150 of FIG. 1, according to various embodiments. As shown, the blendweight associated with a smile expression shape has been varied between −2.5 and 2.5 to generate a set of facial expressions 402 using the face model 150.)
Consider claim 8, where Chandran in view of Li teaches the decoder according to claim 1, wherein the expression data includes data indicating at least one of a facial expression, a head pose, a facial part movement, and a head movement. (See Li ¶40 where Avatar control module 210 may include custom, proprietary, known, and/or after-developed avatar generation processing code (or instruction sets) that are generally well-defined and operable to generate an avatar based on the user's face/head position and/or facial characteristics 206 detected by face detection module 208)
Consider claim 9, where Chandran in view of Li teaches the decoder according to claim 1, wherein the expression data includes data represented by coordinates. (See Chandran ¶36 where the representation of the facial identity that is input into the identity encoder is the difference between a three-dimensional (3D) mesh (e.g., a triangle mesh) of a particular face with a neutral expression and a reference mesh. As used herein, a “neutral” expression refers to a facial expression with neutral positioning of facial features, which is in contrast to other expressions that show stronger emotions such as smiling, crying, etc. The reference mesh can be an average of multiple meshes of faces with neutral expressions, and the difference between the mesh of a particular face and the reference mesh can include displacements between vertices of the two meshes.)
Consider claim 10, where Chandran in view of Li teaches the decoder according to claim 1, wherein the circuitry decodes the at least one profile image. (See Li ¶59-61 where a 3D morphable face model is created from 2D face image inputs.) Therefore, it would have been obvious for one of ordinary skill in the art to recognize that the synthetic face model of Chandran would be created from images of the person as taught by Li. One of ordinary skill in the art would have been motivated to perform the modification for the advantage of/ benefit of using known methods of creating a synthetic face model to yield the desired results.
Consider claim 12, where Chandran in view of Li teaches the decoder according to claim 1, wherein the circuitry reads the at least one profile image from the memory. (See Chandran ¶42 where the trained face model 150 and/or the landmark model may be deployed to any suitable applications that generate faces and use the same. Illustratively, a face generating application 146 is stored in a memory 144, and executes on a processor 142 of the computing device 140. Components of the computing device 140, including the memory 144 and the processor 142 may be similar to corresponding components of the machine learning server 110..)
Consider claim 14, where Chandran in view of Li teaches the decoder according to claim 3, wherein the at least one profile image is composed of one profile image, and the circuitry: derives, by simulating a head movement or an eye movement, a second feature set indicating the head movement or the eye movement; and generates the person equivalent image through the neural network according to the audio data, the second feature set, and the one profile image. (See Li ¶76-77, 48 where The process can continue as in block 44b with decomposing movement of the detected/tracked facial feature points into at least two categories: (1) facial expression movements; and (2) head rigid movements. The former category (facial expression movements) may include non-rigid transformations, for instance, due to facial expressions. The latter category (head rigid movements) may include rigid movements (e.g., translation, rotation, and scaling factors) due to head gestures. This also can be performed, for example, using face detection module 204, as previously discussed.
See Li ¶83 where any detected face/head movements, including movement of and/or changes in one or more of the user's facial characteristics 206 (e.g., eyes, nose, mouth, etc.) can be converted into parameters usable for animating an avatar mesh (e.g., such as is discussed above with reference to the example avatar mesh of FIG. 3C).) Therefore, it would have been obvious for one of ordinary skill in the art that the facial expression data of Chandran would comprise movement of the eyes as taught by Li. One of ordinary skill in the art would have been motivated to perform the modification for the advantage of/ benefit of extracting known parameters used for animating facial features.
Consider claim 16, where Chandran teaches an encoder comprising: circuitry; and memory coupled to the circuitry, wherein in operation, (See Chandran Fig. 1 where there are encoders 152, 154, located in memory 144 coupled to a processor 142 located on a processor 142) the circuitry: encodes expression data indicating information expressed by a person; generates a person equivalent image through a neural network according to the expression data and at least one synthetic face model, the person equivalent image corresponding to the person; and outputs the person equivalent image. (See Chandran ¶43, 35 where The application 146 can retarget the expressions of the individual in the video frames to meshes of faces having various identities by processing the detected sets of landmarks using the mapping module 806, which performs a mapping between 2D facial landmarks and expression codes, and further inputting a representation of a target facial identity into the identity encoder 152 that generates an associated identity code. Thereafter, the application 146 can concatenate the identity code together with the expression codes, and feed the concatenated codes into the decoder 156 to generate representations of faces having the target facial identity and the expressions depicted in the video.)
Chandran teaches a synthetic face model; however Chandran does not explicitly teach a profile image of the person. However, in an analogous field of endeavor Li teaches a profile image of the person. (See Li ¶59-61 where a 3D morphable face model is created from 2D face image inputs.) Therefore, it would have been obvious for one of ordinary skill in the art to recognize that the synthetic face model of Chandran would be created from images of the person as taught by Li. One of ordinary skill in the art would have been motivated to perform the modification for the advantage of/ benefit of using known methods of creating a synthetic face model to yield the desired results.
Consider claim 17, where Chandran in view of Li teaches the encoder according to claim 16, wherein the expression data includes data originated from a video of the person. (See Chandran Fig. 2 and ¶43 where he application 146 could perform blendweight retargeting in which the face model 150 is used to transfer facial expression(s) from an image or video to a new facial identity by determining blendweights associated with the facial expression(s) in the image or video, inputting the blendweights into the expression encoder 154.)
Consider claim 18, where Chandran in view of Li teaches the encoder according to claim 16, wherein the expression data includes audio data of the person. (See Li ¶17 where during the initial video recording, the facial characteristics of the user are detected and changes therein are tracked, as may result, for example, from changing of the user's facial expression, movement of the user's head, etc. Thereafter, those changes are mapped to the selected avatar on a frame-by-frame basis, and the resultant collection of avatar frames can be encoded with the original audio (if any)) Therefore, it would have been obvious for one of ordinary skill in the art to recognize that the video of Chandran may comprise sound taught by Li. One of ordinary skill in the art would have been motivated to perform the modification for the advantage of/ benefit of using existing data in a video.
Consider claim 19, where Chandran in view of Li teaches the encoder according to claim 16, wherein the at least one profile image is composed of a plurality of profile images, and the circuitry: selects one profile image from among the plurality of profile images according to the expression data; and generates the person equivalent image through the neural network according to the one profile image. (See Chandran ¶2 where such models typically generate a tensor of different dimensions that a user is permitted to control, such as the identity and expressions of faces that are being generated. User control over the identity and expressions of faces is oftentimes referred to as having “semantic control” of those facial dimensions. Also See Li ¶42 where avatar control module 210 may include custom, proprietary, known, and/or after-developed graphics processing code (or instruction sets) that are generally well-defined and operable to generate parameters for animating the avatar selected by avatar selection module 208 based on the face/head position and/or facial characteristics 206 detected by face detection module 204.)
Consider claim 20, where Chandran in view of Li teaches the encoder according to claim 19, wherein the expression data includes an index indicating a facial expression of the person, and the plurality of profile images correspond to a plurality of facial expressions of the person. (See Chandran Fig. 4B and ¶57-58 where FIG. 4B illustrates exemplar facial expressions 402 along an expression dimension generated using the face model 150 of FIG. 1, according to various embodiments. As shown, the blendweight associated with a smile expression shape has been varied between −2.5 and 2.5 to generate a set of facial expressions 402 using the face model 150.)
Consider claim 21, where Chandran teaches a decoding method comprising: decoding expression data indicating information expressed by a person; generating a person equivalent image through a neural network according to the expression data and at least one synthetic face model of the person, the person equivalent image corresponding to the person; and outputting the person equivalent image. (See Chandran Fig. 2 and ¶48-49, 35 where As shown, the facial identity code 206 and the facial expression code 208 are concatenated together into a concatenated code 210 that is fed to the decoder 156. The concatenated code 210 can be in the form of a vector having dimension nid+nexp. The decoder 156 is trained to reconstruct a given identity in a desired expression. In some embodiments, each of the identity encoder 152 and the expression encoder 154 may include a deep neural network, such as an encoder from a variational autoencoder (VAE). Similarly, the decoder 156 may also include a deep neural network in some embodiments.)
Chandran teaches a synthetic face model; however Chandran does not explicitly teach a profile image of the person. However, in an analogous field of endeavor Li teaches a profile image of the person. (See Li ¶59-61 where a 3D morphable face model is created from 2D face image inputs.) Therefore, it would have been obvious for one of ordinary skill in the art to recognize that the synthetic face model of Chandran would be created from images of the person as taught by Li. One of ordinary skill in the art would have been motivated to perform the modification for the advantage of/ benefit of using known methods of creating a synthetic face model to yield the desired results.
Consider claim 22, where Chandran teaches an encoding method comprising: encoding expression data indicating information expressed by a person; generating a person equivalent image through a neural network according to the expression data and at least one synthetic face model of the person, the person equivalent image corresponding to the person; and outputting the person equivalent image. (See Chandran ¶43, 35 where The application 146 can retarget the expressions of the individual in the video frames to meshes of faces having various identities by processing the detected sets of landmarks using the mapping module 806, which performs a mapping between 2D facial landmarks and expression codes, and further inputting a representation of a target facial identity into the identity encoder 152 that generates an associated identity code. Thereafter, the application 146 can concatenate the identity code together with the expression codes, and feed the concatenated codes into the decoder 156 to generate representations of faces having the target facial identity and the expressions depicted in the video.)
Chandran teaches a synthetic face model; however Chandran does not explicitly teach a profile image of the person. However, in an analogous field of endeavor Li teaches a profile image of the person. (See Li ¶59-61 where a 3D morphable face model is created from 2D face image inputs.) Therefore, it would have been obvious for one of ordinary skill in the art to recognize that the synthetic face model of Chandran would be created from images of the person as taught by Li. One of ordinary skill in the art would have been motivated to perform the modification for the advantage of/ benefit of using known methods of creating a synthetic face model to yield the desired results.
Claim(s) 6, 7, and 11 is/are rejected under 35 U.S.C. 103 as being unpatentable over Chandran in view of Li as applied to claim 1 above, in further view of Teo et al. (US2018/0192046)
Consider claim 6, where Chandran in view of Li teaches the decoder according to claim 1, wherein the circuitry decodes the expression data from each of data regions in a source video. (See Chandran ¶39 where Such meshes may be obtained from standalone images and/or the image frames of a video.)
Chandran teaches a source video; however Chandran does not explicitly teach a bitstream. However, in an analogous field of endeavor Teo teaches a bitstream. (See Teo ¶369 where control parameters in the encoded video bitstream) Therefore, it would have been obvious for one of ordinary skill in the art that the video of Chandran would be encoded in a bitstream as taught by Teo. One of ordinary skill in the art would have been motivated to perform the modification for the advantage of/ benefit of using known representations of video data. One of ordinary skill in the art would have been motivated to perform the modification for the advantage of/ benefit of using existing standards to yield predictable results.
Consider claim 7, where Chandran in view of Li teaches the decoder according to claim 1, however they do not explicitly teach wherein the circuitry decodes the expression data from a header of a bitstream. However, in an analogous field of endeavor Teo teaches a bitstream. (See Teo ¶369 where control parameters in the encoded video bitstream. a control parameter is present in the header of a slice.) Therefore, it would have been obvious for one of ordinary skill in the art that the video of Chandran would be encoded in a bitstream as taught by Teo. One of ordinary skill in the art would have been motivated to perform the modification for the advantage of/ benefit of using known representations of video data. One of ordinary skill in the art would have been motivated to perform the modification for the advantage of/ benefit of using existing standards to yield predictable results
Consider claim 11, where Chandran in view of Li teaches the decoder according to claim 1, wherein the circuitry: decodes the expression data from a first video; and decodes the at least one profile image from a second video different from the first video. (See Chandran ¶43, 35 where The application 146 can retarget the expressions of the individual in the video frames to meshes of faces having various identities by processing the detected sets of landmarks using the mapping module 806, which performs a mapping between 2D facial landmarks and expression codes, and further inputting a representation of a target facial identity into the identity encoder 152 that generates an associated identity code. Thereafter, the application 146 can concatenate the identity code together with the expression codes, and feed the concatenated codes into the decoder 156 to generate representations of faces having the target facial identity and the expressions depicted in the video.)
Chandran teaches a video; however Chandran does not explicitly teach a bitstream. However, in an analogous field of endeavor Teo teaches a bitstream. (See Teo ¶369 where control parameters in the encoded video bitstream) Therefore, it would have been obvious for one of ordinary skill in the art that the video of Chandran would be encoded in a bitstream as taught by Teo. One of ordinary skill in the art would have been motivated to perform the modification for the advantage of/ benefit of using known representations of video data. One of ordinary skill in the art would have been motivated to perform the modification for the advantage of/ benefit of using existing standards to yield predictable results.
Claim(s) 13, 15 is/are rejected under 35 U.S.C. 103 as being unpatentable over Chandran in view of Li as applied to claim 1 above, in further view of Tian et al. ("Audio2Face: Generating Speech/Face Animation from Single Audio with Attention-Based Bidirectional LSTM Networks," 2019 IEEE International Conference on Multimedia & Expo Workshops (ICMEW), Shanghai, China, 2019, pp. 366-371, doi: 10.1109/ICMEW.2019.00069.)
Consider claim 13, where Chandran in view of Li teaches the decoder according to claim 3, wherein the at least one profile image is composed of one profile image, however they do not explicitly teach the circuitry: derives, from the audio data, a first feature set indicating a mouth movement; and generates the person equivalent image through the neural network according to the first feature set and the one profile image. (See Tian section 3: Approach, where “Then we build several 3D blendshape-based cartoon face models with counterpart parameters to control different parts of the face: eyebrows, eyes, lip, jaw, etc. Further, we bring attention mechanism into bidirectional LSTM networks to map input audio feature to animation parameters. Input a short span of audio, the goal of our framework is to correspondingly return a series of successive lip/facial movements.”) Therefore, it would have been obvious for one of ordinary skill in the art that the blendweights of Chandran can be modified to receive the blendshape derived form the audio of Tian. One of ordinary skill in the art would have been motivated to perform the modification for the advantage of/ benefit of using other known methods of creating plausible lip movements. (See Tian’s introduction where “In this work, our primary goal is to recreate a plausible 3D virtual talking avatar that can make reasonable lip movements and then further generate natural facial movements.”)
Consider claim 15, where Chandran in view of Li teaches the decoder according to claim 3, however they do not explicitly teach wherein the circuitry matches a facial expression in the person equivalent image to a facial expression inferred from the audio data. However, in an analogous field of endeavor Tian teaches wherein the circuitry matches a facial expression in the person equivalent image to a facial expression inferred from the audio data. (See Tian section 3: Approach, where “Then we build several 3D blendshape-based cartoon face models with counterpart parameters to control different parts of the face: eyebrows, eyes, lip, jaw, etc. Further, we bring attention mechanism into bidirectional LSTM networks to map input audio feature to animation parameters. Input a short span of audio, the goal of our framework is to correspondingly return a series of successive lip/facial movements.”) Therefore, it would have been obvious for one of ordinary skill in the art that the blendweights of Chandran can be modified to receive the blendshape derived form the audio of Tian. One of ordinary skill in the art would have been motivated to perform the modification for the advantage of/ benefit of using other known methods of creating plausible lip movements. (See Tian’s introduction where “In this work, our primary goal is to recreate a plausible 3D virtual talking avatar that can make reasonable lip movements and then further generate natural facial movements.”)
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to WILLIAM LU whose telephone number is (571)270-1809. The examiner can normally be reached 10am-6:30pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Matthew Eason can be reached at 571-270-7230. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
WILLIAM LU
Primary Examiner
Art Unit 2624
/WILLIAM LU/Primary Examiner, Art Unit 2624