DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-10 is/are rejected under 35 U.S.C. 103 as being unpatentable over Korobchenko et al. (WO2024010484) in view of Verius et al. (Provisional Patent Application 638655070).
Regarding claim 1, Korobchenko teaches a virtual performer expression adjustment system with emotion-aware, comprising:
An emotional voice database, configured to store voice signals and emotion messages, wherein each of the voice signals corresponds to one of the emotion messages (“a collection of speech performances can be captured of one or more actors….full dataset for use in training” (Korobchenko: 0029));
And a computer host, connected to the emotional voice database and comprising (“computer system 900 may include, without limitation, processor 902 that may include, without limitation, one or more execution units 908 to perform machine learning model training and/or inferencing” (Korobchenko: 0112)):
A non-transitory computer-readable storage medium, configured to store computer readable instructions (Korobchenko: 0197);
And a hardware processor, electrically connected to the non-transitory computer-readable storage medium, and configured to execute the computer readable instructions to operate (Korobchenko: 0197):
Loading the voice signals and the emotion message corresponding to the loaded voice signals as training data from the emotional voice database, and inputting the training data into an artificial intelligence model to perform training to generate an emotion recognition model (Korobchenko teaches providing audio and emotion input as data to a deep neural network 206 (Korobchenko: 0030));
Receiving a user voice, performing feature extraction (“extract raw formant information using fixed-function autocorrelation analysis… the convolutional layers learn to extract short-term features that are relevant for facial animation, such as intonation, emphasis, and specific phonemes” (Korobchenko: 0050)), standardization (each vocal track can be normalized (Korobchenko: 0054)) and dimensionality reduction processes on the user voice (An autocorrelation layer can convert the input audio window to a compact 2D representation for the subsequent convolutional layers. (Korobchenko: 0054)),
However, Korobchenko does not expressly teach but Verius teaches,
and inputting the processed user voice into the emotion recognition model to obtain an emotional status (page 4 of Verius teaches audio data need for the Audio2Emotion engine);
And executing a facial expression generation calculation to generate facial landmarks based on the emotional status (the notes section of page 3 of Verius teach emotional status represented by different floating point values. The chart on the same page shows the inference of different face poses based on audio data),
and adjusting face model parameters of a virtual performer based on the facial landmarks in real time to dynamically display a facial expression of the virtual performer (“generate realistic human avatars in real-time that can visibly express emotion appropriate for corresponding audio content” (Verius: para 3)).
Therefore, before the effective filing date of the claimed invention, it would have been obvious to one of an ordinary skill in the art to modify the teachings of Korobchenko such as to apply the audio to emotion engine of Verius, because this enables efficient AI processing.
Regarding claim 2, the combined teachings teach the virtual performer expression adjustment system with emotion-aware according to claim 1, wherein the emotion recognition model comprises a convolutional neural network (CNN), and a recurrent neural network (RNN) (“A deep neural network, such as may correspond to a U-Net, frame-based convolutional neural network (CNN), or recurrent neural network (RNN)” (Korobchenko: 0020)),
and during a training process, the emotion recognition model is allowed to receive the emotion message containing text, images, videos, audio, or a combination thereof as the training data in multimodal forms (the neural network above can take as input a segment of audio data (Korobchenko: 0020)).
Regarding claim 3, the combined teachings teach the virtual performer expression adjustment system with emotion-aware according to claim 1, wherein the facial expression generation calculation is performed by at least one of a generative adversarial network (GAN), deep learning (DL) and reinforcement learning (RL) (“generative neural network such as a generative adversarial network (GAN) can be used to directly infer image data” (Korobchenko: 0034)), and is configured to generate an face image corresponding to the emotional status and extract the facial landmarks from the face image (modeling facial components according to the emotion vector (Korobchenko: 0031)).
Regarding claim 4, the combined teachings teach the virtual performer expression adjustment system with emotion-aware according to claim 1, wherein the virtual performer has a predefined facial expression model (A fixed-topology template mesh that forma neutral pose (Korobchenko: 0062)),
the facial expression model comprises the face model parameters (the different facial parts or components that can be specified (Korobchenko: 0036, 0025)),
after the facial landmarks are generated (“provide output, such as motion, vertex, and/or deformation data” (Korobchenko: 0020)), a change of the facial landmarks is smoothed through a filter (“additional smoothing to be applied, such as where a user may be able to specify one or more smoothing parameters” (Korobchenko: 0034). Also refer to the smoothing process disclosed in para 0034 of Korobchenko), a boundary check is executed to remove an unnatural expression (the culling process for discarding unnatural motion corresponds to the “boundary check” as presently claimed (Korobchenko: 0078)),
and the processed facial landmarks are mapped to the face model parameters to modify the facial expression model (“the data. The first layer maps the set of input features to the weights of a linear basis, and the set of second layers calculate the final PCA coefficients for face and tongue, rotation values for eyeballs, and the translational displacements for jaw and head” (Korobchenko: 0053). Also note, “per-vertex difference vectors from a neutral pose in a fixed-topology face mesh” (Korobchenko: 0049)).
Regarding claim 5, the combined teachings teach the virtual performer expression adjustment system with emotion-aware according to claim 1, wherein the hardware processor further operates:
Detecting whether the facial landmarks match an inappropriate expression feature (“vectors can be discarded that result in subdued or spurious, unnatural motion, indicating that the vector may be tainted with short-term effects” (Korobchenko: 0078));
And when the facial landmarks match the inappropriate expression feature, prohibit using the facial landmarks to adjust the face model parameters of the virtual performer, and initializing the face model parameters (“vectors can be discarded” (Korobchenko: 0078)).
Claim(s) 6-10 are corresponding method claim(s) of claim(s) 1-5. The limitations of claim(s) 6-10 are substantially similar to the limitations of claim(s) 1-5. Therefore, it has been analyzed and rejected substantially similar to claim(s) 6-10.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to David H Chu whose telephone number is (571)272-8079. The examiner can normally be reached M-F: 9:30 - 1:30pm, 3:30-8:30pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Daniel F Hajnik can be reached at (571) 272-7642. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/DAVID H CHU/Primary Examiner, Art Unit 2616