DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Continued Examination Under 37 CFR 1.114
A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on April 30, 2026 has been entered.
Response to Arguments
Applicant’s arguments, see p. 18-19, filed April 30, 2026, with respect to Claims 6 and 19 have been fully considered and are persuasive. The 35 U.S.C. 103 rejections of Claims 6 and 19 have been withdrawn.
Applicant's arguments filed April 30, 2026, with respect to Claims 1-5, 7, 8, 18, and 20 have been fully considered but they are not persuasive.
As per Claim 1, Applicant argues that Li (US 20170243387A1) does not suggest that the eye convolution neural network 311 is trained to determine the eye animation based on “audio data” while refraining determining the mouth animation based on the “audio data”. Li does not suggest that the mouth convolution neural network 312 is trained to determine the mouth animation based on “audio data” while refraining from determining the eye animation based on the “audio data.” Li does not suggest that the eye convolution neural network is trained to “refrain” from animating the mouth and/or the mouth convolution neural network is trained to “refrain” from animating the eyes (p. 14).
In reply, the Examiner points out that Li describes “For the eyes convolutional neural network 311 speech animation control weights 313 derived during the training process are applied” [0049]. Li describes “capture audio recitation of phonemes during training” [0026]. Li describes “deriving extremely accurate facial and speech animation by replying upon a corresponding training dataset of a series of users speaking known phonemes” [0019]. Thus, the phonemes are based on audio data, and the speech animation control weights 313 are based on the phonemes, and thus the eye convolution neural network 311 is trained to determine the eye animation based on speech animation control weights 313 (based on phonemes, which are based on the audio data) [0049, 0026, 0019]. Li describes “audio and image data…reciting…sentences which may be used to derive visemes suitable for use in composing animations form the audio and image data” [0059]. Li describes “’viseme’ as used herein means the visual facial expressions of one pronouncing a given phoneme” [0044]. Li describes “For the mouth convolutional neural network 312, mouth animation control weights 314 are applied. Both of these rely upon the trained datasets 315 created by the training system 200” [0049]. Thus, visemes are based on the audio data, and thus the mouth convolutional neural network 312 is trained to determine the mouth animation based on visemes (based on the audio data) [0059, 0044, 0049]. The Examiner was not able to find any description in Applicant’s disclosure that the eye convolution neural network is trained to “refrain” from animating the mouth and/or the mouth convolution neural network is trained to “refrain” from animating the eyes. The closest description that the Examiner found in Applicant’s disclosure appears to be paragraphs [0025-0026] on p. 7-8, which does describe that the first neural network is trained based on displacing an upper portion of a facial representation of the animated character based on the audio data; and the second neural network is trained based on displacing the lower portion of the facial representation based on the audio data. This is what Li teaches, and thus Li teaches the claimed limitations in the manner that is supported by Applicant’s disclosure.
Applicant’s arguments with respect to claim(s) 9-14, 16, 17, and 21 have been considered but are moot because new grounds of rejection are made in view of Villanueva Aylagas (US012406419B1).
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claim(s) 1, 7, 8, 18, and 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Beith (US 20240078732A1) and Li (US 20170243387A1).
As per Claim 1, Beith teaches a method comprising: obtaining audio data (204) representative of one or more words (microphones are configured to generate audio data 204, microphone configured to capture speech of the user, [0074]) to be output by an animated character; generating, using one or more first neural networks (222) and based at least on the audio data, a first output indicating an emotional state for the animated character, wherein the one or more first neural networks are trained based on displacing first points of an upper portion of a facial representation of the animated character based on the audio data; generating, using one or more neural networks and based at least on the audio data and the first output, a second output associated with a facial animation for the animated character, wherein the one or more neural networks are trained based on displacing the second points of the lower portion of the facial representation based on the audio data; and causing, based on the second output, the animated character to perform the facial animation while outputting the one or more words (audio unit 222 includes a deep learning neural network, [0080], semantical context is based on an emotion 270 associated with the speech 258 represented in the audio data 204, processors 116 are configured to process the audio data 204 to predict the emotion 270, in addition to detecting emotion associated with the meanings of words of the user’s speech 258, the audio unit 222 can include machine learning models that are configured to detect audible emotions based on the speaking characteristics of the user, feature data generator 120 may be configured to associate particular facial expressions with various audible emotions, adjusted face data 134 causes the avatar facial expression 156 to represent the emotion 270 (e.g., smiling to express happiness, eyes narrowed to express anger, eyes widened to express surprise, etc.), [0083], feature data generator 120 includes an image unit 226 that is configured to generate a facial representation 228 based on the image data 208 and that may indicate the semantical context 122, by processing the image data 208 using neural networks of the image unit 226, the resulting adjusted face data 134 can provide a more accurate and realistic facial expression of the avatar 154, facial representation 228 includes an indication of expressions, movements of the user, [0086], context-based future speech prediction network 1210 processes the audio data 204 to determine a predicted word in context 1220, context-based future speech prediction network 120 includes neural network, [0135], representation generator 1230 is configured to generate a representation 1250 of the predicted word in context 1220, the representation 1250 may be concatenated to the image-based features 322 to generate the feature data 124, [0136], context-based future speech prediction network 1210 and the representation generator 1230 enable prediction, based on a context of spoken words, of what a word will be, which is used to predict an avatar’s facial image or to ensure compliance frame-to-frame, to ensure that the image of the avatar pronouncing words is transitioning correctly over time, [0137]).
However, Beith does not teach that the one or more first neural networks are trained based on displacing the first points of the upper portion of the facial representation of the animated character based on the audio data while refraining from displacing second points of a lower portion of the facial representation based on the audio data; generating, using one or more second neural networks, the second output associated with the facial animation for the animated character, wherein the one or more second neural networks are trained based on displacing the second points of the lower portion of the facial representation based on the audio data while refraining from displacing the first points of the upper portion of the facial animation based on the audio data. However, Li describes “For the eyes convolutional neural network 311 speech animation control weights 313 derived during the training process are applied” [0049]. Li describes “capture audio recitation of phonemes during training” [0026]. Li describes “deriving extremely accurate facial and speech animation by replying upon a corresponding training dataset of a series of users speaking known phonemes” [0019]. Thus, the phonemes are based on audio data, and the speech animation control weights 313 are based on the phonemes, and thus the eye convolution neural network 311 is trained to determine the eye animation based on speech animation control weights 313 (based on phonemes, which are based on the audio data) [0049, 0026, 0019]. Li describes “audio and image data…reciting…sentences which may be used to derive visemes suitable for use in composing animations form the audio and image data” [0059]. Li describes “’viseme’ as used herein means the visual facial expressions of one pronouncing a given phoneme” [0044]. Li describes “For the mouth convolutional neural network 312, mouth animation control weights 314 are applied. Both of these rely upon the trained datasets 315 created by the training system 200” [0049]. Thus, visemes are based on the audio data, and thus the mouth convolutional neural network 312 is trained to determine the mouth animation based on visemes (based on the audio data) [0059, 0044, 0049]. Thus, Li teaches wherein the one or more first neural networks (eyes convolutional neural network 311) are trained based on displacing the first points of the upper portion of the facial representation of the animated character based on the audio data while refraining from displacing second points of a lower portion of the facial representation based on the audio data [0049, 0026, 0019]; generating, using one or more second neural networks (312), the second output associated with the facial animation for the animated character, wherein the one or more second neural networks are trained based on displacing the second points of the lower portion of the facial representation based on the audio data while refraining from displacing the first points of the upper portion of the facial animation based on the audio data [0059, 0044, 0049].
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Beith so one or more first neural networks are trained based on displacing first points of upper portion of facial representation of animated character based on audio data while refraining from displacing second points of lower portion of the facial representation based on the audio data; generating, using one or more second neural networks, the second output associated with the facial animation for the animated character, wherein the one or more second neural networks are trained based on displacing the second points of the lower portion of the facial representation based on the audio data while refraining from displacing the first points of the upper portion of the facial animation based on the audio data because Li suggests that properly-trained neural networks can be applied to the mouth image data and the eyes image data to derive extremely accurate facial and speech animation [0019].
As per Claim 7, Beith does not teach wherein: the upper portion of the facial representation includes at least one of one or more eyes, a nose, or one or more eyebrows of the facial representation; and the lower portion of the facial representation includes at least one of a mouth, one or more cheeks, or a chin of the facial representation. However, Li teaches wherein: the upper portion of the facial representation includes at least one of one or more eyes, a nose, or one or more eyebrows of the facial representation [0047]; and the lower portion of the facial representation includes at least one of a mouth, one or more cheeks, or a chin of the facial representation [0049]. This would be obvious for the reasons given in the rejection for Claim 1.
As per Claim 8, Beith does not teach wherein: the one or more first neural networks are trained based on displacing the first points of the upper portion of the facial representation of the animated character based on processing input data representing the one or more words represented by the audio data while refraining from displacing the second points of the lower portion of the facial representation based on the input data; and the one or more second neural networks are trained based on displacing the second points of the lower portion of the facial representation of the animated character based on processing the input data while refraining from displacing the first points of the upper portion of the facial representation based on the input data. However, Li describes “For the eyes convolutional neural network 311 speech animation control weights 313 derived during the training process are applied” [0049]. Li describes “capture audio recitation of phonemes during training” [0026]. Li describes “deriving extremely accurate facial and speech animation by replying upon a corresponding training dataset of a series of users speaking known phonemes” [0019]. Thus, the phonemes are based on audio data, and the speech animation control weights 313 are based on the phonemes, and thus the eye convolution neural network 311 is trained to determine the eye animation based on speech animation control weights 313 (based on phonemes, which are based on the audio data) [0049, 0026, 0019]. Li describes “audio and image data…reciting…sentences which may be used to derive visemes suitable for use in composing animations form the audio and image data” [0059]. Li describes “’viseme’ as used herein means the visual facial expressions of one pronouncing a given phoneme” [0044]. Li describes “For the mouth convolutional neural network 312, mouth animation control weights 314 are applied. Both of these rely upon the trained datasets 315 created by the training system 200” [0049]. Thus, visemes are based on the audio data, and thus the mouth convolutional neural network 312 is trained to determine the mouth animation based on visemes (based on the audio data) [0059, 0044, 0049]. Thus, Li teaches wherein: the one or more first neural networks (311) are trained based on displacing the first points of the upper portion of the facial representation of the animated character based on processing input data representing the one or more words represented by the audio data while refraining from displacing the second points of the lower portion of the facial representation based on the input data [0049, 0026, 0019]; and the one or more second neural networks (312) are trained based on displacing the second points of the lower portion of the facial representation of the animated character based on processing the input data while refraining from displacing the first points of the upper portion of the facial representation based on the input data [0059, 0044, 0049]. This would be obvious for the reasons given in the rejection for Claim 1.
As per Claim 18, Claim 18 is similar in scope to Claim 1, and therefore is rejected under the same rationale.
As per Claim 20, Beith teaches wherein the processor is comprised in at least one of: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing simulation operations; a system for performing digital twin operations; a system for performing light transport simulation; a system for performing collaborative content creation for 3D assets; a system for performing deep learning operations; a system implemented using an edge device; a system implemented using a robot; a system implemented using one or more large language models; a system for performing conversational AI operations; a system for generating synthetic data; a system incorporating one or more virtual machines (VMs); a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources (deep learning architecture neural network, [0078]).
Claim(s) 2-4 is/are rejected under 35 U.S.C. 103 as being unpatentable over Beith (US 20240078732A1) and Li (US 20170243387A1) in view of Karras (US 20180336464A1).
As per Claim 2, Beith and Li are relied upon for the teachings as discussed above relative to Claim 1.
However, Beith and Li do not teach the second output corresponds to locations of a plurality of vertices associated with one or more of the first points or the second points, an individual vertex of the plurality of vertices representing a three-dimensional point associated with the facial representation of the animated character. However, Karras teaches wherein the second output corresponds to locations of a plurality of vertices associated with one or more of the first points or the second points, an individual vertex of the plurality of vertices representing a three-dimensional point associated with the facial representation of the animated character (creating facial animation sequences in a vertex mesh based on audio input, [0079], produce the final 3D positions of a plurality of control vertices of a mesh 808, the facial mesh includes 5022 control vertices that can be moved in 3D positions, [0090]).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Beith and Li so that the second output corresponds to locations of a plurality of vertices associated with one or more of the first points or the second points, an individual vertex of the plurality of vertices representing a three-dimensional point associated with the facial representation of the animated character because Karras suggests that this way, a 3D realistic looking face can be animated [0079, 0090].
As per Claim 3, Beith and Li do not teach further comprising: determining, using one or more third neural networks and based at least on the audio data, a third output, wherein: the determining the first output indicating the emotional state is based at least on the third output; and the determining the second output associated with the facial animation for the animated character is based at least on the first output and the third output. However, Karras teaches further comprising: determining, using one or more third neural networks (810) and based at least on the audio data (802), a third output, wherein: the determining the first output indicating the emotional state is based at least on the third output; and the determining the second output associated with the facial animation for the animated character is based at least on the first output and the third output (first part of the deep neural network is a formant analysis network 810, the formant analysis network 810 receives an audio input 802 and produces a time-varying sequence of speech features that are passed to the articulation network 820, the articulation network 820 analyze the temporal evolution of the features and output a single abstract feature vector that describes the facial pose, the articulation network 820 outputs a set of 256+E abstract features, where E is the number of components of the emotional state vector 804, the set of abstract features output by the articulation network 820 is fed to an output network 830, which generates the final 3D positions of a plurality of vertices in a mesh output 808, [0083], [0090, 0079]).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Beith and Li to include determining, using one or more third neural networks and based at least on the audio data, a third output, wherein: the determining the first output indicating the emotional state is based at least on the third output; and the determining the second output associated with the facial animation for the animated character is based at least on the first output and the third output because Karras suggests that this is useful for generating plausible and expressive 3D facial animation based exclusively on a vocal audio track [0081].
As per Claim 4, Beith and Li do not teach wherein the determining the second output associated with the facial animation comprises: determining, using at least one third neural network of the one or more second neural networks, and based at least on the audio data and the first output, a third output; and determining, using at least one fourth neural network of the one or more second neural networks, and based at least on the third output, the second output associated with the facial animation for the animated character. However, Karras teaches wherein the determining the second output associated with the facial animation comprises: determining, using at least one third neural network (820) of the one or more second neural networks, and based at least on the audio data and the first output, a third output; and determining, using at least one fourth neural network (830) of the one or more second neural networks, and based at least on the third output, the second output associated with the facial animation for the animated character [0083, 0090, 0079]. This would be obvious for the reasons given in the rejection for Claim 3.
Claim(s) 5 is/are rejected under 35 U.S.C. 103 as being unpatentable over Beith (US 20240078732A1) and Li (US 20170243387A1) in view of Bolzoni (US 20240021196A1).
Beith and Li are relied upon for the teachings as discussed above relative to Claim 1. Beith teaches generating, using one or more first neural networks and based at least on the audio data, a first output indicating an emotional state; determining, using one or more second neural networks and based at least on the audio data and the first output, a second output associated with a facial animation; and causing, based at least on the second output, the animated character to perform the facial animation, as discussed in the rejection for Claim 1.
However, Beith and Li do not teach generating using the one or more first neural networks and based at least on second audio data corresponding to one or more second words, a third output associated with at least one of the emotional state or a second emotional state; determining, using the one or more second neural networks and based at least on the second audio data and the third output, a fourth output associated with a second facial animation; and causing, based at least on the fourth output, the animated character to perform the second facial animation. However, Bolzoni teaches that the user’s first response (first audio data) is “I’m very worried. It’s Peter, we got into an argument again” [0046]. The user’s second response (second audio data) is “I believe he doesn’t love me anymore so I dumped him” [0048]. Thus, Bolzoni teaches further comprising: generating, using the one or more first neural networks (150) and based at least on second audio data correspond to one or more second words, a third output associated with at least one of the emotional state or a second emotional state (the user may reply by stating “I believe he doesn’t love me anymore so I dumped him”, [0048], once again, the input analyzer 150 may convert the user’s audio input into text and may extract metadata therefrom, the input analyzer 150 may enrich the text with the metadata and transmit the enriched text to the state machine 155, the state machine 155 may classify the user’s response of “I believe he doesn’t love me anymore so I dumped him” during the second iteration as feeling (sadness) and anxiety, [0049], [0025]); determining, using the one or more second neural networks (155) and based at least on the second audio data and the third output, a fourth output associated with a face; and causing, based at least on the fourth output, a character with the face ([0036], based on the classification of the user’s response as feeling (anger), the state machine 155 may determine that the guidance state would be an appropriate state to transition to, the state machine 155 may transition to the guidance state, [0050], the third iteration may begin with the state machine 155 outputting the base prompt to the output composer 160 along with the extracted metadata, based on the metadata extracted from the user’s response during the second iteration, the output composer 160 may modify the base prompt of the inquiry, the output composer 160 may then output the modified prompt to the speaker and display for presentation to the user via the avatar 350, [0051], Fig. 3B). Since Beith teaches generating, using one or more first neural networks and based at least on the audio data, a first output indicating an emotional state; determining, using one or more second neural networks and based at least on the audio data and the first output, a second output associated with a facial animation; and causing, based at least on the second output, the animated character to perform the facial animation, as discussed in the rejection for Claim 1, this teaching from Bolzoni of performing the processing for second audio data can be implemented into the device of Beith to include a fourth output associated with a second facial animation; and causing, based at least on the fourth output, the animated character to perform the second facial animation.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Beith and Li to include generating using the one or more first neural networks and based at least on second audio data corresponding to one or more second words, a third output associated with at least one of the emotional state or a second emotional state; determining, using the one or more second neural networks and based at least on the second audio data and the third output, a fourth output associated with a second facial animation; and causing, based at least on the fourth output, the animated character to perform the second facial animation because Bolzoni suggests that this way, the conversation can continue based on the next sentence, and thus this provides an interactive conversation platform that can engage in conversation with a user in a manner that simulates humanistic interaction including learned understanding of users [0020].
Claim(s) 9, 14, and 17 is/are rejected under 35 U.S.C. 103 as being unpatentable over Beith (US 20240078732A1) and Villanueva Aylagas (US012406419B1).
As per Claim 9, Beith teaches a system comprising: one or more processors to: determine, using one or more neural networks (222) and based at least on audio data corresponding to one or more sounds to be output by a character, a first output indicating a first emotional state for the character [0080, 0083]; determine, based at least on one or more user inputs associated with updating the first emotional state, a second emotional state (process sensor data 106 to generate feature data 124, the sensor data 106 includes image data, feature data generator 120 is configured to process the sensor data 106 to determine a semantical context 122 associated with the sensor data 106, a semantical context refers to emotions that can be determined based on the sensor data 106, the sensor data 106 includes image data, and the semantical context 122 is based on an emotion associated with an expression on a user’s face represented in the image, [0062], feature data generator 120 includes an image unit 226 that is configured to generate a facial representation 228 based on the image data 208 and that may indicate the semantical context 122, by processing the image data 208 using neural networks of the image unit 226, the resulting adjusted face data 134 can provide a more accurate and realistic facial expression of the avatar 154, facial representation 228 includes an indication of expressions, movements of the user, [0086]); determine, using the one or more neural networks and based at least on the audio data and the second emotional state, a second output associated with a facial animation for the character, wherein a greater displacement in an upper portion of a facial representation represented by the facial animation is caused by the second emotional state as compared to the audio data; and cause, based at least on the second output, the facial animation of the character while outputting the one or more sounds [0080, 0083, 0086, 0135, 0136, 0137].
However, Beith does not teach that the first output represents a first vector in a latent space that corresponds to the first emotional state for the character; determining a second vector in the latent space that corresponds to the second emotional state. However, Villanueva Aylagas teaches that the first output represents a first vector in a latent space that corresponds to the first emotional state for the character; determining a second vector in the latent space that corresponds to the second emotional state (latent vector 502-1 corresponding to a sad speech emotion, latent vector 502-3 corresponding to a happy speech emotion, respective effects in the resulting facial animation when applying each latent vector 502 as a conditioning input to the facial animation generative model, col. 14, lines 40-47).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Beith so that the first output represents a first vector in a latent space that corresponds to the first emotional state for the character; determining a second vector in the latent space that corresponds to the second emotional state as suggested by Villanueva Aylagas. It is well-known in the art that latent vectors offer efficient compression, noise reduction, and semantic representation.
As per Claim 14, Beith does not teach wherein the one or more processors are further to determine, based at least on the first output, the first vector in the latent space that corresponds to the first emotional state. However, Villanueva Aylagas teaches wherein the one or more processors are further to determine, based at least on the first output, the first vector in the latent space that corresponds to the first emotional state (conditioning inputs in the form of latent vectors corresponding to particular speech emotions can be determined, if inputting a particular latent vector with any speech audio clip to trained decoder generally causes the resulting facial animations to portray a happy expression, the particular latent vector may be stored and associated with an indication of happy as a speech emotion, in this way, conditioning inputs comprising latent vectors corresponding to various speech emotions (e.g. happy, angry, sad, neutral) can be identified, col. 11, lines 11-30). This would be obvious for the reasons given in the rejection for Claim 9.
As per Claim 17, Beith teaches wherein the system is comprised in at least one of: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing simulation operations; a system for performing digital twin operations; a system for performing light transport simulation; a system for performing collaborative content creation for 3D assets; a system for performing deep learning operations; a system implemented using an edge device; a system implemented using a robot; a system implemented using one or more large language models; a system for performing conversational AI operations; a system for generating synthetic data; a system incorporating one or more virtual machines (VMs); a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources (deep learning architecture neural network, [0078]).
Claim(s) 10-11 is/are rejected under 35 U.S.C. 103 as being unpatentable over Beith (US 20240078732A1) and Villanueva Aylagas (US012406419B1) in view of Karras (US 20180336464A1).
As per Claim 10, Claim 10 is similar in scope to Claim 2, and therefore is rejected under the same rationale.
As per Claim 11, Beith and Villanueva Aylagas do not teach the one or more processors are to determine, using the one or more neural networks and based at least on the audio data, a third output, wherein the first output is determined based at least on the third output, and wherein the second output associated with the facial animation for the character is determined based at least on the third output and the second output. However, Karras teaches the one or more processing units are to determine, using the one or more neural networks (820) and based at least on the audio data, a third output, wherein the first output is determined based at least on the third output, and wherein the second output associated with the facial animation for the character is determined based at least on the third output and the second output [0083, 0090, 0079]. This would be obvious for the reasons given in the rejection for Claim 3.
Claim(s) 12 is/are rejected under 35 U.S.C. 103 as being unpatentable over Beith (US 20240078732A1) and Villanueva Aylagas (US012406419B1) in view of Bolzoni (US 20240021196A1).
Claim 12 is similar in scope to Claim 5, and therefore is rejected under the same rationale.
Claim(s) 13 is/are rejected under 35 U.S.C. 103 as being unpatentable over Beith (US 20240078732A1) and Villanueva Aylagas (US012406419B1) in view of Li (US 20170243387A1).
Claim 13 is similar in scope to Claim 1, and therefore is rejected under the same rationale.
Claim(s) 16 is/are rejected under 35 U.S.C. 103 as being unpatentable over Beith (US 20240078732A1) and Villanueva Aylagas (US012406419B1) in view of Li (US 20170243387A1) and Karras (US 20180336464A1).
Beith and Villanueva Aylagas are relied on for the teachings as discussed above relative to Claim 9.
However, Beith and Villanueva Aylagas do not teach wherein: the one or more neural networks are trained using a first loss function that is associated with the upper portion of the facial representation of the character; and the one or more neural networks are trained using a second loss function that is associated with a lower portion of the facial representation of the character. However, Li teaches wherein: the one or more first neural networks are trained based at least on animating the upper portion of the facial representation of the character [0047]; and the one or more second neural networks are trained based at least on animating a lower portion of the facial representation of the character [0049]. This would be obvious for the reasons given in the rejection for Claim 1.
However, Beith, Villanueva Aylagas, and Li do not teach wherein: the one or more neural networks are trained using a first loss function that is associated with the upper portion of the facial representation of the character; and the one or more neural networks are trained using a second loss function that is associated with a lower portion of the facial representation of the character. However, Karras teaches the one or more neural networks are trained using a loss function that is associated with the facial representation of the character (training the network involves the steps of: comparing the output of the network with a desired target output in the training dataset using a loss function, the network parameters are then updated based on the result of the loss function, [0092], [0083, 0090]). Thus, this teaching of the loss function from Karras can be implemented into the neural networks of Li so that the one or more neural networks are trained using a first loss function that is associated with the upper portion of the facial representation of the character; and the one or more neural networks are trained using a second loss function that is associated with a lower portion of the facial representation of the character.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Beith, Villanueva Aylagas, and Li so that the one or more neural networks are trained using a first loss function that is associated with the upper portion of the facial representation of the character; and the one or more neural networks are trained using a second loss function that is associated with a lower portion of the facial representation of the character because Karras suggests that this way, the neural network can be trained to be more accurate [0092].
Claim(s) 21 is/are rejected under 35 U.S.C. 103 as being unpatentable over Beith (US 20240078732A1) and Villanueva Aylagas (US012406419B1) in view of McDuff (US 20200279553A1).
Beith and Villanueva Aylagas are relied upon for the teachings as discussed above relative to Claim 9. Beith teaches wherein the one or more neural networks are trained such that: the second emotional state drives animation of the upper portion of the facial representation of the character [0080, 0083]; and the audio data drives animation of a lower portion of the facial representation of the character [0135, 0137].
However, Beith and Villanueva Aylagas do not teach the second emotional state mostly drives animation of the upper portion of the facial representation of the character; and the audio data mostly drives animation of a lower portion of the facial representation of the character. However, McDuff teaches the second emotional state mostly drives animation of the upper portion of the facial representation of the character [0080]; and the audio data mostly drives animation of a lower portion of the facial representation of the character [0069].
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Beith and Villanueva Aylagas so that the second emotional state mostly drives animation of the upper portion of the facial representation of the character; and the audio data mostly drives animation of a lower portion of the facial representation of the character because McDuff suggests that the emotions are expressed more by the upper face [0080], and the audio is used to animate the lips so that it looks like the conversational agent is saying the words in the audio [0069].
Allowable Subject Matter
Claims 6 and 19 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
The following is a statement of reasons for the indication of allowable subject matter:
The prior art taken singly or in combination do not teach or suggest the combination of all the limitations of Claim 6 and base Claim 1, and in particular, do not teach wherein: the one or more second neural networks are further trained based at least on displacing the second points of the upper portion of the facial representation based on the emotional state while refraining from displacing the first points of the lower portion of the facial animation based on the emotional state.
The prior art taken singly or in combination do not teach or suggest the combination of all the limitations of Claim 19 and base Claim 18, and in particular, do not teach determining, based at least on the first output, a first vector corresponding to the emotional state; and determine, based at least on one or more user inputs associated with updating the emotional state, a second vector corresponding to a second emotional state, wherein the second output is determined using the one or more second neural networks and based at least on the audio data and the second vector.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JONI HSU whose telephone number is (571)272-7785. The examiner can normally be reached M-F 10am-6:30pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kee Tung can be reached at (571)272-7794. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
JH
/JONI HSU/Primary Examiner, Art Unit 2611