Prosecution Insights
Last updated: August 08, 2026
Application No. 18/456,002

MODEL LEARNING SYSTEM, MODEL LEARNING METHOD, A NON-TRANSITORY COMPUTER-READABLE RECORDING MEDIUM, AN ANIMATION GENERATION SYSTEM, AND AN ANIMATION GENERATION METHOD

Non-Final OA §103
Filed
Aug 25, 2023
Priority
Aug 25, 2022 — JP 2022-133949
Examiner
BECKER, TYLER JUSTIN
Art Unit
2657
Tech Center
2600 — Communications
Assignee
Square Enix Co., Ltd.
OA Round
3 (Non-Final)
74%
Grant Probability
Favorable
3-4
OA Rounds
0m
Est. Remaining
80%
With Interview

Examiner Intelligence

Grants 74% — above average
74%
Career Allowance Rate
17 granted / 23 resolved
+11.9% vs TC avg
Moderate +6% lift
Without
With
+6.3%
Interview Lift
resolved cases with interview
Typical timeline
2y 8m
Avg Prosecution
15 currently pending
Career history
45
Total Applications
across all art units

Statute-Specific Performance

§101
19.6%
-20.4% vs TC avg
§103
50.3%
+10.3% vs TC avg
§102
14.5%
-25.5% vs TC avg
§112
15.6%
-24.4% vs TC avg
Black line = Tech Center average estimate • Based on career data from 23 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Continued Examination Under 37 CFR 1.114 A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on February 20th, 2026 has been entered. Response to Amendment The amendment filed February 20th, 2026 has been entered. Claims 1, 3, and 4 have been amended. Claims 1-7 are pending and have been examined. Response to Arguments Applicant’s arguments with respect to claim(s) 1-7 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1-7 is/are rejected under 35 U.S.C. 103 as being unpatentable over Yu et al. (US Pat. Pub. No. 2021/0375260 A1 hereinafter Yu), in view of Brady et al. (US Pat. Pub. No. 2020/0034025 A1 hereinafter Brady) and Comer et al. (US Pat. Pub. No. 2021/0012549 A1 hereinafter Comer). Regarding claim 1, Yu discloses a model learning system, comprising: one or more processors, and a non-transitory computer readable medium storing computer-executable instructions which, when executed, cause the one or more processors to perform operations comprising (Yu, [0012]: "The device includes a memory, storing computer-executable instructions; and a processor, coupled with the memory and, when the computer-executable instructions being executed, configured to:): voice model learning including: extracting an acoustic feature value by executing predetermined acoustic signal processing with respect to voice data including human voice (Yu, [0057]: "In certain embodiments, parameters of the neural network may be generated by a training process configured to receive training data containing a plurality of audio signals and ground-truth phoneme annotations."; [0054]: " Step S220 is to transform the audio input signal into frequency-domain audio features."); and extracting a voice feature value by executing first transformation processing with respect to first input information including the extracted acoustic feature value (Yu, [0057]: "Step S230 is to perform neural-network processing on the frequency-domain audio features to recognize phonemes."); and rig model learning including: extracting a frame feature value by executing second transformation processing with respect to second input information including the extracted voice feature value (Yu, [0061]: "According to certain embodiments, the phoneme recognition process does not directly predict phoneme label but instead predicts HMM tied states since a combination of neural network and HMM decoder may further improve system performance. The predicted HMM tied-states in B.sub.2 may be used as input to the HMM tri-phone decoder. For each time interval br.Math.A, the decoder may take all the predicted data in B.sub.2, calculate corresponding phonemes [P.sub.t, . . . , P.sub.t−br+1], and store the phonemes it in buffer B.sub.3."; [0062]: "An additional buffer, B.sub.3, may be used to control the final output of P of the phoneme recognition system. This output predicted phoneme with a timestamp may be stored into B.sub.3. The method may continuously take the first phoneme from B.sub.3 and use it as input to an animation step every A ms time interval."), and outputting character control information for controlling a character from the extracted frame feature value (Yu, [0063]: "Step S240 of the speech animation generation method is to produce an animation according to the recognized phoneme."). However, Yu fails to expressly recite wherein the acoustic feature value is normalized based on style information including a language, a character, and a rig; and outputting character control information for controlling a character from the extracted frame feature value, wherein outputting the character control information includes outputting the character control information based on pose weights for generating a pose set corresponding to an emotion of the character and emotion weights corresponding to all possible emotions. Brady teaches wherein the acoustic feature value is normalized based on style information including a language, a character, and a rig (Brady, fig. 4; [0058]: “At step 422, the one or more customizations selections modify the expressed phrase. Modifying the expressed phrase using the selections can comprise changing suprasegmentals of the expressed phrase. A suprasegmental is a feature of speech which may comprise stress, tone, word juncture, prominence, pitch or any other desirable suprasegmental. Every selection, the gender selection 514, the age selection 516, the emotion selection 518, the race selection 520, the location selection 522, the nationality selection 524, and the language selection 526 may modify the expressed phrase using the elements of suprasegmentals. The gender selection 514 may change the gender of the expressed phrase. If the user selects a female then the expressed phrase will be spoken by a female voice. The age selection 516 may change the pitch of the expressed phrase. If the user desires a child to speak the expressed phrase then increasing the pitch simulates the expressed phrase being spoken by a young child. The emotion selection 518 may change the emotion of the expressed phrase by adding the user desired emotion to the expressed phrase. If the user desires the expressed phrase to be spoken by someone happy, then the user will select the emotion happiness and the expressed phrase will be spoken to simulate happiness. The race selection 520 may change the expressed phrase to simulate the expressed phrase being spoken by someone of the desired race. The location selection 522 may further adjust suprasegmentals of the expressed phrase by linking certain dialects of speech to the location selection. The nationality selection 524 may adjust the suprasegmentals of the expressed phrase by modifying the expressed to simulate different nationalities throughout the world. The language selection 526 may modify the expressed phrase by changing the language of the expressed phrase. Any language may be selected such as Chinese, Spanish, English, Hindi, Arabic, Portuguese, German, or any other desirable language.”; [0056]: “At step 418, the one or more customizations selections modify the avatar.”; Here, modifying the expressed phrase is seen as normalizing the acoustic feature value. Further, the customization selections are seen as rig and character information since they are also used to modify the avatar.). Yu and Brady are analogous arts because they both belong to the same field of virtual animation systems. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the speech animation generation system of Yu to incorporate the teachings of Brady to normalize the voice feature value based on a variety of other data. Being able to modify the voice feature based on a corresponding avatar can help overcome communication barriers with people of different backgrounds (Brady, [0002]). This ensures the user can communicate more easily with many different people. However, Yu, in view of Brady, Fails to expressly recite outputting character control information for controlling a character from the extracted frame feature value, wherein outputting the character control information includes outputting the character control information based on pose weights for generating a pose set corresponding to an emotion of the character and emotion weights corresponding to all possible emotions. Comer teaches outputting character control information for controlling a character from the extracted frame feature value, wherein outputting the character control information includes outputting the character control information based on pose weights for generating a pose set corresponding to an emotion of the character and emotion weights corresponding to all possible emotions (Comer, [0132]: “A skeletal system for a virtual character can be defined with joints at appropriate positions, and with appropriate local axes of rotation, degrees of freedom, etc., to allow for a desired set of mesh deformations to be carried out. Once a skeletal system has been defined for a virtual character, each joint can be assigned, in a process called “skinning,” an amount of influence over the various vertices in the mesh. This can be done by assigning a weight value to each vertex for each joint in the skeletal system. When a transform is applied to any given joint, the vertices under its influence can be moved, or otherwise altered, automatically based on that joint transform by amounts which can be dependent upon their respective weight values.”; [0174]: “FIG. 12 schematically illustrates an example of a facial rig 1200 that can be used to animate an avatar in an AR/VR/MR system. The facial rig 1200 can be thought of as a digital puppet in which points of articulation are parameterized into a plurality of facial rig parameters 1204, 1208. The facial rig 1200 electronically adjusts values for the points of articulation to generate a desired facial expression or emotion of the avatar.”; [0178]: “Although many of the AU variants may be based on FACS, some expressions may be different from traditional FACS groupings. For example, the disclosed FACS rig system may utilize different AUs or different intensity scales to represent an emotion. Accordingly, the facial rig parameters 1204, 1208 can be coarsely mapped in real time. For example, if a real person smiles, the Facial Solver may interpret this as happy, and the person's avatar can be manipulated as a “happy” category (rather than point-to-point matching of the real person). This may simplify the computational load of the system.”). Yu, Brady, and Comer are analogous arts because they each belong to the same field of virtual animation systems. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the speech animation generation system of Yu, as modified by the system for multisensory semiotic communications of Brady, to incorporate the teachings of Comer to control a character based on pose weights and emotion information. Using pose weights and emotion information helps animate a high quality or high fidelity avatar while still being time and resource efficient (Comer, [0029]). This ensures that an avatar is controlled accurately and efficiently for a user. PNG media_image1.png 784 628 media_image1.png Greyscale Brady, Fig. 4 for reference Regarding claim 2, the rejection of claim 1 is incorporated. Yu, in view of Brady and Comer, discloses all of the elements of the current invention as stated above. Yu further discloses a training data storage configured to store training data including voice and information related to an animation of the character as an answer (Yu, [0057]: "In certain embodiments, parameters of the neural network may be generated by a training process configured to receive training data containing a plurality of audio signals and ground-truth phoneme annotations."), wherein the operations further comprise updating parameters of the voice model and the rig model based on a difference between the information related to the animation of the character included in the training data and the character control information output using the training data (Yu, [0084]: "Following a standard TIMIT set up, a standard TIMIT training set including 3696 utterances from 462 speakers, is used to train the neural network. The performance of the neural network is then evaluated on a TIMIT core test set which consists of 192 utterances."; [0085]: " For network training, 10 epochs are configured as the minimum training epoch and early stop is enabled (if validate loss change in epochs is less than 0.001) during the training."). Regarding claim 3, Yu discloses a model learning method comprising: extracting an acoustic feature value by executing predetermined acoustic signal processing with respect to voice data including human voice (Yu, [0054]: " Step S220 is to transform the audio input signal into frequency-domain audio features."); extracting a voice feature value by executing first transformation processing with respect to first input information including the extracted acoustic feature value (Yu, [0057]: "Step S230 is to perform neural-network processing on the frequency-domain audio features to recognize phonemes."); extracting a frame feature value by executing second transformation processing with respect to second input information including the extracted voice feature value (Yu, [0061]: "According to certain embodiments, the phoneme recognition process does not directly predict phoneme label but instead predicts HMM tied states since a combination of neural network and HMM decoder may further improve system performance. The predicted HMM tied-states in B.sub.2 may be used as input to the HMM tri-phone decoder. For each time interval br.Math.A, the decoder may take all the predicted data in B.sub.2, calculate corresponding phonemes [P.sub.t, . . . , P.sub.t−br+1], and store the phonemes it in buffer B.sub.3."; [0062]: "An additional buffer, B.sub.3, may be used to control the final output of P of the phoneme recognition system. This output predicted phoneme with a timestamp may be stored into B.sub.3. The method may continuously take the first phoneme from B.sub.3 and use it as input to an animation step every A ms time interval."); and outputting character control information for controlling a character from the extracted frame feature value (Yu, [0063]: "Step S240 of the speech animation generation method is to produce an animation according to the recognized phoneme."). However, Yu fails to expressly recite wherein the acoustic feature value is normalized based on style information including a language, a character, and a rig; and outputting character control information for controlling a character from the extracted frame feature value, wherein the method further comprises outputting the character control information based on pose weights for generating a pose set corresponding to an emotion of the character and emotion weights corresponding to all possible emotions. Brady teaches wherein the acoustic feature value is normalized based on style information including a language, a character, and a rig (Brady, fig. 4; [0058]: “At step 422, the one or more customizations selections modify the expressed phrase. Modifying the expressed phrase using the selections can comprise changing suprasegmentals of the expressed phrase. A suprasegmental is a feature of speech which may comprise stress, tone, word juncture, prominence, pitch or any other desirable suprasegmental. Every selection, the gender selection 514, the age selection 516, the emotion selection 518, the race selection 520, the location selection 522, the nationality selection 524, and the language selection 526 may modify the expressed phrase using the elements of suprasegmentals. The gender selection 514 may change the gender of the expressed phrase. If the user selects a female then the expressed phrase will be spoken by a female voice. The age selection 516 may change the pitch of the expressed phrase. If the user desires a child to speak the expressed phrase then increasing the pitch simulates the expressed phrase being spoken by a young child. The emotion selection 518 may change the emotion of the expressed phrase by adding the user desired emotion to the expressed phrase. If the user desires the expressed phrase to be spoken by someone happy, then the user will select the emotion happiness and the expressed phrase will be spoken to simulate happiness. The race selection 520 may change the expressed phrase to simulate the expressed phrase being spoken by someone of the desired race. The location selection 522 may further adjust suprasegmentals of the expressed phrase by linking certain dialects of speech to the location selection. The nationality selection 524 may adjust the suprasegmentals of the expressed phrase by modifying the expressed to simulate different nationalities throughout the world. The language selection 526 may modify the expressed phrase by changing the language of the expressed phrase. Any language may be selected such as Chinese, Spanish, English, Hindi, Arabic, Portuguese, German, or any other desirable language.”; [0056]: “At step 418, the one or more customizations selections modify the avatar.”; Here, modifying the expressed phrase is seen as normalizing the acoustic feature value. Further, the customization selections are seen as rig and character information since they are also used to modify the avatar.). Yu and Brady are analogous arts because they both belong to the same field of virtual animation systems. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the speech animation generation system of Yu to incorporate the teachings of Brady to normalize the voice feature value based on a variety of other data. Being able to modify the voice feature based on a corresponding avatar can help overcome communication barriers with people of different backgrounds (Brady, [0002]). This ensures the user can communicate more easily with many different people. However, Yu, in view of Brady, fails to expressly recite outputting character control information for controlling a character from the extracted frame feature value, wherein the method further comprises outputting the character control information based on pose weights for generating a pose set corresponding to an emotion of the character and emotion weights corresponding to all possible. Comer teaches outputting character control information for controlling a character from the extracted frame feature value, wherein the method further comprises outputting the character control information based on pose weights for generating a pose set corresponding to an emotion of the character and emotion weights corresponding to all possible (Comer, [0132]: “A skeletal system for a virtual character can be defined with joints at appropriate positions, and with appropriate local axes of rotation, degrees of freedom, etc., to allow for a desired set of mesh deformations to be carried out. Once a skeletal system has been defined for a virtual character, each joint can be assigned, in a process called “skinning,” an amount of influence over the various vertices in the mesh. This can be done by assigning a weight value to each vertex for each joint in the skeletal system. When a transform is applied to any given joint, the vertices under its influence can be moved, or otherwise altered, automatically based on that joint transform by amounts which can be dependent upon their respective weight values.”; [0174]: “FIG. 12 schematically illustrates an example of a facial rig 1200 that can be used to animate an avatar in an AR/VR/MR system. The facial rig 1200 can be thought of as a digital puppet in which points of articulation are parameterized into a plurality of facial rig parameters 1204, 1208. The facial rig 1200 electronically adjusts values for the points of articulation to generate a desired facial expression or emotion of the avatar.”; [0178]: “Although many of the AU variants may be based on FACS, some expressions may be different from traditional FACS groupings. For example, the disclosed FACS rig system may utilize different AUs or different intensity scales to represent an emotion. Accordingly, the facial rig parameters 1204, 1208 can be coarsely mapped in real time. For example, if a real person smiles, the Facial Solver may interpret this as happy, and the person's avatar can be manipulated as a “happy” category (rather than point-to-point matching of the real person). This may simplify the computational load of the system.”). Yu, Brady, and Comer are analogous arts because they each belong to the same field of virtual animation systems. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the speech animation generation system of Yu, as modified by the system for multisensory semiotic communications of Brady, to incorporate the teachings of Comer to control a character based on pose weights and emotion information. Using pose weights and emotion information helps animate a high quality or high fidelity avatar while still being time and resource efficient (Comer, [0029]). This ensures that an avatar is controlled accurately and efficiently for a user. Regarding claim 4, Yu discloses a non-transitory computer-readable recording medium having recorded thereon instructions that when executed by a computer apparatus, cause the computer apparatus to perform operations comprising (Yu, [0012]: "The device includes a memory, storing computer-executable instructions; and a processor, coupled with the memory and, when the computer-executable instructions being executed, configured to:): voice model learning including: extracting an acoustic feature value by executing predetermined acoustic signal processing with respect to voice data including human voice (Yu, [0057]: "In certain embodiments, parameters of the neural network may be generated by a training process configured to receive training data containing a plurality of audio signals and ground-truth phoneme annotations."; [0054]: " Step S220 is to transform the audio input signal into frequency-domain audio features."); and extracting a voice feature value by executing first transformation processing with respect to first input information including the extracted acoustic feature value (Yu, [0057]: "Step S230 is to perform neural-network processing on the frequency-domain audio features to recognize phonemes."); and rig model learning including: extracting a frame feature value by executing second transformation processing with respect to second input information including the extracted voice feature value (Yu, [0061]: "According to certain embodiments, the phoneme recognition process does not directly predict phoneme label but instead predicts HMM tied states since a combination of neural network and HMM decoder may further improve system performance. The predicted HMM tied-states in B.sub.2 may be used as input to the HMM tri-phone decoder. For each time interval br.Math.A, the decoder may take all the predicted data in B.sub.2, calculate corresponding phonemes [P.sub.t, . . . , P.sub.t−br+1], and store the phonemes it in buffer B.sub.3."; [0062]: "An additional buffer, B.sub.3, may be used to control the final output of P of the phoneme recognition system. This output predicted phoneme with a timestamp may be stored into B.sub.3. The method may continuously take the first phoneme from B.sub.3 and use it as input to an animation step every A ms time interval."); and outputting character control information for controlling a character from the extracted frame feature value (Yu, [0063]: "Step S240 of the speech animation generation method is to produce an animation according to the recognized phoneme."). However, Yu fails to expressly recite wherein the acoustic feature value is normalized based on style information including a language, a character, and a rig; and outputting character control information for controlling a character from the extracted frame feature value, wherein outputting the character control information includes outputting the character control information based on pose weights for generating a pose set corresponding to an emotion of the character and emotion weights corresponding to all possible emotions. Brady teaches wherein the acoustic feature value is normalized based on style information including a language, a character, and a rig (Brady, fig. 4; [0058]: “At step 422, the one or more customizations selections modify the expressed phrase. Modifying the expressed phrase using the selections can comprise changing suprasegmentals of the expressed phrase. A suprasegmental is a feature of speech which may comprise stress, tone, word juncture, prominence, pitch or any other desirable suprasegmental. Every selection, the gender selection 514, the age selection 516, the emotion selection 518, the race selection 520, the location selection 522, the nationality selection 524, and the language selection 526 may modify the expressed phrase using the elements of suprasegmentals. The gender selection 514 may change the gender of the expressed phrase. If the user selects a female then the expressed phrase will be spoken by a female voice. The age selection 516 may change the pitch of the expressed phrase. If the user desires a child to speak the expressed phrase then increasing the pitch simulates the expressed phrase being spoken by a young child. The emotion selection 518 may change the emotion of the expressed phrase by adding the user desired emotion to the expressed phrase. If the user desires the expressed phrase to be spoken by someone happy, then the user will select the emotion happiness and the expressed phrase will be spoken to simulate happiness. The race selection 520 may change the expressed phrase to simulate the expressed phrase being spoken by someone of the desired race. The location selection 522 may further adjust suprasegmentals of the expressed phrase by linking certain dialects of speech to the location selection. The nationality selection 524 may adjust the suprasegmentals of the expressed phrase by modifying the expressed to simulate different nationalities throughout the world. The language selection 526 may modify the expressed phrase by changing the language of the expressed phrase. Any language may be selected such as Chinese, Spanish, English, Hindi, Arabic, Portuguese, German, or any other desirable language.”; [0056]: “At step 418, the one or more customizations selections modify the avatar.”; Here, modifying the expressed phrase is seen as normalizing the acoustic feature value. Further, the customization selections are seen as rig and character information since they are also used to modify the avatar.). Yu and Brady are analogous arts because they both belong to the same field of virtual animation systems. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the speech animation generation system of Yu to incorporate the teachings of Brady to normalize the voice feature value based on a variety of other data. Being able to modify the voice feature based on a corresponding avatar can help overcome communication barriers with people of different backgrounds (Brady, [0002]). This ensures the user can communicate more easily with many different people. However, Yu, in view of Brady, fails to expressly recite outputting character control information for controlling a character from the extracted frame feature value, wherein outputting the character control information includes outputting the character control information based on pose weights for generating a pose set corresponding to an emotion of the character and emotion weights corresponding to all possible emotions. Comer teaches outputting character control information for controlling a character from the extracted frame feature value, wherein outputting the character control information includes outputting the character control information based on pose weights for generating a pose set corresponding to an emotion of the character and emotion weights corresponding to all possible emotions (Comer, [0132]: “A skeletal system for a virtual character can be defined with joints at appropriate positions, and with appropriate local axes of rotation, degrees of freedom, etc., to allow for a desired set of mesh deformations to be carried out. Once a skeletal system has been defined for a virtual character, each joint can be assigned, in a process called “skinning,” an amount of influence over the various vertices in the mesh. This can be done by assigning a weight value to each vertex for each joint in the skeletal system. When a transform is applied to any given joint, the vertices under its influence can be moved, or otherwise altered, automatically based on that joint transform by amounts which can be dependent upon their respective weight values.”; [0174]: “FIG. 12 schematically illustrates an example of a facial rig 1200 that can be used to animate an avatar in an AR/VR/MR system. The facial rig 1200 can be thought of as a digital puppet in which points of articulation are parameterized into a plurality of facial rig parameters 1204, 1208. The facial rig 1200 electronically adjusts values for the points of articulation to generate a desired facial expression or emotion of the avatar.”; [0178]: “Although many of the AU variants may be based on FACS, some expressions may be different from traditional FACS groupings. For example, the disclosed FACS rig system may utilize different AUs or different intensity scales to represent an emotion. Accordingly, the facial rig parameters 1204, 1208 can be coarsely mapped in real time. For example, if a real person smiles, the Facial Solver may interpret this as happy, and the person's avatar can be manipulated as a “happy” category (rather than point-to-point matching of the real person). This may simplify the computational load of the system.”). Yu, Brady, and Comer are analogous arts because they each belong to the same field of virtual animation systems. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the speech animation generation system of Yu, as modified by the system for multisensory semiotic communications of Brady, to incorporate the teachings of Comer to control a character based on pose weights and emotion information. Using pose weights and emotion information helps animate a high quality or high fidelity avatar while still being time and resource efficient (Comer, [0029]). This ensures that an avatar is controlled accurately and efficiently for a user. Regarding claim 5, the rejection of claim 1 is incorporated. Yu, in view of Brady and Comer, discloses all of the elements of the current invention as stated above. Yu further discloses one or more processors, and a non-transitory computer readable medium storing computer-executable instructions which, when executed, cause the one or more processors to perform operations comprising (Yu, [0012]: "The device includes a memory, storing computer-executable instructions; and a processor, coupled with the memory and, when the computer-executable instructions being executed, configured to:"): extracting a voice feature value using a voice model, the voice model configured to receive voice data comprising human voice as input and further configured to extract the voice feature value from the voice data, the voice model being trained by the model learning system of claim 1 (Yu, [0054]: " Step S220 is to transform the audio input signal into frequency-domain audio features."; [0057]: "Step S230 is to perform neural-network processing on the frequency-domain audio features to recognize phonemes."); outputting character control information using a rig model, the rig model configured to receive second input information comprising the voice feature value as input and further configured to output the character control information for controlling a character from the second input information including the voice feature value, the rig model being trained by the model learning system of claim 1 (Yu, [0061]: "According to certain embodiments, the phoneme recognition process does not directly predict phoneme label but instead predicts HMM tied states since a combination of neural network and HMM decoder may further improve system performance. The predicted HMM tied-states in B.sub.2 may be used as input to the HMM tri-phone decoder. For each time interval br.Math.A, the decoder may take all the predicted data in B.sub.2, calculate corresponding phonemes [P.sub.t, . . . , P.sub.t−br+1], and store the phonemes it in buffer B.sub.3."; [0062]: "An additional buffer, B.sub.3, may be used to control the final output of P of the phoneme recognition system. This output predicted phoneme with a timestamp may be stored into B.sub.3. The method may continuously take the first phoneme from B.sub.3 and use it as input to an animation step every A ms time interval."); and generating an animation related to the character based on the character control information (Yu, [0063]: "Step S240 of the speech animation generation method is to produce an animation according to the recognized phoneme."). Regarding claim 6, the rejection of claim 3 is incorporated. Yu, in view of Brady and Comer, discloses all of the elements of the current invention as stated above. Yu further discloses extracting a voice feature value using a voice model, the voice model configured to receive voice data including human voice as input and further configured to extract the voice feature value from the voice data, the voice model being trained using the model learning method of claim 3 (Yu, [0054]: " Step S220 is to transform the audio input signal into frequency-domain audio features."; [0057]: "Step S230 is to perform neural-network processing on the frequency-domain audio features to recognize phonemes."); outputting character control information using a rig model, the rig model configured to receive second input information including the voice feature value as input and further configured to output the character control information for controlling a character from the second input information including the voice feature value, the rig model being trained using the model learning method according to claim 3 (Yu, [0061]: "According to certain embodiments, the phoneme recognition process does not directly predict phoneme label but instead predicts HMM tied states since a combination of neural network and HMM decoder may further improve system performance. The predicted HMM tied-states in B.sub.2 may be used as input to the HMM tri-phone decoder. For each time interval br.Math.A, the decoder may take all the predicted data in B.sub.2, calculate corresponding phonemes [P.sub.t, . . . , P.sub.t−br+1], and store the phonemes it in buffer B.sub.3."; [0062]: "An additional buffer, B.sub.3, may be used to control the final output of P of the phoneme recognition system. This output predicted phoneme with a timestamp may be stored into B.sub.3. The method may continuously take the first phoneme from B.sub.3 and use it as input to an animation step every A ms time interval."); and generating an animation related to the character based on the character control information (Yu, [0063]: "Step S240 of the speech animation generation method is to produce an animation according to the recognized phoneme."). Regarding claim 7, the rejection of claim 4 is incorporated. Yu, in view of Brady and Comer, discloses all of the elements of the current invention as stated above. Yu further discloses extracting a voice feature value using a voice model configured to receive voice data including human voice as input and further configured to extract the voice feature value from the voice data, the voice model being trained by the voice model learning of claim 4 (Yu, [0054]: " Step S220 is to transform the audio input signal into frequency-domain audio features."; [0057]: "Step S230 is to perform neural-network processing on the frequency-domain audio features to recognize phonemes."); outputting character control information using a rig model configured to receive second input information including the voice feature value as input and further configured to output the character control information for controlling a character from the second input information including the voice feature value, the rig model being trained by the rig model learning of claim 4 (Yu, [0061]: "According to certain embodiments, the phoneme recognition process does not directly predict phoneme label but instead predicts HMM tied states since a combination of neural network and HMM decoder may further improve system performance. The predicted HMM tied-states in B.sub.2 may be used as input to the HMM tri-phone decoder. For each time interval br.Math.A, the decoder may take all the predicted data in B.sub.2, calculate corresponding phonemes [P.sub.t, . . . , P.sub.t−br+1], and store the phonemes it in buffer B.sub.3."; [0062]: "An additional buffer, B.sub.3, may be used to control the final output of P of the phoneme recognition system. This output predicted phoneme with a timestamp may be stored into B.sub.3. The method may continuously take the first phoneme from B.sub.3 and use it as input to an animation step every A ms time interval."); and generating an animation related to the character based on the character control information (Yu, [0063]: "Step S240 of the speech animation generation method is to produce an animation according to the recognized phoneme."). Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Kolluru et al. (US Pat. Pub. No. 2015/0052084 A1) discloses a system for computer generated emulation of a subject. Any inquiry concerning this communication or earlier communications from the examiner should be directed to TYLER J BECKER whose telephone number is (703)756-1271. The examiner can normally be reached M-Th, 7:15am-5:45pm PT. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Daniel Washburn can be reached at (571) 272-5551. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /TYLER BECKER/ Examiner, Art Unit 2657 /DANIEL C WASHBURN/ Supervisory Patent Examiner, Art Unit 2657
Read full office action

Prosecution Timeline

Aug 25, 2023
Application Filed
Jun 13, 2025
Non-Final Rejection mailed — §103
Sep 04, 2025
Response Filed
Nov 21, 2025
Final Rejection mailed — §103
Feb 20, 2026
Request for Continued Examination
Feb 24, 2026
Response after Non-Final Action
May 05, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12694228
REAL-TIME USER COMMUNICATION SENTIMENT DETECTION FOR DYNAMIC ANOMALY DETECTION AND MITIGATION
3y 4m to grant Granted Jul 28, 2026
Patent 12682113
SYSTEMS, METHODS, AND APPARATUSES FOR GENERATING STRUCTURED DATA FROM UNSTRUCTURED DATA USING NATURAL LANGUAGE PROCESSING TO GENERATE A SECURE MEDICAL DASHBOARD
3y 0m to grant Granted Jul 14, 2026
Patent 12651592
SYSTEM, METHOD, AND COMPUTER PROGRAM FOR REAL-TIME LANGUAGE TRANSLATION USING GENERATIVE ARTIFICIAL INTELLIGENCE
3y 0m to grant Granted Jun 09, 2026
Patent 12632657
Joint Speech and Text Streaming Model for ASR
2y 10m to grant Granted May 19, 2026
Patent 12614560
REVERBERATION REMOVAL DEVICE, PARAMETER ESTIMATION DEVICE, REVERBERATION REMOVAL METHOD, PARAMETER ESTIMATION METHOD, AND PROGRAM
2y 9m to grant Granted Apr 28, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
74%
Grant Probability
80%
With Interview (+6.3%)
2y 8m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 23 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month