DETAILED ACTION
Notice of Pre-AIA or AIA Status
1. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Arguments/Amendments
2. With respect to Claim Rejection 35 U.S.C § 102/103 towards Claims 1-5, Applicant’s arguments have been considered but are moot because the new ground to rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenge in the argument.
With respect to Claims 6-20, all of the rejections in the most recent Office action are overcome. Thus, the rejections have been withdrawn.
Claim Rejections - 35 USC § 103
3. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
4. Claims 1, 4-5 are rejected under 35 U.S.C.103 as being unpatentable over Park et al. (US 2025/0200855 A1) in view of Kitagishi (US 2014/0019137 A1.)
With respect to Claim 1, Park et al disclose
A method comprising:
generating, by one or more first models and based at least on processing user data (Park et al. [0012] capture a facial image of the user) and input data representing a prompt (Park et al. [0050-0051] receiving the question of the user basically include a text input), first output data representative of one or more emotion attributes associated with a user and one or more response attributes associated with a response to the prompt (Park et al. [0073] describes model trained based on a machine-learning (ML)-based modeling method for detecting the facial area from a facial image, [0012] describes recognizing an emotion of the user from the facial image of the user, [0094] describes determining an emotion of the response);
generating, by one or more second models and based at least on processing the input data and the first output data, second output data representative of the response to the prompt and one or more tags corresponding to a voice that is related to the one or more emotion attributes and the one or more response attributes (Park et al. [0022 and 0094] describes receiving a user’s question (e.g., a prompt), processing the received user’s input and the emotion of the user from the facial image of the user to generate the response having the content in compliance with an emotion of the user by reflecting the multimodal recognition result with respect to the emotion of the user, [0080] determining one or more intensity value associated with the one or more labels (e.g., the valence of an emotion is a positive force of 7). See paragraphs [0091-0092 and 0096]);
generating, based at least on processing the second output data, audio data representative of speech corresponding to the response and expressed using the voice (Park et al. [0094] a response with respect to the word of the user “hello” as a prompt sentence of the LLM may be generated, wherein an emotion of the response should be the arousal of A and the valence of B (A and B are the arousal and valence states of the user, estimated in the multimodal emotion recognition in B. of operation S305 above). The content of the word generated by the LLM may be converted into the voice through the TTS technique. Here, the tone and the manner of the voice of the virtual human may have to have the same arousal and valence as the content of the word generated by the LLM. When a gap occurs between the content and the arousal and the valence, a feeling of distance may occur between the content of the word and the voice expression, and thus, a hearer may feel it unnatural. For example, when the virtual human says “hello!” the tone and the manner must not be too dark or slow. Tagging the recognized emotion of the user into the response text to synthesize the audio response); and
Park et al. fail to explicitly teach
causing an output of an animation of a character and audio corresponding to the speech represented by the audio data.
However, Kitagishi teaches
causing an output of an animation of a character and audio corresponding to the speech represented by the audio data (Kitagishi [0146] While a speech synthesis system according to this embodiment is basically the same as the speech synthesis system according to the ninth embodiment, in a case where the application operating by the application operating unit is a generation animation, the speech synthesis terminal synchronizes output timing of the animation and the output timing of the synthesized speech synthesized by the speech synthesis unit are synchronized. By employing such a configuration of this embodiment having such a feature, in a voice animation, synthesized speech can be output with a feeling of character’s naturally speaking.)
Park et al. and Kitagishi are analogous art because they are from a similar field of endeavor in the Speech Processing techniques and applications. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the steps of generating a response to the user’s input based on the content of the user’s input and the user’s emotion as taught by Park et al., using teaching of synchronizing as taught by Kitagishi for the benefit of outputting the synthesized speech with the feeling of character’s naturally speaking (Kitagishi [0146] While a speech synthesis system according to this embodiment is basically the same as the speech synthesis system according to the ninth embodiment, in a case where the application operating by the application operating unit is a generation animation, the speech synthesis terminal synchronizes output timing of the animation and the output timing of the synthesized speech synthesized by the speech synthesis unit are synchronized. By employing such a configuration of this embodiment having such a feature, in a voice animation, synthesized speech can be output with a feeling of character’s naturally speaking.)
With respect to Claim 4, Park et al. in view of Kitagishi teach
further comprising:
obtaining dialogue data representative of at least one of one or more previous prompts (Park et al. [0047] describes the database store personal information, and previous input from the user) or one or more previous responses associated with the one or more previous prompts,
wherein the generating the first output data is further based at least on the one or more first models processing the dialogue data (Park et al. [0047] describes conversation based on the previous input from the user.)
With respect to Claim 5, Park et al. in view of Kitagishi teach
wherein the user data comprises at least one of:
text data representative of text describing one or more emotional states associated with a user;
audio data representative of user speech corresponding to the user (Park et al. [0077] receives the user’s utterance via a microphone);
video data representative of one or more videos corresponding to the user; or
image data representative of one or more images corresponding to the user.
Allowable Subject Matter
5. Claims 2-3 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
The following is a statement of reasons for the indication of allowable subject matter: the prior art(s) taken alone or in combination fail(s) to teach the following element(s) in the combination with the other recited elements in the claim(s).
“generating, based at least on the one or more first models processing character data, third output data representative of one or more second emotion attributes associated with the character,
wherein the generating of the second output data is further based at least on the one or more second models processing the third output data.” as recited in Claim 2.
“obtaining character data representative of one or more attributes associated with the character that is to output the speech, wherein the generating of the second output data is further based at least on the one or more second models processing the character data.” as recited in Claim 3.
Claims 6-20 are allowed.
The following is an examiner’s statement of reasons for allowance: the prior art(s) taken alone or in combination fail(s) to teach the following element(s) in the combination with the other recited elements in the claim(s).
“generate, by one or more first models and based at least on processing user data and character data, first output data representative of one or more first emotional states associated with a user and one or more second emotional states associated with a character;
generate, by one or more second models and based at least on processing the first output data and input data representative of first text, second output data representative of second text and information associated with a voice related to the one or more first emotional states and the one or more second emotional states;
generate, based at least on the second output data, audio data representative of speech corresponding to the second text and expressed using the voice; and
cause an output of an animation of the character and audio corresponding to the speech represented by the audio data.” as recited in Claim 6.
“generate, by one or more first models and based at least on processing user data and character data, first output data representative of one or more first attributes associated with a user and one or more second attributes associated with a character;
generate, by one or more second models and based at least on processing the first output data and input data representative of a prompt, second output data representative of a response for the character and information associated with a voice related to the response;
generate, based at least on the second output data, audio data representative of speech in the voice; and
cause an output of an animation of the character and audio corresponding to the speech represented by the audio data.” as recited in Claim 18.
None of the prior arts of records teach and/or suggest processing character data of a character animation combination with processing user data as input to generate emotional states/emotion attributes/second attributes associated with a character in order to generate a response for the character.
Conclusion
6. The prior art made of record and not relied upon is considered pertinent to application’s disclosure. See PTO-892.
a. Bonar et al. (US 2024/0169974 A1.) In this reference, Bonar et al. disclose the large language model to select an appropriate sentiment to attach to the text response based on the sentiment determined from the user input.
b. Jayaraman et al. (US 2024/0163232 A1.) In this reference, Jayaraman et al. disclose user the recorded user sentiment to determine future chat bot response.
c. Wang (US 2024/0096329 A1.) In this reference, Wang disclose adjusting the emotion of the response to be displayed using the cloned character voice model based on the detected emotion of the user.
7. Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any extension fee pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the date of this final action.
8. Any inquiry concerning this communication or earlier communications from the examiner should be directed to THUYKHANH LE whose telephone number is (571)272-6429. The examiner can normally be reached Mon-Fri: 9am-5pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Andrew C. Flanders can be reached on 571-272-7516. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/THUYKHANH LE/Primary Examiner, Art Unit 2655