DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Arguments
Applicant’s arguments, see p. 22 and 24, filed May 13, 2026, with respect to Claims 10-13, 17, 18, and 21 have been fully considered and are persuasive. The 35 U.S.C. 103 rejections of Claims 10-13, 17, 18, and 21 have been withdrawn.
As per Claim 10, Applicant argues that Chen (US 20240005905A1) does not teach any type of token that represents an emotion. The combination of Lembersky (US 20190095775A1) and Chen does not teach obtaining first input data representative of one or more learned tokens that cause one or more language models to generate one or more values corresponding to one or more speech characteristics for a character; generating, based at least on the one or more language models processing the one or more learned tokens and the one or more text tokens, output data representative of an emotional state associated with the one or more words and the one or more values corresponding to the one or more speech characteristics associated with the one or more words (p. 22, 1st paragraph).
In reply, the Examiner agrees. The rejections of Claims 10-13, 17, 18, and 21 have been withdrawn.
Applicant’s arguments with respect to claim(s) 1-9, 19, 20, 22, and 23 have been considered but are moot because new grounds of rejection are made in view of Baik (US 20180047391A1).
As per Claim 1, Applicant argues that Lembersky does not suggest applying both input data representing the user’s converted text and output data representing the answer to the AI engine, let alone output data from a different machine model. Lembersky does not suggest that a machine learning model processes such input data and output data to determine an emotional state associated with the answer. Lembersky does not suggest that the AI engine processes the user’s converted text and the answer to determine the specific emotion (p. 15-18).
In reply, the Examiner points out that new grounds of rejection are made in view of Baik (US 20180047391A1) to combine with Lembersky to teach generating, by a first machine learning model, the first output data; and applying to a second machine learning model that is different from the first machine learning model and as input, the first output data, and that a machine learning model processes such input data and output data to determine an emotional state associated with the answer.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claim(s) 1-3, 6, 9, 19, 20, and 22 is/are rejected under 35 U.S.C. 103 as being unpatentable over Lembersky (US 20190095775A1), Chen (US 20240005905A1), and Baik (US 20180047391A1).
As per Claim 1, Lembersky teaches a method comprising: receiving first input data representative of first text corresponding to a query (receive user input indicative of a user’s speech through an audio processor (speech-to-text), user’s converted text (speech), [0022]); generating, by a machine learning model and based at least on the first text corresponding to the query, first output data representative of second text (114) corresponding to a response to the query for output by a character (user’s converted text (speech) may then be passed to an AI engine 112 to determine a proper response 114 to the user (an answer to a question), which results in the proper text and emotional response being sent to a processor, which then translates the responsive text back to synthesized speech 118, and also triggers visual display blend shapes to morph a face of the AI character into a proper facial expression to convey the appropriate emotional response and mouth movement for the response, [0022]). Lembersky describes “determine an emotion of the user based on the aforementioned audio user input (e.g., the words themselves and/or how they are spoken) or visual user input (e.g., accomplished via facial recognition API)” [0101]. Lembersky describes “determine one or more avatar characteristics based on the one or both of the audio user input and the visual user input of the user. For instance, in various embodiments, the device can be configured to modify features of an avatar that is to be presented to the user based on the user(s) themselves, such as how to respond, with what emotion to display, with what words to say, with what tone to speak” [0102]. Thus, Lembersky teaches that in one embodiment, emotion could be determined by using facial recognition API. However, that is only one of the embodiments. Lembersky teaches in another embodiment, emotion is determined based only on words themselves. Lembersky teaches user’s converted text (speech) is then passed to AI engine 112 to determine a proper response 114 to the user (e.g., an answer to a question and specific emotion) [0022]. Thus, the first text (user’s converted text) is applied to machine learning model (AI engine 112). Since AI engine 112 determines the second text (answer to a question) and specific emotion [0022], this means that AI engine 112 processes the second text (answer to a question) in order to generate specific emotion (emotional state associated with the second text). Thus, Lembersky teaches applying, to a machine learning model (AI engine 112), the first output data and second input data representative of at least the first text (user’s converted text); generating, by the machine learning model (AI engine 112) and based at least on processing the first output data (answer to a question) and the second input data (user’s converted text), second output data representative of an emotional state associated with the second text (specific emotion) [0022]; generating, using one or more text-to-speech models and based at least on the second output data, audio data (118) representative of the speech that is associated with the second text and based at least on the emotional state (which results in the proper text and emotional response being sent to a processor, which then translates the responsive text back to synthesized speech 118, [0022]); and causing the character to be animated using at least the speech (triggers visual display blend shapes to morph a face of the AI character or avatar into a proper facial expression to convey the appropriate emotional response and mouth movement (lip synching) for the response, [0022]).
However, Lembersky does not teach generating, by the machine learning model processing the first output data and the second input data, one or more variables associated with at least one of the emotional state or speech associated with the second text; the speech is based on the one or more variables. However, Chen teaches generating, by the machine learning model processing data representative of the text, second output data representative of an emotional state associated with text and one or more variables associated with at least one of the emotional state or speech associated with the text; generating, using one or more text-to-speech models and based at least on the second output data, audio data representative of the speech that is associated with the text and based at least on the emotional state and the one or more variables (acoustic model, an optimal emotion intensity of the sample text input may be considered, and a mel spectrum closest to the optimal emotion intensity may be selected, which makes the emotion intensity of generated speech more reasonable and more in line with an actual need, [0085], emotion intensity extraction model 490 may be a machine learning model for determining the sample emotion intensity, [0198], sample may include the sample text input corresponding to multiple languages, so that the trained acoustic model may be capable of processing text input in multiple languages, [0068], generating a prediction speech corresponding to the text input based on the prediction mel spectrum, [0014]). Since Lembersky teaches generating, by the machine learning model and based at least on processing the first output data and the second input data, second output data representative of an emotional state associated with the second text [0022, 0005]; generating, using one or more text-to-speech models and based at least on the second output data, audio data (118) representative of the speech that is associated with the second text and based at least on the emotional state [0022], this teaching of one or more variables associated with at least one of the emotional state or speech from Chen can be implemented on the first text and the second text of Lembersky so that it generates, by the machine learning model and based at least on processing the first output data and the second input data, one or more variables associated with at least one of the emotional state or speech associated with the second text; the speech is based on the one or more variables.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Lembersky to include generating, by the machine learning model processing the first output data and the second input data, one or more variables associated with at least one of the emotional state or speech associated with the second text; the speech is based on the one or more variables because Chen suggests that this improves the naturalness and emotion richness of a synthesized speech [0003].
However, Lembersky and Chen do not teach generating, by a first machine learning model and based at least on the first text corresponding to the query, first output data corresponding to a request to the query for output by a character; applying, to a second machine learning model that is different from the first machine learning model and as input, the first output data and second input data representative of at least the first text; generating, by the second machine learning model and based at least on processing the first output data and the second input data, second output data representative of an emotional state associated with the second text. However, Baik teaches generating, by a first machine learning model (AI answering server 220) and based at least on the first text corresponding to the query, first output data corresponding to a request to the query for output by a character (AI answering server 220 may analyze a voice command included in the user speech, based on the text data from speech recognition server 210, based on the analysis result, AI answering server 220 may deduce artificial intelligent answer and generate data for answers in response to the voice command, [0069]); applying, to a character generating server 230 that is different from the first machine learning model (AI answering server 220) and as input, the first output data; generating, by the character generating server 230 and based at least on processing the first output data, second output data representative of an emotional state associated with a response to the query for output by a character (character generating server 230 may receive AI answer from AI answering server 220, determines an emotion type, an action type, and a sentence type based on the AI answer, the voice command, [0070]). Thus, this teaching of the AI answering server 220 from Baik can be implemented into the device of Lembersky so that the AI answering server is the first machine learning model and AI engine 112 is the second machine learning model that is different from the first machine learning model (AI answering server).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Lembersky and Chen to include generating, by a first machine learning model and based at least on the first text corresponding to the query, first output data corresponding to a request to the query for output by a character; applying, to a second machine learning model that is different from the first machine learning model and as input, the first output data and second input data representative of at least the first text; generating, by the second machine learning model and based at least on processing the first output data and the second input data, second output data representative of an emotional state associated with the second text because Baik suggests that an AI answering server can more accurately generate answers [0069], and using character generating server that is different from the AI answering server can more accurately determine the emotional state [0145].
As per Claim 2, Lembersky teaches wherein the machine learning model (AI engine 112) is trained to determine the emotional state based at least on the query and the response (AI engine 112 may thus be configured to consider all inputs available to it in order to learn determinations of user emotion, [0028], machine learning on top of AI so the virtual assistant can learn more about the user and make appropriate responses based on their past experiences, [0031], emotional responses will change over time based as the AI learns through multiple user interactions with machine learning to improve the response, [0057]).
However, Lembersky and Chen do not teach the first machine learning model is trained to determine the response based at least on the query; and the second machine learning model is trained to determine the emotional state based at least on the query and the response. However, Baik teaches the machine learning model (AI answering server 220) is trained to determine the response based at least on the query [0069]. Thus, this teaching of the AI answering server 220 from Baik can be implemented into the device of Lembersky so that the AI answering server is the first machine learning model and the AI engine 112 is the second machine learning model. This would be obvious for the reasons given in the rejection for Claim 1.
As per Claim 3, Lembersky does not teach wherein: the one or more variables include at least an intensity associated with the emotional state; and the second output data further represents a value associated with the intensity. However, Chen teaches wherein: the one or more variables include at least an intensity associated with the emotional state; and the second output data further represents a value associated with the intensity ([0085], emotion intensity may be quantified as a scale of 0-10, and a greater scale indicates a stronger corresponding emotion intensity, [0072]). This would be obvious for the reasons given in the rejection for Claim 1.
As per Claim 6, Lembersky teaches wherein the generating the second output data is further based at least on the machine learning model processing: third input data associated with the character, the third input data representative of at least one of one or more characteristics associated with the character, one or more situations associated with the character, or one or more interactions associated with the character, or one or more past communications associated with the character ([0005], virtual assistant can learn more about the user and make appropriate responses based on their past experiences, collect a historical activity database, the sentiment from the user using facial recognition, and stores this in their emotional history in the database of emotions and responses for a particular user, machine learning tools and techniques may then be used to improve the virtual assistant’s responses based on the user’s past experiences, the user will then be able to receive personalized greetings and suggestions, [0031], based on the user’s profile the virtual assistant can be a targeted personal advertisement directed at the user from the stores, for example, the virtual assistant could suggest salad place to eat based on John’s information and given an excited look and encouraged tone to stay on the diet, [0032]).
However, Lembersky and Chen do not teach that the machine learning model is the second machine learning model. However, Baik teaches the second machine learning model, as discussed in the rejection for Claim 1.
As per Claim 9, Lembersky teaches wherein: the second text includes one or more words; and the speech includes the one or more words spoken using the emotional state [0022].
However, Lembersky does not teach the speech includes the one or more words spoken based at least on the one or more variables. However, Chen teaches the speech includes the one or more words spoken based at least on the one or more variables [0085, 0198, 0068]. This would be obvious for the reasons given in the rejection for Claim 1.
As per Claim 19, Claim 19 is similar in scope to Claim 3, and therefore is rejected under the same rationale.
As per Claim 20, Lembersky teaches wherein the system is comprised in at least one of: a control system for an autonomous or semi-autonomous machine; a system for performing one or more simulation operations; a system for performing one or more digital twin operations; a system for performing light transport simulation; a system for performing collaborative content creation for 3D assets; a system for performing one or more deep learning operations; a system implemented using an edge device; a system implemented using a robot; a system for performing one or more generative AI operations; a system for performing operations using one or more large language models (LLMs); a system for performing one or more conversational AI operations; a system for generating synthetic data; a system for presenting at least one of virtual reality content, augmented reality content, or mixed reality content; a system incorporating one or more virtual machines (VMs); a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources [0022].
As per Claim 22, Lembersky teaches wherein the generation of the second audio data uses one or more speech-to-text models that process the second output data [0022].
Claim(s) 4-5 is/are rejected under 35 U.S.C. 103 as being unpatentable over Lembersky (US 20190095775A1), Chen (US 20240005905A1), and Baik (US 20180047391A1) in view of van der Meulen (US010360716B1).
As per Claim 4, Lembersky, Chen, and Baik are relied upon for the teachings as discussed above relative to Claim 1.
However, Lembersky, Chen, and Baik do not teach wherein: the one or more variables include one or more characteristics associated with the speech, the one or more characteristics including at least one of a volume, a rate, a pitch, or an emphasis associated with the speech; and the second output data further represents one or more values associated with the one or more characteristics. However, van der Meulen teaches wherein: the one or more variables include one or more characteristics associated with the speech, the one or more characteristics including at least one of a volume, a rate, a pitch, or an emphasis associated with the speech; and the second output data further represents one or more values associated with the one or more characteristics (message analyzer 208 may receive messages or other text and determine characteristics, through analysis, which may be used to enhance animation of an avatar that speaks the text, col. 5, lines 42-45; message analyzer 208 may markup text to indicate special audio features and/or special visual features, message analyzer 208 may identify information in the text, as well as a context of the message, message analyzer 208 may determine a speed, a pitch, a volume, and other attributes of speech, which may be included in the audio features and/or visual features, col. 8, lines 26-29, 42-45).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Lembersky, Chen, and Baik so the one or more variables include one or more characteristics associated with the speech, the one or more characteristics including at least one of volume, rate, a pitch, or an emphasis associated with the speech; and the second output data further represents one or more values associated with the one or more characteristics because van der Meulen suggests that for example if the message analyzer determines that the text has the emotion of surprise, then the speech rate is increased, since a person usually has an increased speech rate when they are surprised, and thus the speech that is output sounds more like when a real person is speaking when they are surprised (col. 9, line 61-col. 10, line 7).
As per Claim 5, Lembersky does not teach wherein: the one or more variables include at least an intensity level associated with the emotional state; the second output data further represents a first value associated with the intensity level; and the generating the audio data representative of the speech comprises generating, using the one or more text-to-speech models and based at least one the emotional state and the first value, the audio data such that the speech expresses the emotional state using the intensity level. However, Chen teaches wherein: the one or more variables include at least an intensity level associated with the emotional state; the second output data further represents a first value associated with the intensity level; and the generating the audio data representative of the speech comprises generating, using the one or more text-to-speech models and based at least one the emotional state and the first value, the audio data such that the speech expresses the emotional state using the intensity level [0072, 0085, 0198, 0014]. This would be obvious for the reasons given in the rejection for Claim 1.
However, Lembersky, Chen, and Baik do not teach wherein: the one or more variables include one or more characteristics associated with the speech; the second output data further represents one or more second values associated with one or more levels of the one or more characteristics; and the generating the audio data representative of the speech comprises generating, based on the one or more second values, the audio data such that the speech expresses the emotional state using the one or more levels of the one or more characteristics. However, van der Meulen teaches wherein: the one or more variables include one or more characteristics associated with the speech; the second output data further represents one or more second values associated with one or more levels of the one or more characteristics (col. 5, lines 42-45, col. 8, lines 26-29, 42-45); and the generating the audio data representative of the speech comprises generating, based on the one or more second values, the audio data such that the speech expresses the emotional state using the one or more levels of the one or more characteristics (ASML module 212 may create an audio indicator of <lower pitch for “step 2”>, which may be inserted into an output directed to the audio processor, col. 9, lines 20-23; ASML module 212 may create an audio indicator of “<speech rate +2>did you hear the news”, which may be inserted into an output directed to the audio processor, col. 10, lines 4-7; information that indicates an emotion of an originator of the message may be identified from text, punctuation, formatting, emoticons, word choice, and other information in the message or about the message and may be associated with one or more audio features, col. 11, lines 43-48). This would be obvious for reasons given in rejection for Claim 4.
Claim(s) 7-8 is/are rejected under 35 U.S.C. 103 as being unpatentable over Lembersky (US 20190095775A1), Chen (US 20240005905A1), and Baik (US 20180047391A1) in view of Yang (US 20200035215A1).
As per Claim 7, Lembersky, Chen, and Baik are relied upon for the teachings as discussed above relative to Claim 1.
However, Lembersky does not teach the second input data further represents one or more first values associated with the one or more variables. However, Chen teaches wherein: the second input data further represents one or more first values associated with the one or more variables [0085, 0198]. This would be obvious for the reasons given in the rejection for Claim 1.
However, Lembersky and Chen do not teach that the machine learning model is the second machine learning model. However, Baik teaches the second machine learning model, as discussed in the rejection for Claim 1.
However, Lembersky, Chen, and Baik do not teach the method further comprises generating, using the second machine learning model and based at least on third input data representative of one or more inputs, third output data representative of a second emotional state associated with the second text and one or more second values associated with the one or more variables. However, Yang teaches the method further comprises generating, using the machine learning model and based at least on third input data representative of one or more inputs (previous sentence), third output data representative of a second emotional state (second emotion vector) associated with the second text (second text is the current sentence, first text is the previous sentence) and one or more second values associated with the one or more variables (entire weight applied to “love”) (plurality of sentences in received text data and different second emotion vectors may be set for the current sentence “Where are you?”, although weight “1” may be applied to the emotion item “neutral” when determination is carried out only using the current sentence “Where are you?”, a larger weight can be applied to the emotion item “love” or “happy” for the current sentence “Where are you” when the previous sentence “I miss you” is considered through context analysis, the entire weight is applied to “love”, [0288], emotion vector can be generated through DNN learning for situation explanation, deep learning model is used for learning with respect to emotion expression on the basis of situation explanation, [0250]). Since the combination of Lembersky and Baik teaches the second machine learning model, as discussed in the rejection for Claim 1, this teaching from Yang 1 can be implemented into the second machine learning model of the combination of Lembersky and Baik so that the method further comprises generating, using the second machine learning model and based at least on third input data representative of one or more inputs, third output data representative of a second emotional state associated with the second text and one or more second values associated with the one or more variables.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Lembersky, Chen, and Baik to include generating, using the second machine learning model and based at least on third input data representative of one or more inputs, third output data representative of a second emotional state associated with the second text and one or more second values associated with the one or more variables because Yang suggests that this way, the context of the sentence can be determined from the previous sentence, and the context is used to more accurately determine the emotion of the sentence [0288].
As per Claim 8, Lembersky teaches the second output data is associated with a first portion of the second text; and causing the character to be animated using the first speech [0022]. It would be obvious that after this first speech, there is second speech, and the character is animated using at least the second speech [0022]. Thus, Lembersky teaches causing the character to be animated using at least the second speech [0022].
However, Lembersky does not teach the second output data further represents one or more first values for the one or more variables. However, Chen teaches wherein: the second output data further represents one or more first values for the one or more variables [0085, 0198]. This would be obvious for the reasons given in the rejection for Claim 1.
However, Lembersky and Chen do not teach that the machine learning model is the second machine learning model. However, Baik teaches the second machine learning model, as discussed in the rejection for Claim 1.
However, Lembersky, Chen, and Baik do not teach the method further comprises: generating, using the second machine learning model and based at least on the second input data, third output data associated with a second portion of the second text, the third output data representative of a second emotional state and one or more second values associated with the one or more variables; generating, based at least on the third output data, second audio data representative of second speech associated with the second portion of the second text, the second speech being based at least on the second emotional state and the one or more second values associated with the one or more variables; and causing the character to be animated using at least the second speech. However, Yang teaches a plurality of sentences in received text data and different second emotion vectors are set for the current sentence “Where are you?”. Although weight “1” is applied to the emotion item “neutral” when determination is carried out only using the current sentence “Where are you?”, a larger weight is applied to the emotion item “love” or “happy” for the current sentence “Where are you” when the previous sentence “I miss you” is considered through context analysis. The entire weight is applied to “love” [0288]. It would be obvious that the context analysis was performed on the previous sentence “I miss you” before it was performed on the current sentence “Where are you?” Thus, the context analysis was performed on the previous sentence “I miss you”, and it generated a weight for the emotion item for it. Thus, Yang teaches wherein: the second output data is associated with a first portion of the text (previous sentence “I miss you”) and further represents one or more first values for the one or more variables (weight for the emotion item); the method further comprises: generating, using the machine learning model and based at least on the second input data, third output data associated with a second portion of the text (current sentence “Where are you?”), the third output data representative of a second emotional state (second emotion vector) and one or more second values associated with the one or more variables (entire weight applied to “love”) [0288, 0250]; generating, based at least on the third output data, second audio data representative of second speech associated with the second portion of the text, the second speech being based at least on the second emotional state and the one or more second values associated with the one or more variables ([0288], speech synthesis apparatus may generate second metadata corresponding to the second emotion information corresponding to the sum of the first emotion vector and the second emotion vector, [0279], the generated second metadata may be transmitted to the speech synthesis engine, and the speech synthesis engine may add the second metadata to speech synthesis target text of the received data to perform speech synthesis, [0280], speech synthesis for outputting lively speech, [0006]). Since Lembersky teaches the second text [0022], and the combination of Lembersky and Baik teaches the second machine learning model, as discussed in the rejection for Claim 1, this teaching from Yang can be implemented on the second text of Lembersky so that it generates, using the second machine learning model and based at least on the second input data, third output data associated with a second portion of the second text, the third output data representative of a second emotional state and one or more second values associated with the one or more variables; generating, based at least on the third output data, second audio data representative of second speech associated with the second portion of the second text, the second speech being based at least on the second emotional state and the one or more second values associated with the one or more variables; and causing the character to be animated using at least the second speech. This would be obvious for the reasons given in the rejection for Claim 7.
Claim(s) 23 is/are rejected under 35 U.S.C. 103 as being unpatentable over Lembersky (US 20190095775A1), Chen (US 20240005905A1), and Baik (US 20180047391A1) in view of Wu (US 20180189857A1).
Lembersky, Chen, and Baik are relied upon for the teachings as discussed above relative to Claim 1. Lembersky teaches determining one or more characteristics associated with the character, the one or more characteristics including personal habits of the character, wherein the second input data further represents the one or more characteristics (analyze sentiment of the user, and may correspondingly adjust the response, tone of the AI character, for instance, a child lost in a mall, and based on detecting a worried child, may appear as a calming and concerned cartoon character, if an adult user approaches the system, and the user is correlated to a user that frequents the athletic store in the mall, [0029], machine learning tools 122 may then be used to improve the virtual assistant’s responses based on the user’s past experiences such as shopping and dining habits, the user will then be able to receive personalized greetings and suggestions, [0031], the virtual assistant may then suggest lighter food if the user John set in his preferences to help him to watch after his diet, the virtual assistant could suggest salad place to eat based on John’s information and give encouraged tone to stay on the diet, [0032]).
However, Lembersky, Chen, and Baik do not expressly teach the one or more characteristics including at least one of a profession of the character, a relationship associated with the character, or one or more personal traits of the character. However, Wu teaches the one or more characteristics including at least one of a profession of the character, a relationship associated with the character, or one or more personal traits of the character (user query and profile modeling for providing tailored product recommendations that take into account user traits, examples of user traits include: gender, age, country affiliation, among others, [0018], conversational AI environment 100 for providing product and service recommendations, conversational AI environment 100 includes user profile 108, conversational AI systems may provide conversational responses to user input through chat bots which may be associated with businesses and the like, recommendations may be provided to a user, [0032]).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Lembersky, Chen, and Baik so that the one or more characteristics include at least one of a profession of the character, a relationship associated with the character, or one or more personal traits of the character because Wu suggests that this way, the conversational AI system can provide conversation responses that are even more personalized and tailored for the user [0018, 0032].
Allowable Subject Matter
Claims 10-13, 17, and 18 are allowed.
Claim 21 is objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
The following is a statement of reasons for the indication of allowable subject matter:
The prior art taken singly or in combination do not teach the combination of all the limitations of Claim 10, and in particular, do not teach obtaining first input data representative of one or more learned tokens that cause one or more language models to generate one or more values corresponding to one or more speech characteristics for a character; generate, based at least on first text associated with a query, second input data representative of one or more second text tokens corresponding to one or more words to be output by the character; generate, based at least on the one or more language models processing the one or more learned tokens and the one or more text tokens, output data representative of an emotional state associated with the one or more words and the one or more values corresponding to the one or more speech characteristics associated with the one or more words. Claims 11-13, 17, and 18 depend from Claim 10, and therefore also contain allowable subject matter.
The prior art taken singly or in combination do not teach the combination of all the limitations of Claim 21 and base Claim 1, and in particular, do not teach wherein the second output data represents at least one or more first tokens associated with the second text, one or more second tokens that indicate the emotional state associated with the second text, and one or more third tokens associated with the one or more variables.
One prior art (Beaver (US011822888B2)) teaches receives the query in the form of a string of text. Preprocesses the string by identifying tokens within the string. Tokenizing the string of text. Mapping tokens from the original string of text to vocab items (col. 16, lines 18-35). Tag Express Emotion is used for emotional statement (col. 9, lines 22-23). However, Beaver does not teach wherein the second output data represents at least one or more first tokens associated with the second text, one or more second tokens that indicate the emotional state associated with the second text, and one or more third tokens associated with the one or more variables.
Prior Art of Record
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Beaver (US011822888B2) teaches receives the query in the form of a string of text. Preprocesses the string by identifying tokens within the string. Tokenizing the string of text. Mapping tokens from the original string of text to vocab items (col. 16, lines 18-35).
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JONI HSU whose telephone number is (571)272-7785. The examiner can normally be reached M-F 10am-6:30pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kee Tung can be reached at (571)272-7794. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
JH
/JONI HSU/Primary Examiner, Art Unit 2611