DETAILED ACTION
Notice of Pre-AIA or AIA Status
1. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Continued Examination Under 37 CFR 1.114
2. A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant’s submission filed on 08/25/2026 has been entered.
Response to Arguments/Amendments
3. With respect to 35 U.S.C. § 101, Applicant argues on pages 2-3 of the Remarks that
“Claims 1-6 and 8-21 stand rejected under 35 U.S.C. § 101 because the claimed invention is directed to a judicial exception (i.e., a law of nature, a natural phenomenon, or an abstract idea) without significantly more.
The Examiner asserts that the entire logic described in the claims is entirely equivalent to the conventional thinking behavior of an ordinary person conversing with another person, and that the interface and controls in the claims are merely tools for implementing an abstract thought process. Applicant respectfully disagrees.
Applicant respectfully relies on the Federal Circuit's decision in McRO, Inc. v. Bandai Namco Games Am. Inc., 837 F.3d 1299 (Fed. Cir. 2016) in support of the patent eligibility of the present claims. As McRO established, the key standard for determining whether a claim is directed to an abstract concept lies in whether the claim focuses on a specific means or method that improves the relevant technology, or merely claims the result or effect of the abstract concept itself. The claims of the present application recite specific technical means for improving computer voice interaction functions through the use of a generative model, voice features, and pause duration control, rather than merely pursuing the abstract result of "natural conversation."
As further established in McRO, in determining whether a method is patent-eligible, the court must examine the claim as an ordered combination, without ignoring the specific requirements of the individual steps. By summarizing the claims of the present application as merely the "conventional thinking behavior of an ordinary person conversing with another person," the Examiner has committed the exact "oversimplifying the claims" error criticized by the McRO court. The Examiner has improperly abstracted a specific technical improvement (i.e., a generative model combined with a response speed feature and pause duration control) into the high-level concept of "human conversation," untethered from the actual language of the claims.
Furthermore, there is no evidentiary support demonstrating that, in the field of voice interaction technology, there exists prior art applying the specific technical solution of determining a pause duration based on voice features and controlling the agent's voice output to solve the technical problem of improving voice interaction naturalness.
When viewed as a whole, the claims of the present application constitute a concrete technical improvement to existing computer voice interaction systems. Through specific technical means, the claims achieve natural conversational pacing control in AI voice interactions, rather than merely moving human conversational behavior onto a generic computer. Therefore, the claims are directed to a specific technological solution that improves the functioning of computer voice interaction systems and are patent-eligible under 35 U.S.C. § 101.”
In response, Applicant should note that McRO was found to be eligible under step 2A prong 1 because the claim’s construction incorporated computer-specific rules for animation that improved an existing technological process. While certain steps of McRO could be performed by a human animator, the claimed invention of McRO pertained to computer-specific animation of a talking character. The instant claims contain no such technological as all of the claimed steps are recited at a high level and could be practically performed by a human. Claims recite “a generative model for outputting audio”. Outputting audio is a mental process. The human could communicate (e.g., listen or response) with the user by voice. The generative model is recited at high level of generality. The generative model is used to generally apply the abstract idea without placing any limits on how the generative model functions. Rather, the limitation does not include any details about how the “generating” is accomplished. Accordingly, Applicant’s attempt to draw an analogy between McRO and the instant claims are not persuasive.
Applicant is reminded that lack of novelty under 35 U.S.C. 101 or obviousness under 35 U.S.C. 103 of a claimed invention does not necessarily indicate that additional elements are well-understood, routine, conventional elements. Because they are separate and distinct requirements from eligibility, patentability of the claimed invention under U.S.C. 102 and 103 with respect to the prior art is neither required for, nor a guarantee of, patent eligibility under 35 U.S.C 101. The distinction between eligibility (under 35 U.S.C. 101) and patentability over the art (under 35 U.S.C. 102 and/or 103) is further discussed in MPEP § 2106.05(d). In the present case, the process steps were categorized as mental processes and the remaining component (e.g., a generative model, memory, processor, non-transitory computer-readable storage medium) were the particular left over limitations shown to be merely instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea.
Applicant’s arguments are not persuasive, thus for these reasons, Examiner respectfully disagrees.
Claim Rejections - 35 USC § 101
4. 35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
5. Claims 1-6 and 8-21 are rejected under 35 U.S.C. 101 because the claimed invention is directed to a judicial exception (i.e., a law of nature, a natural phenomenon, or an abstract idea) without significantly more.
All the claims are directed towards the statutory category of a machine, apparatus or process.
Claim 1 recites
1. (Currently Amended) A communication method, comprising:
receiving, from a first object, an input instruction of the first object in an interaction interface between the first object and a second object, wherein the interaction interface is a voice communication interface, and the input instruction is received through a voice input control or by triggering a control in the interaction interface;
determining, based on the input instruction, a target scene mode from one or more scene modes configured for the second object, wherein the second object is an agent, each of the scene modes is configured with a voice feature, the voice feature comprises a response speed feature; and
controlling, during a voice communication process, the second object to perform a voice interaction with the first object based on the voice feature of the target scene mode, comprising:
receiving, through the voice communication interface, a voice from the first object;
generating, by a generative model for outputting audio, a voice of the second object using the voice of the first object and the voice feature during the voice interaction;
determining a waiting duration for the second object based on the response speed feature;
keeping the second object to be silent and waiting for receiving a subsequent voice from the first object through the voice communication interface, in response to an interval between a time the first object last spoke and a current time not exceeding the waiting duration; and
playing the voice of the second object in response to an interval between a time the first object last spoke and the current time exceeding the waiting duration.
The independent Claims 1, 16 recite substantially the same concept but do so in the context of a method and a device.
Claim 20 carries out the communication method of Claim 1.
The limitations recited in the independent claims 1, 16 and 20 as drafted covers a mental process. More specifically, the underlying abstract idea revolved around what happen once a human is talking with a user. The human could listen to the user’s selection for a specific scene (e.g., controlling a device, chatting with a virtual agent), determine which scene mode based on the selected scene. Further, the human starts conversation with the user, the human also could determine a waiting time for the user response based on how fast the user’s talking.
Claim 11 recites
11. (Currently Amended) A communication method, comprising:
receiving, from a first object, an input instruction of the first object in an interaction interface between the first object and a second object, wherein the interaction interface is a voice communication interface, and the input instruction is received through a voice input control or by triggering a control in the interaction interface;
determining, based on the input instruction, a target scene mode from one or more scene modes configured for the second object, wherein the second object is an agent, each of the scene modes is configured with a voice feature; and
controlling, during voice communication process, the second object to perform a voice interaction with the first object based on a voice feature of the target scene mode, wherein
in response to the target scene mode being a third mode, the voice feature of the third mode comprises a pause duration, and the controlling the second object to perform the voice interaction with the first object based on the voice feature of the target scene mode comprises:
receiving, through the voice communication interface, a voice from the first object;
generating, by a generative model for outputting audio, a first voice of the second object using the voice of the first object and the voice feature during the voice interaction, content of the first voice corresponding to the third mode;
playing the first voice generated for the second object;
keeping the second object to be silent and waiting for receiving a subsequent voice from the first object through the voice communication interface, in response to the pause duration being not elapsed after the first voice is played; and
generating, by the generative model for outputting audio, a second voice of the second object using the voice of the first object and the voice feature during the voice interaction, content of the second voice corresponding to the third mode; and
playing, in response to the first object remaining silent and after the pause duration elapses after the first voice is played the second voice generated for the second object.
Claim 21 carries out the communication method of Claim 11.
The limitations recited in the independent claims 11 and 21 as drafted cover a mental process. More specifically, the underlying abstract idea revolved around what happens once a human is talking with a user. The human could listen to the user’s selection for a specific scene (e.g., controlling a device, chatting with a virtual agent), determine which scene mode based on the selected scene. From now on, the human could start conversation with the user in the selected scene. In the selected scene, the human listens to the user, response to the user using the voice of the user and the determined voice feature. The human waits for the subsequent voice from the user, and finally verbally reminds the user after a predetermined waiting time. Besides, the human could keep talking about the same topic/domain with the user from the beginning to the end of the conversation.
The judicial exception is not integrated into a practical application. Particularly, claims recite the additional limitations of a memory, a processor, a non-transitory computer-readable storage medium, a voice input control and a voice communication interface. The additional element(s) or combination of elements such as a memory, a processor, a non-transitory computer-readable storage medium and a voice communication interface in the claim(s) other than the abstract idea per se amount(s) to no more than (i) mere instructions to implement the idea on a computer, and/or (ii) recitation of generic computer structure that serves to perform generic computer functions that are well-understood, routine, and conventional activities previously known to the pertinent industry. Viewed as a whole, these additional claim element(s) do not provide meaningful limitation(s) to transform the abstract idea into a patent eligible application of the abstract idea such that the claim(s) amounts to significantly more than the abstract idea itself. Therefore, the claim(s) are rejected under 35 U.S.C. 101 as being directed to non-statutory subject matter. There is no further improvement to the computing device other than response to the user’s input in the same mode of the user’s voice feature. The mere recitation of a memory and a processor and/or the like is akin of adding the word “apply it” and/or “use it” with a computer in conjunction with the abstract idea. The paragraphs [0006-0008] of the specification disclose “According to the some embodiment of this disclosure, there is provided an electronic device, including: at least one memory; at least one processor coupled to the memory, the processor configured to executed the communication method provided in any embodiment of the present disclosure based on instruction stored in the memory, [0007] According to the some embodiment of this disclosure, there is provided a non-transitory computer-readable storage medium stored there on a computer program that, when executed by a processor, performed the communication method provided by any embodiment of the present disclosure, [0008] According to the some embodiment of this disclosure, there is provided a non-transitory computer program product that, when running on a computer, causes the computer to perform to perform the communication method provided by any embodiment of the present disclosure.”
Claims recite “a generative model for outputting audio”. The generative model is recited at high level of generality. The generative model is used to generally apply the abstract idea without placing any limits on how the generative model functions. Rather, the limitation does not include any details about how the “generating” is accomplished.
As filed in the specification, the computer is listed as a general-purpose computer and are mainly used as an application thereof. Accordingly, these additional elements do not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea.
The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to the integration of the abstract idea into a practical application, the additional element of using a computer is noted as a general computer. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. The claims are not patent eligible.
The dependent claims further do not remedy the issues noted above. More specifically, Claims 2 and 17 recites generating a text based on the input of the first object and an attribute of the second object, generating an playing a voice of the second object based on the voice feature of the target scene mode and the chat text. This reads on the human could write down what he would like to say to the first object on the paper and read the text in the target voice feature. There is no additional limitation presented. Claims 3 and 18 recites generating the voice of the second object corresponding to the chat text based on the sound feature of the target scene mode. This reads on the human could read the text in a specific voice characteristic, language style. There is no additional limitation presented. Claim 4 defines the sound feature. There is no additional limitation presented. Claims 5 and 19 recite adjusting the chat text and generating the voice based on the adjusted text. This reads on the human could change the chat text and read the adjusted text. There is no additional limitation presented. Claim 6 recites determining whether response to the first object. This reads on the human listen to the user and determine if it is question from the user, if so, response to the user’s question. There is no additional limitation presented. Claim 8 recites generating a chat text for the second object based on a voice of the first object, the content feature and an attribute of the second object, generating a voice of the second object based on the chat text, and playing the voice of the second object. This reads on the human could write down what he wants to say on the paper and reads the generated text. There is no additional limitation presented. Claim 9 recites terminating a communication between the first object and the second object in response to an interval between a time the first object last spoke and a current time exceeding a first threshold. This reads on the human could end the conversation with the user because the waiting time exceeds the threshold. There is no additional limitation presented. Claim 10 recites the voice feature of the target scene mode comprises a language recognition instruction, recognizing a voice of the first object based on the language recognition instruction to determine a language used by the first object, generating a voice of the second object using the language used by the first object, playing the voice of the second object. This reads on the human could determine which language the user is speaking, response to the user in the determined language. There is no additional limitation presented. Claim 12 recites generating a voice feature for each scene mode of the scene modes based on an attribute of the scene mode and an attribute of the second object. This reads on the human responses to the user verbally. There is no additional limitation presented. Claim 13 recites displaying the one or more scene modes in the communication interface and determining the target scene mode based on an operation for selecting a scene mode of the first object. This reads on the human could determine which mode is the best mode to talk with the user. There is no additional limitation presented. Claim 14 recites receiving the input of the first object, performing semantic understanding the input, and determining a scene mode. This reads on the human listens to the user, understands what the user just said and determines the target mode (e.g., topic of the conversation) to keep conversation with the user. There is no additional limitation presented. Claim 15 recites who is/are the first object and the second object. There is no additional limitation presented.
For at least the supra provided reasons, claims 1-6 and 8-21 are rejected under 35 U.S.C. 101 as being directed to non-statutory subject matter.
Allowable Subject Matter
6. Claims 1-6 and 8-21 are allowed in view of the prior art of record. The claims stand rejected under 101 Abstract idea, and for the application to pass to allowance this rejection needs to be overcome. Any amendments to overcome the 101 Abstract idea rejection that results in any change in scope require further search and/or consideration to determine its allowability.
The following is a statement of reasons for the indication of allowable subject matter: the prior art(s) taken alone or in combination fail(s) to teach the following element(s) in combination with the other recited elements in the claim(s).
“determining a waiting duration for the second object based on the response speed feature;
keeping the second object to be silent and waiting for receiving a subsequent voice from the first object through the voice communication interface, in response to an interval between a time the first object last spoke and a current time not exceeding the waiting duration; and
playing the voice of the second object in response to an interval between a time the first object last spoke and the current time exceeding the waiting duration.” as recited in Claim 1.
Claim 16 recites similar features as Claim 1. Claim 20 carries out the communication method of Claim 1.
“controlling, during voice communication process, the second object to perform a voice interaction with the first object based on a voice feature of the target scene mode, wherein
in response to the target scene mode being a third mode, the voice feature of the third mode comprises a pause duration, and the controlling the second object to perform the voice interaction with the first object based on the voice feature of the target scene mode comprises:
receiving, through the voice communication interface, a voice from the first object;
generating, by a generative model for outputting audio, a first voice of the second object using the voice of the first object and the voice feature during the voice interaction, content of the first voice corresponding to the third mode;
playing the first voice generated for the second object;
keeping the second object to be silent and waiting for receiving a subsequent voice from the first object through the voice communication interface, in response to the pause duration being not elapsed after the first voice is played; and
generating, by the generative model for outputting audio, a second voice of the second object using the voice of the first object and the voice feature during the voice interaction, content of the second voice corresponding to the third mode; and
playing, in response to the first object remaining silent and after the pause duration elapses after the first voice is played the second voice generated for the second object.” as recited in Claim 11.
Claim 21 carries out the communication method of Claim 11.
The closest prior arts have found as follows.
a. Aher et al. (US 2024/0153483 A1.) In this reference, Aher et al. disclose a method and a system for synthesizing a response to a user’s voice input based on the user’s voice characteristic (Aher et al. [0004] The system determines a response to the voice input, and then generates the synthesized speech response to the voice input for output. The synthesized speech response includes prosodic characteristics that the system determines based on the response itself and the prosodic character of the voice input. The prosodic character of a voice input or response includes pitch, note, duration, prominence, timbre, rate, rhythm, any other suitable metric, and any combination thereof that affect the sound of the utterance, [0009] The system receives a plurality of voice inputs, each associated with at least one respective voice input prosodic metric, and a plurality of responses, each associated with at least one respective response prosodic metric. The plurality of voice inputs and the plurality of responses are associated in a database, and the system may retrieve the voice inputs and responses from the database. To illustrate, in an embodiment, each voice input of the plurality of voice inputs is linked with a respective set of responses of the plurality of responses. The system trains the model based on the plurality of voice inputs, the plurality of responses, the voice input prosodic metrics, and the response prosodic metrics such that the model outputs information used to generate the synthesized speech response to the voice input. For example, the voice input and response prosodic metrics each include pitch, note, duration, prominence, timbre, rate, rhythm, any other suitable metrics, or any combination thereof, [0084] For example, based on user profile information, the voice application may determine that a spoken voice input associated with the user is faster and less verbose when that user is in a hurry, and accordingly, the voice application may provide a faster, less wordy response (e.g., altering both the word content and prosodic character thereof). In this reference, Aher et al. disclose that if the user is in a hurry, the synthesized response may provide a faster, less wordy response. Aher et al. disclose how to response to the user when the user is in a hurry. Aher et al. does not disclose determining a waiting duration based on the user’s speed feature and playing the synthesized response based on the determined waiting duration as recited in Claims 1, 16 and 20. Aher et al. does not disclose controlling the second object based on the voice features of the target scene mode, generating a first voice and a second voice as claimed in Claims 11 and 21. Thus, Aher et al. fails to teach and/or disclose the allowable subject matter noted above.
b. Ghoche et al. (US 2024/0386214 A1.) In this reference, Ghoche et al. disclose a method and a system for generating a natural language workflow policy for a workflow for customer support of emails (Ghoche et al. [0176] FIG. 27 is a variation of the example of FIG. 2A with examples of different modules to automatically generate workflow template answers. In one implementation, the workflow builder 230 includes a workflow template answer engine 2730. In one implementation, the workflow template answer engine 2730 includes a workflow automation recommendations module 2702 to generate recommendations for workflows to be automated. For example, in one implementation this may include identifying topics for which a workflow has not yet been automated. This may include, for example, determining potential cost savings for automating the generation of workflows for one or more topics. In one implementation, a template answer text generation module 2706 generates suggested template text for responding to specific topics. For example, for a topic corresponding to a customer request for a refund, the template text may be generated based on the monitored text answers agents use to respond to the topic.) Ghoche et al. generates a template text may be generated based on the monitored text answers agents use to respond to the topic. Ghoche et al. does not disclose determining a waiting duration based on the user’s speed feature and playing the synthesized response based on the determined waiting duration as recited in Claims 1, 16 and 20. Ghoche et al. does not disclose controlling the second object based on the voice features of the target scene mode, generating a first voice and a second voice as claimed in Claims 11 and 21. Thus, Ghoche et al. fails to teach and/or disclose the allowable subject matter noted above.
c. Kim et al. (US 2021/0383794 A1). In this reference, Kim et al. disclose determining the speech features of the user. The speech features of the user may include the gender of the user, the pitch of the user, the tone of the user, the subject to the user speech, the speed of the user’s speech, and the user’s volume. Kim et al. also considers the waiting time to end the interaction mode or disabling the command recognition function (Kim et al. [0166] In this case, after recognizing the wakeup word, the speech agent changed the command recognition function from an inactive state to an active state and recognizes the command. Then, the speech agent processes the command and disables the command recognition function again if a short speech waiting period is ended after processing the command, [0143] The processor 180 may determine the speech feature of the user by using one or more of the text data or the power spectrum 430 transmitted from the audio processor 181, [0144] The speech feature of the user may include the gender of the user, the pitch of the user, the tone of the user, the subject of the user's speech, the speed of the user's speech, and the user's volume.) Kim et al. does not disclose determining a waiting duration based on the user’s speed feature and playing the synthesized response based on the determined waiting duration as recited in Claims 1, 16 and 20. Kim et al. does not disclose controlling the second object based on the voice features of the target scene mode, generating a first voice and a second voice as claimed in Claims 11 and 21. Thus, Kim et al. fails to teach and/or disclose the allowable subject matter noted above.
Conclusion
7. The prior art made of record and not relied upon is considered pertinent to applicant’s disclosure. See PTO-892.
a. Muntasir et al. (US 2025/0118298 A1.) In this reference, Muntasir et al. discloses a method and a system for performing speech synthesis on the generated responses further comprises storing TTS models and TTS parameters such as, speaking rate, pitch, volume, intonation, and preferred responses corresponding to the user interaction session.
b. Lezzoum et al. (US 12,087,284 B1). In this reference, Lezzoum et al. discloses a method and a system for selecting a gender, an accent, and a language for the synthesized speech.
c. Mennicken et al. (US 2021/0104220 A1.) In this reference, Mennicken et al. discloses a method and a system for generating synthesized speech as a voice output.
8. Any inquiry concerning this communication or earlier communications from the examiner should be directed to THUYKHANH LE whose telephone number is (571)272-6429. The examiner can normally be reached Mon-Fri: 9am-5pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Andrew C. Flanders can be reached on 571-272-7516. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/THUYKHANH LE/Primary Examiner, Art Unit 2655