Prosecution Insights
Last updated: August 17, 2026
Application No. 18/981,180

Voice Interaction Method, System, Terminal Device and Medium

Non-Final OA §103§112§DP
Filed
Dec 13, 2024
Priority
Aug 29, 2019 — CN 201910808807.0 +2 more
Examiner
YEN, ERIC L
Art Unit
Tech Center
Assignee
BOE Technology Group Co., Ltd.
OA Round
1 (Non-Final)
85%
Grant Probability
Favorable
1-2
OA Rounds
1y 0m
Est. Remaining
97%
With Interview

Examiner Intelligence

Grants 85% — above average
85%
Career Allowance Rate
661 granted / 777 resolved
+25.1% vs TC avg
Moderate +12% lift
Without
With
+11.7%
Interview Lift
resolved cases with interview
Typical timeline
2y 9m
Avg Prosecution
18 currently pending
Career history
783
Total Applications
across all art units

Statute-Specific Performance

§101
17.5%
-22.5% vs TC avg
§103
32.8%
-7.2% vs TC avg
§102
3.9%
-36.1% vs TC avg
§112
35.1%
-4.9% vs TC avg
Black line = Tech Center average estimate • Based on career data from 777 resolved cases

Office Action

§103 §112 §DP
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 3, 5, 16, and 17, are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. As per Claim 3 (and similarly claims 5 and 16-17): “the sample sentence” in lines 2-3 of claim 3, and in the last line of claim 3, is ambiguous (Applicant fairly clearly meant to refer to “the preconfigured sample sentence” [Applicant references “the preconfigured sample sentence” in line 5 of claim 3] but as claimed “the sample sentence” can also refer to “a sample sentence having the same or similar semantics as the input sentence” in lines 6-7 of claim 1, and can also refer to any one of “cached sample sentences” in lines 3-4 of claim 1) “the response content of the cached sample sentences” in lines 4-5 of claim 3 lacks antecedent basis (claim 1 only recites “cached response content corresponding to the sample sentence having the same or similar semantics as the input sentence”, and does not explicitly recite where any other sample sentences of the cached sample sentences also have corresponding response content) Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1, 11, 14, 19, and 20, is/are rejected under 35 U.S.C. 103 as being unpatentable over Guo et al. (US 2019/0354630), hereafter Guo, in view of Petrov et al. (US 2015/0161996), hereafter Petrov, Ducatel et al. (US 2012/0303358), hereafter Ducatel, and Tseretopoulos et al. (US 2019/0103102), hereafter Tseretopoulos. As per Claim 1, 14, and 20, Guo suggests (along with its terminal device and medium equivalents) A method performed by a terminal device, comprising:… matching between… input sentence and cached sample sentence… to determine whether there is a sample sentence having the same or similar semantics as the input sentence among the cached sample sentences; in response to determining that there is a sample sentence having the same or similar semantics as the input sentence among the cached sample sentences, acquiring cached response content corresponding to the sample sentence having the same or similar semantics as the input sentence as first response content of the input sentence; in response to determining that there is no sample sentence having the same or similar semantics as the input sentence among the cached sample sentences, sending at least one of the input sentence or the collected voice signals to a server, and receiving transmitted response content of the at least one of the input sentence or the collected voice signals from the server as the first response content of the input sentence… responding to the input sentence according to the first response content of the input sentence… (Figures 3 and 5; paragraphs 21, 41, 43, 59-65, 67, 73-78; [all paragraphs and Figures are cited for each limitation with “key” paragraphs and Figures pertaining to each limitation identified below, i.e. all other paragraphs and Figures not specifically referenced for any particular limitation are eligible to provide context and additional support] “A method performed by a terminal device,”: paragraphs 41, 62, 67, 73-78, Figures 3 and 5; a Q&A device [“terminal device”] performs a method that includes receiving a question from a user and providing an answer to the question to the user, where the terminal device can be a computer. Paragraphs 73-78 describe where functions/acts “in the flowchart and/or block diagram block or blocks” [paragraphs 77-78] are implemented by a processor of a computer executing instructions, where instructions can be stored in a computer readable storage medium. Figure 5 and paragraphs 62 and 67 depicts a flowchart including, among other things, receiving a query from a user and providing an answer to a query to the user [i.e. functions of the Q&A device of paragraph 41]. These portions thus suggest an embodiment of the Q&A device [“terminal device”] which is a computer which includes a memory storing instructions and a processor which executes the instructions in the memory to perform the functions of the Q&A device [and thus suggests the terminal device and non-transitory computer-readable storage medium of claims 9 and 20] “comprising:… matching between… input sentence and cached sample sentence… to determine whether there is a sample sentence having the same or similar semantics as the input sentence among the cached sample sentences; in response to determining that there is a sample sentence having the same or similar semantics as the input sentence among the cached sample sentences, acquiring cached response content corresponding to the sample sentence having the same or similar semantics as the input sentence as first response content of the input sentence;”: Figures 3 and 5; paragraphs 21, 41, 59-63; determining whether there is a match to a question in a query from a user among one of a plurality of frequently asked questions [in a local cache of the Q&A device] that each have an associated answer/”response content” [where each frequently asked question and its associated answer are a question-answer pair in the local cache], and in response to determining that there is a matching frequently asked question in the local cache, providing an answer in the local cache [“cached response content”] associated with the matching frequently asked question to the user [as an answer to the user’s question, i.e. “as response content of the input sentence”]. Paragraph 21 describes determining whether a question from a user matches a question in a local cache and “If the question is found in the local” [at least suggested to be where the user question matches a question in the local cache] “an associated answer” [at least suggested to be an answer associated with the local cache question that matches the user question] “is provided to the user”, and paragraph 62 describes determining whether a query included a question from a user is matched to a question/answer pair in a local cache and where the query is answered with “the matching answer” [where paragraphs 21 and 62 suggest where the user’s question is matched to a question in a question/answer pair in the local cache and then the answer of the question/answer pair is provided to the user as an answer]. Figure 3 depicts where the question/answer cache is in the Q&A device. Paragraphs 60-61 describes examples of a user question which can be a sentence [such that a sentence user question can be interpreted as an “input sentence”]. Paragraphs 21 and 59 describe where the questions in the local cache are “frequently asked questions”, and matching the user question [which can be a sentence] to one of the frequently asked questions [paragraphs 21 and 62] suggests where the frequently asked questions are also “cached… sentences”, because the frequently asked questions are in the local cache, because questions can be sentences [paragraphs 60-61], and because “match”, given its literal meaning, means that the user question and a local cache frequently asked question are the same [i.e. the same words and thus the same sentences], and in order for the matching frequently asked question’s associated answer to be appropriate for the user question, the matching frequently asked question logically needs to at least be semantically similar to the user question. The suggested sentence frequently asked questions can be interpreted as “cached sample sentences” in the sense that they are “sentences” in the local “cache” that are “samples” of frequently asked questions that may be asked by the user. Paragraphs 21, 62-63 also suggests an embodiment where the question in the query from the user is compared/”matched” to all of the frequently asked questions in the local cache, because, in order to determine that “a match is not found between the question including in the query” [paragraph 63, at least suggested to be a contrasting “If a match is” not “found between the question included in the query and a question/answer pair in cache 306” condition relative to the “If a match is found between the question included in the query and a question/answer pair in cache 306” condition of paragraph 62], the Q&A device logically needs to compare the question in the query to every question in the local cache and determine that none of them match the question in the query [see paragraph 21 which describes where it is determined “whether the question matches a question in the local cache”]. Also of relevance is that paragraph 59 describes where a user can interact with the Q&A device using voice input or text input and where a question is “in the input” and “receiv[ing] an input, such as a voice or text input” [suggesting that the user’s question can be in text format]. “in response to determining that there is no sample sentence having the same or similar semantics as the input sentence among the cached sample sentences, sending at least one of the input sentence or the collected voice signals to a server, and receiving transmitted response content of the at least one of the input sentence or the collected voice signals from the server as the first response content of the input sentence… “: Figures 3 and 5; paragraphs 21, 41, 43, 59-65, and 67; if a match is not found between the user question and any frequently asked question in the local cache, the Q&A device sends the user’s question to a “cloud service” at a server and receives an answer from the “cloud service” at the server to be provided to the user [where the answer is an answer to the user’s question, i.e. “response content of the input sentence”]. Paragraph 21 describes where “Otherwise” [relative to “If the question is found in the local”, which logically means when the question is not found in the local [at least suggested to be because there is no matching question in the local cache], “the question” [i.e. the user question] is sent to a cloud service. Figure 5 and paragraphs 63-65 and 67 similarly describe sending a request to a cloud service requesting an answer to the question [paragraph 65] if there is no match [paragraph 63, where Figure 5 depicts where no match at element 508 leads, eventually, to elements 518, 526 and 528]. Paragraph 43, 59, and 67, and Figure 3, describe where Q&A cloud services are implemented by a Q&A system in a server [suggesting that sending the question to a cloud service sends the question to the server and where the answer from the cloud service comes from the server] and where the answer provided to the user is “received from the cloud server[s]”. “and responding to the input sentence according to the first response content of the input sentence…”: Figures 3 and 5; paragraphs 21, 41, 43, 59-65, and 67; providing the answer in the local cache [as an answer to the user’s question, i.e. “as response content of the input sentence”] corresponding to the frequently asked question that matches the user question [if there is a matching frequently asked question in the local cache, see paragraph 62] and providing the answer from the server [as an answer to the user’s question, i.e. “as response content of the input sentence”, if there is no matching frequently asked question in the local cache, see paragraphs 63-65 and 67]. Also of relevance [for claims 7 and 15] is that paragraph 21 at least suggests where a response provided by a device [provided to a user in response to a question] is provided via audio output.) Guo does not, but Petrov suggests A method performed by a terminal device, comprising: performing voice recognition on collected voice signals to acquire an input sentence;… matching between the input sentence and cached sample sentence… to determine whether there is a sample sentence having the same or similar semantics as the input sentence among the cached sample sentences; in response to determining that there is a sample sentence having the same or similar semantics as the input sentence among the cached sample sentences, acquiring cached response content corresponding to the sample sentence having the same or similar semantics as the input sentence as first response content of the input sentence; in response to determining that there is no sample sentence having the same or similar semantics as the input sentence among the cached sample sentences, sending at least one of the input sentence or the collected voice signals to a server, and receiving transmitted response content of the at least one of the input sentence or the collected voice signals from the server as the first response content of the input sentence,…and responding to the input sentence according to the first response content of the input sentence… (Figure 1; paragraphs 35-36; Guo [in paragraph 59] describes where a user can interact with the Q&A device via voice input or text input and where a question is “in the input” where “an input” can be, for example, “a voice or text input” [suggesting where the user question can either be in text or voice form]. Petrov [Figure 1 and paragraphs 35-36] describes where a text input representing a question [that may be input by a user] can be obtained from speech-to-text of a speech input [where speech to text is a common form of speech/”voice” recognition], where a question which may be a string of characters [at least suggested to be text based on antecedent basis of “the question” in line 4 of paragraph 36] may be transmitted to a computing device 200 [suggesting that speech-to-text processing is performed at the computing device 104, and where the text question generated by the speech-to-text processing is sent to the computing device 200] and where an answer to the question is determined based on analyzing the text representing the question. Petrov thus suggests “A method performed by a terminal device, comprising: performing voice recognition on collected voice signals to acquire an input sentence;… matching between the input sentence and cached sample sentence… to determine whether there is a sample sentence having the same or similar semantics as the input sentence among the cached sample sentences; in response to determining that there is a sample sentence having the same or similar semantics as the input sentence among the cached sample sentences, acquiring cached response content corresponding to the sample sentence having the same or similar semantics as the input sentence as first response content of the input sentence; in response to determining that there is no sample sentence having the same or similar semantics as the input sentence among the cached sample sentences, sending at least one of the input sentence or the collected voice signals to a server, and receiving transmitted response content of the at least one of the input sentence or the collected voice signals from the server as the first response content of the input sentence,…and responding to the input sentence according to the first response content of the input sentence…”: where the user’s question [which can be in text form and which is compared/matched to frequently asked questions in the local cache or is sent to a cloud server to obtain, from the cloud server, an answer to the question] in Guo is more specifically a text question [“input sentence”] generated by performing speech-to-text [“voice recognition”] on a question spoken by a user at the Q&A device [i.e. by the “terminal device”]) Therefore, it would have been obvious to one of ordinary skill in the art at the time of effective filing to perform a simple substitution of one type of user-input text question with another because the prior art teaches the claimed invention except for the substitution of a user-input text question which is not necessarily generated by client-side speech-to-text with a user-input text question which is. Petrov suggests that a user-input text question which is generated by client-side speech-to-text was known in the art. One of ordinary skill in the art could have substituted one type of user-input text question with another to obtain the predictable results of a Q&A user computer device which receives a input question from a user, determines if there is a matching frequently asked question that matches the input question in a local cache containing frequently-asked-question/answer pairs, and if there is a matching frequently asked question in the local cache, provides an answer associated with the matching frequently asked question to the user, and if there is no matching frequently asked question in the local cache, sends the input question to a cloud server, receives an answer to the input question from the cloud server, and provides the answer received from the cloud server to the user (as per Guo) where the input question is an input text question that is generated by performing speech-to-text processing on a question spoken by the user at the Q&A user computer device (as per Petrov). Guo, in view of Petrov, do not, but Ducatel suggests performing semantic matching between the input sentence and cached sample sentences to determine whether there is a sample sentence having the same or similar semantics as the input sentence among the cached sample sentences; (paragraphs 10, 17, 40-44 and 56; Ducatel describes [in paragraphs 10 and 56] where a submitted question/query can be analyzed and compared with queries held in a document retrieval system to determine if the submitted question has sufficient semantic similarity to one [or more] of the queries held in the document retrieval system, and if so, documents which are pertinent to the query [or queries] are returned to the user who submitted the question [similar to how the associated answer for a matching frequently asked question in Guo is provided to the user]. “If the submitted question has sufficient semantic similarity to one [or more] of the queries held within the document retrieval system” [paragraph 56] suggests where the submitted question is compared with every/each one of the multiple “number of queries” that are held within the document retrieval system, because, in order to accurately assess which of the “number of queries” held in the document retrieval system do or do not match the submitted question [thereby accurately determining which one or more of “the” queries held in the document retrieval system match the submitted question], the submitted question logically needs to be compared with every/each one of the multiple queries held within the document retrieval system. Ducatel [paragraphs 17 and 40-44] further describes where semantic similarity between text passages can be numeric quantities [a PSSM from 0 to 1], and where text passages can be a complete sentence [paragraph 17], and where a numerical indication of semantic similarity [the PSSM] can be calculated [paragraph 40]. Ducatel thus suggests “performing semantic matching between the input sentence and cached sample sentences to determine whether there is a sample sentence having the same or similar semantics as the input sentence among the cached sample sentences;”: where the matching/comparison of the user input question [a speech-to-text-generated text input question in the Guo/Petrov combination] to frequently asked question[s] in the local cache is more specifically a semantic similarity matching/comparison [i.e. “semantically matching”] that matches/compares, for each/every one of the frequently asked questions in the local cache [“cached sample sentences”], the user input question to the frequently asked question in the local cache by determining the semantic similarity between the user input question and the frequently asked question, and by determining whether the semantic similarity between the user input question and the frequently asked question is sufficient for the user input question and the frequently asked question to be considered a match) Therefore, it would have been obvious to one of ordinary skill in the art at the time of effective filing to perform a simple substitution of one type of user-question-to-stored-question comparison with another because the prior art teaches the claimed invention except for the substitution of user-question-to-stored-question comparison which does not necessarily determine, for each of a plurality of stored questions, how semantically similar the user question and the stored question are with user-question-to-stored-question comparison which does. Ducatel teaches that user-question-to-stored-question comparison which determines, for each of a plurality of stored questions, how semantically similar the user question and the stored question are was known in the art. One of ordinary skill in the art could have substituted one type of user-question-to-stored-question comparison with another to obtain the predictable results of a Q&A user computer device which receives a input question from a user, determines if there is a matching frequently asked question that matches the input question in a local cache containing frequently-asked-question/answer pairs, and if there is a matching frequently asked question in the local cache, provides an answer associated with the matching frequently asked question to the user, and if there is no matching frequently asked question in the local cache, sends the input question to a cloud server, receives an answer to the input question from the cloud server, and provides the answer received from the cloud server to the user (as per Guo) where the input question is an input text question that is generated by performing speech-to-text processing on a question spoken by the user at the Q&A user computer device (as per Petrov) where the determining of if there is a matching frequently asked question that matches the input question in a local cache containing frequently-asked-question/answer pairs determines, for each of the frequently asked questions in the local cache, if there is sufficient semantic similarity between the input question and the frequently asked question in the local cache (as suggested by Ducatel). Guo, in view of Petrov and Ducatel, do not, but Tseretopoulos suggests wherein the transmitted response content of the at least one of the input sentence or the collected voice signals from the server is acquired by the server through semantic understanding of the at least one of the input sentence or the collected voice signals according to a knowledge base stored on the server (Figure 1; paragraphs 38, 40, 49, 52, 76-77; Tseretopoulos describes where a “server-side” system [conversational analysis system 102 in Figure 1] receives an input from a client device input which can be text or audio is analyzed to determine an intent [e.g., a particular question to which a response may be generated] and generating a response based on the determined intent [paragraph 40], where a result of intent determination can be a set of lexical semantics of a received input which can be provided to a response analysis engine [paragraph 49], where a determined intent of the received input can be provided to a response analysis engine which determines an ideal response, where the response analysis engine determines a responsive answer to a question when the determined intent is a question, where the answer is provided to an NLG engine for generation of a sentence [at least suggested to be an answer sentence], and generates a response to the determined intent [paragraph 52] and where a conversational input which is analyzed to determine an intent can be text or audio [paragraphs 76-77] and where intent can be used to determine a meaning of a particular request and may be used to determine a particular response to be generated [paragraph 77]. Paragraphs 49 and 52 suggest where the determined intent includes semantic information of the received input because paragraph 52 describes where the determined intent is provided to a response analysis engine and paragraph 49 describes where “the result of the intent determination can be a set of lexical semantics of the received input” [semantic information] “can then be provided to the response analysis engine” [i.e. both the set of lexical semantics which are the result of the intent determination and the determined intent are provided to the same response analysis engine, which suggests that the determined intent is “the result of the intent determination which includes the set of lexical semantics]. Paragraph 77 further suggests where the determined intent [which is suggested to be semantic information as per paragraphs 49 and 52] reflects/indicates the meaning of the input because it states that “The intent associated with the input… can be used to determine meaning of a particular request associated with the conversational input” [the meaning of the input logically can’t be determined from the intent if the intent contained no information about the meaning of the input]. The determined intent is thus suggested to be a form of “understanding” of the “semantics” of the input which is determined. Paragraph 46 describes where operations of the conversational analysis system are performed by a processor included in the conversational analysis system executing instructions, and paragraph 47 describes where software instructions can be stored in a non-transitory medium, and paragraph 24 describes where a system comprising “at least one process” [at least suggested to be “at least one processor”] “and a memory operatively coupled to the at least one processor where the memory stores instructions that when executed cause the at least one processor to perform the operations”. These portions suggest where the “server-side” conversational analysis system includes/”stores” a “base” of “knowledge” containing all data for all functions/operations, and where the functions of the conversational analysis system [including the intent determination] are performed “according to” the conversational analysis system’s “knowledge base” [each function is performed according to respective software-specific software instructions stored in the “knowledge base”]. Additionally/alternatively, all data that exists in the conversational analysis system can be interpreted as a “base” of “knowledge” that is “stored” in the conversational analysis system. Additionally/alternatively, the specific portion of the memory storing software instructions used particularly to perform the intent determination can be interpreted as the claimed “knowledge base” Tseretopoulos suggests “wherein the transmitted response content of the at least one of the input sentence or the collected voice signals from the server is acquired by the server through semantic understanding of the at least one of the input sentence or the collected voice signals according to a knowledge base stored on the server”: where, in the Guo/Petrov/Ducatel combination, the cloud server which receives the speech-to-text-generated text user’s question [“the input sentence”] which is not matched to [i.e. not determined to be sufficiently semantically similar to] a frequently asked question in the local cache and provides an answer to the user’s question to the Q&A device more specifically determines the intent of the speech-to-text-generated text user’s question [i.e. the intent of “the input sentence”, where the intent includes semantic information indicating the meaning of the particular question that the user has asked], and determines/generates/”acquires”, based on the determined intent, the answer received by the Q&A device from the cloud server [i.e. “the response content of the input sentence” if there is no matching frequently asked question in the local cache]. The determining of the intent of the speech-to-text-generated text user’s question is suggested to be a form of “semantic understanding of the input sentence” in the sense that the determined intent [suggested to be semantic information that indicates the meaning of input text, as discussed above] reflects what the server “understood” the meaning/“semantics” of the speech-to-text-generated user’s question to be [such that the process that determines the intent is suggested to be a process the semantically understands the speech-to-text-generated user’s question]. The determining of the intent of the speech-to-text-generated text user’s question is also suggested to be “according to a knowledge base stored on the server” [as discussed above, performed according to intent determination software instructions stored in a particular portion of the server’s memory, where any of the particular potion, the entire memory storing all data for all functions, and all data in the cloud server, can be interpreted as the claimed “knowledge base stored on the server”]) Therefore, it would have been obvious to one of ordinary skill in the art at the time of effective filing to perform a simple substitution of one type of server-side question answering with another because the prior art teaches the claimed invention except for the substitution of server-side question answering which does not necessarily determine an answer to a question based on semantic understanding of the question with server-side question answering which does. Tseretopoulos suggests that server-side question answering which determines an answer to a question based on semantic understanding of the question was known in the art. One of ordinary skill in the art could have substituted one type of server-side question answering with another to obtain the predictable results of a Q&A user computer device which receives a input question from a user, determines if there is a matching frequently asked question that matches the input question in a local cache containing frequently-asked-question/answer pairs, and if there is a matching frequently asked question in the local cache, provides an answer associated with the matching frequently asked question to the user, and if there is no matching frequently asked question in the local cache, sends the input question to a cloud server, receives an answer to the input question from the cloud server, and provides the answer received from the cloud server to the user (as per Guo) where the input question is an input text question that is generated by performing speech-to-text processing on a question spoken by the user at the Q&A user computer device (as per Petrov) where the determining of if there is a matching frequently asked question that matches the input question in a local cache containing frequently-asked-question/answer pairs determines, for each of the frequently asked questions in the local cache, if there is sufficient semantic similarity between the input question and the frequently asked question in the local cache (as suggested by Ducatel) where the cloud server determines intent of the input question and determines the provided answer received from the cloud server based on the determined intent (as per Tseretopoulos). As per Claim 19, Guo suggests A voice interaction system, comprising: a terminal device configured to:… matching between… input sentence and cached sample sentence… to determine whether there is a sample sentence having the same or similar semantics as the input sentence among the cached sample sentences, in response to determining that there is a sample sentence having the same or similar semantics as the input sentence among the cached sample sentences, acquire cached response content corresponding to the sample sentence having the same or similar semantics as the input sentence as first response content of the input sentence, in response to determining that there is no sample sentence having the same or similar semantics as the input sentence among the cached sample sentences, send at least one of the input -37-sentence or the collected voice signals to a server, and receive transmitted response content of the at least one of the input sentence or the collected voice signals from the server as the first response content of the input sentence, respond to the input sentence according to the first response content of the input sentence…; and the server, wherein the server is configured to: receive the at least one of the input sentence or the collected voice signals from the terminal device,…and send the transmitted response content of the at least one of the input sentence or the collected voice signals to the terminal device (Figures 3 and 5; paragraphs 21, 41, 43, 59-65, 67, 73-78; [all paragraphs and Figures are cited for each limitation with “key” paragraphs and Figures pertaining to each limitation identified below, i.e. all other paragraphs and Figures not specifically referenced for any particular limitation are eligible to provide context and additional support] “A voice interaction system comprising: a terminal device, configured to:”: paragraphs 21, 41, 43, 59-65, 67, 73-78, Figures 3 and 5; a Q&A device [“terminal device”] is configured to perform a method that includes receiving a question from a user and providing an answer to the question to the user, where the terminal device can be a computer, and where the terminal device, together with a cloud server [that, when a user’s question does not match a frequently asked question in a local cache of the Q&A device, receives a user’s question from the Q&A device and provides an answer to the user’s question to the Q&A device] are collectively a “system”, where paragraph 59 describes where the user interface of the Q&A device “allow[s] user… to interact with Q&A device… by a voice input” [such that the “system” can be interpreted as a “voice interaction system”]. Paragraphs 73-78 describe where functions/acts “in the flowchart and/or block diagram block or blocks” [paragraphs 77-78] are implemented by a processor of a computer executing instructions, where instructions can be stored in a computer readable storage medium. Figure 5 and paragraphs 62 and 67 depicts a flowchart including, among other things, receiving a query from a user and providing an answer to a query to the user [i.e. functions of the Q&A device of paragraph 41]. These portions thus suggest an embodiment of the Q&A device [“terminal device”] which is a computer which includes a memory storing instructions and a processor which executes the instructions in the memory to perform the functions of the Q&A device [and thus suggests the “terminal device” which is “configured to” perform the functions of the Q&A device]. Paragraph 43, 59, and 67, and Figure 3, describe where Q&A cloud services are implemented by a Q&A system in a server [suggesting that sending the question to a cloud service sends the question to the server and where the answer from the cloud service comes from the server] and where the answer provided to the user is “received from the cloud server[s]”. “… matching between… input sentence and cached sample sentence… to determine whether there is a sample sentence having the same or similar semantics as the input sentence among the cached sample sentences; in response to determining that there is a sample sentence having the same or similar semantics as the input sentence among the cached sample sentences, acquire cached response content corresponding to the sample sentence having the same or similar semantics as the input sentence as first response content of the input sentence;”: Figures 3 and 5; paragraphs 21, 41, 59-63; determining whether there is a match to a question in a query from a user among one of a plurality of frequently asked questions [in a local cache of the Q&A device] that each have an associated answer/”response content” [where each frequently asked question and its associated answer are a question-answer pair in the local cache], and in response to determining that there is a matching frequently asked question in the local cache, providing an answer in the local cache [“cached response content”] associated with the matching frequently asked question to the user [as an answer to the user’s question, i.e. “as response content of the input sentence”]. Paragraph 21 describes determining whether a question from a user matches a question in a local cache and “If the question is found in the local” [at least suggested to be where the user question matches a question in the local cache] “an associated answer” [at least suggested to be an answer associated with the local cache question that matches the user question] “is provided to the user”, and paragraph 62 describes determining whether a query included a question from a user is matched to a question/answer pair in a local cache and where the query is answered with “the matching answer” [where paragraphs 21 and 62 suggest where the user’s question is matched to a question in a question/answer pair in the local cache and then the answer of the question/answer pair is provided to the user as an answer]. Figure 3 depicts where the question/answer cache is in the Q&A device. Paragraphs 60-61 describes examples of a user question which can be a sentence [such that a sentence user question can be interpreted as an “input sentence”]. Paragraphs 21 and 59 describe where the questions in the local cache are “frequently asked questions”, and matching the user question [which can be a sentence] to one of the frequently asked questions [paragraphs 21 and 62] suggests where the frequently asked questions are also “cached… sentences”, because the frequently asked questions are in the local cache, because questions can be sentences [paragraphs 60-61], and because “match”, given its literal meaning, means that the user question and a local cache frequently asked question are the same [i.e. the same words and thus the same sentences], and in order for the matching frequently asked question’s associated answer to be appropriate for the user question, the matching frequently asked question logically needs to at least be semantically similar to the user question. The suggested sentence frequently asked questions can be interpreted as “cached sample sentences” in the sense that they are “sentences” in the local “cache” that are “samples” of frequently asked questions that may be asked by the user. Paragraphs 21, 62-63 also suggests an embodiment where the question in the query from the user is compared/”matched” to all of the frequently asked questions in the local cache, because, in order to determine that “a match is not found between the question including in the query” [paragraph 63, at least suggested to be a contrasting “If a match is” not “found between the question included in the query and a question/answer pair in cache 306” condition relative to the “If a match is found between the question included in the query and a question/answer pair in cache 306” condition of paragraph 62], the Q&A device logically needs to compare the question in the query to every question in the local cache and determine that none of them match the question in the query [see paragraph 21 which describes where it is determined “whether the question matches a question in the local cache”]. Also of relevance is that paragraph 59 describes where a user can interact with the Q&A device using voice input or text input and where a question is “in the input” and “receiv[ing] an input, such as a voice or text input” [suggesting that the user’s question can be in text format]. “in response to determining that there is no sample sentence having the same or similar semantics as the input sentence among the cached sample sentences, send at least one of the input -37-sentence or the collected voice signals to a server, and receive transmitted response content of the at least one of the input sentence or the collected voice signals from the server as the first response content of the input sentence”: Figures 3 and 5; paragraphs 21, 41, 43, 59-65, and 67; if a match is not found between the user question and any frequently asked question in the local cache, the Q&A device sends the user’s question to a “cloud service” at a server and receives an answer from the “cloud service” at the server to be provided to the user [where the answer is an answer to the user’s question, i.e. “response content of the input sentence”]. Paragraph 21 describes where “Otherwise” [relative to “If the question is found in the local”, which logically means when the question is not found in the local [at least suggested to be because there is no matching question in the local cache], “the question” [i.e. the user question] is sent to a cloud service. Figure 5 and paragraphs 63-65 and 67 similarly describe sending a request to a cloud service requesting an answer to the question [paragraph 65] if there is no match [paragraph 63, where Figure 5 depicts where no match at element 508 leads, eventually, to elements 518, 526 and 528]. Paragraph 43, 59, and 67, and Figure 3, describe where Q&A cloud services are implemented by a Q&A system in a server [suggesting that sending the question to a cloud service sends the question to the server and where the answer from the cloud service comes from the server] and where the answer provided to the user is “received from the cloud server[s]”. “respond to the input sentence according to the first response content of the input sentence…”: Figures 3 and 5; paragraphs 21, 41, 43, 59-65, and 67; providing the answer in the local cache [as an answer to the user’s question, i.e. “as response content of the input sentence”] corresponding to the frequently asked question that matches the user question [if there is a matching frequently asked question in the local cache, see paragraph 62] and providing the answer from the server [as an answer to the user’s question, i.e. “as response content of the input sentence”, if there is no matching frequently asked question in the local cache, see paragraphs 63-65 and 67] “and the server, wherein the server is configured to: receive the at least one of the input sentence or the collected voice signals from the terminal device,…and send the transmitted response content of the at least one of the input sentence or the collected voice signals to the terminal device”: Figures 3 and 5; paragraphs 21, 41, 43, 59-65, and 67; if a match is not found between the user question and any frequently asked question in the local cache, the Q&A device sends the user’s question to a “cloud service” at a server [such that the server is configured to “receive the input sentence from the terminal device”] and receives [as a result of the server sending “the response content of the input sentence to the terminal device”] an answer from the “cloud service” at the server to be provided to the user [where the answer is an answer to the user’s question, i.e. “response content of the input sentence”]. Paragraph 21 describes where “Otherwise” [relative to “If the question is found in the local”, which logically means when the question is not found in the local [at least suggested to be because there is no matching question in the local cache], “the question” [i.e. the user question] is sent to a cloud service. Figure 5 and paragraphs 63-65 and 67 similarly describe sending a request to a cloud service requesting an answer to the question [paragraph 65] if there is no match [paragraph 63, where Figure 5 depicts where no match at element 508 leads, eventually, to elements 518, 526 and 528]. Paragraph 43, 59, and 67, and Figure 3, describe where Q&A cloud services are implemented by a Q&A system in a server [suggesting that sending the question to a cloud service sends the question to the server and where the answer from the cloud service comes from the server] and where the answer provided to the user is “received from the cloud server[s]”.) Guo does not, but Petrov suggests A voice interaction system, comprising: a terminal device configured to: perform voice recognition on collected voice signals to acquire an input sentence,… matching between the input sentence and cached sample sentence… to determine whether there is a sample sentence having the same or similar semantics as the input sentence among the cached sample sentences, in response to determining that there is a sample sentence having the same or similar semantics as the input sentence among the cached sample sentences, acquire cached response content corresponding to the sample sentence having the same or similar semantics as the input sentence as first response content of the input sentence, in response to determining that there is no sample sentence having the same or similar semantics as the input sentence among the cached sample sentences, send at least one of the input -37-sentence or the collected voice signals to a server, and receive transmitted response content of the at least one of the input sentence or the collected voice signals from the server as the first response content of the input sentence, respond to the input sentence according to the response content of the input sentence,…; and the server, wherein the server is configured to: receive the at least one of the input sentence or the collected voice signals from the terminal device,… and send the transmitted response content of the at least one of the input sentence or the collected voice signals to the terminal device (Figure 1; paragraphs 35-36; Guo [in paragraph 59] describes where a user can interact with the Q&A device via voice input or text input and where a question is “in the input” where “an input” can be, for example, “a voice or text input” [suggesting where the user question can either be in text or voice form]. Petrov [Figure 1 and paragraphs 35-36] describes where a text input representing a question [that may be input by a user] can be obtained from speech-to-text of a speech input [where speech to text is a common form of speech/”voice” recognition], where a question which may be a string of characters [at least suggested to be text based on antecedent basis of “the question” in line 4 of paragraph 36] may be transmitted to a computing device 200 [suggesting that speech-to-text processing is performed at the computing device 104, and where the text question generated by the speech-to-text processing is sent to the computing device 200] and where an answer to the question is determined based on analyzing the text representing the question. Petrov thus suggests “A voice interaction system, comprising: a terminal device configured to: perform voice recognition on collected voice signals to acquire an input sentence,… matching between the input sentence and cached sample sentence… to determine whether there is a sample sentence having the same or similar semantics as the input sentence among the cached sample sentences, in response to determining that there is a sample sentence having the same or similar semantics as the input sentence among the cached sample sentences, acquire cached response content corresponding to the sample sentence having the same or similar semantics as the input sentence as first response content of the input sentence, in response to determining that there is no sample sentence having the same or similar semantics as the input sentence among the cached sample sentences, send at least one of the input -37-sentence or the collected voice signals to a server, and receive transmitted response content of the at least one of the input sentence or the collected voice signals from the server as the first response content of the input sentence, respond to the input sentence according to the response content of the input sentence,…; and the server, wherein the server is configured to: receive the at least one of the input sentence or the collected voice signals from the terminal device,… and send the transmitted response content of the at least one of the input sentence or the collected voice signals to the terminal device”: where the user’s question [which can be in text form and which is compared/matched to frequently asked questions in the local cache or is sent to a cloud server to obtain, from the cloud server, an answer to the question] in Guo is more specifically a text question [“input sentence”] generated by performing speech-to-text [“voice recognition”] on a question spoken by a user at the Q&A device [i.e. by the “terminal device”], where the cloud server is “configured to: receive the input sentence from the terminal device,… and send the response content of the input sentence to the terminal device” [i.e. the speech-to-text-generated text question “input sentence” which is sent from the Q&A device to the cloud server is received by the cloud server, and the answer to the speech-to-text-generated text question is sent by the cloud server to the Q&A device]) Therefore, it would have been obvious to one of ordinary skill in the art at the time of effective filing to perform a simple substitution of one type of user-input text question with another because the prior art teaches the claimed invention except for the substitution of a user-input text question which is not necessarily generated by client-side speech-to-text with a user-input text question which is. Petrov suggests that a user-input text question which is generated by client-side speech-to-text was known in the art. One of ordinary skill in the art could have substituted one type of user-input text question with another to obtain the predictable results of a system including a cloud server and a Q&A user computer device which receives a input question from a user, determines if there is a matching frequently asked question that matches the input question in a local cache containing frequently-asked-question/answer pairs, and if there is a matching frequently asked question in the local cache, provides an answer associated with the matching frequently asked question to the user, and if there is no matching frequently asked question in the local cache, sends the input question to the cloud server, receives an answer to the input question from the cloud server, and provides the answer received from the cloud server to the user (as per Guo) where the input question is an input text question that is generated by performing speech-to-text processing on a question spoken by the user at the Q&A user computer device (as per Petrov). Guo, in view of Petrov, do not, but Ducatel suggests perform semantic matching between the input sentence and cached sample sentences to determine whether there is a sample sentence having the same or similar semantics as the input sentence among the cached sample sentences; (paragraphs 10, 17, 40-44 and 56; Ducatel describes [in paragraphs 10 and 56] where a submitted question/query can be analyzed and compared with “the queries held in a document retrieval system” [i.e. “a number of queries” to which a submitted question/question is compared with, as per paragraph 56] to determine if the submitted question has sufficient semantic similarity to one [or more] of the queries held in the document retrieval system, and if so, documents which are pertinent to the query [or queries] are returned to the user who submitted the question [similar to how the associated answer for a matching frequently asked question in Guo is provided to the user]. “If the submitted question has sufficient semantic similarity to one [or more] of the queries held within the document retrieval system” [paragraph 56] suggests where the submitted question is compared with every/each one of the multiple “number of queries” that are held within the document retrieval system, because, in order to accurately assess which of the “number of queries” held in the document retrieval system do or do not match the submitted question [thereby accurately determining which one or more of “the” queries held in the document retrieval system match the submitted question], the submitted question logically needs to be compared with every/each one of the multiple queries held within the document retrieval system. Ducatel [paragraphs 17 and 40-44] further describes where semantic similarity between text passages can be numeric quantities [a PSSM from 0 to 1], and where text passages can be a complete sentence [paragraph 17], and where a numerical indication of semantic similarity [the PSSM] can be calculated [paragraph 40]. Ducatel thus suggests “perform semantic matching between the input sentence and cached sample sentences to determine whether there is a sample sentence having the same or similar semantics as the input sentence among the cached sample sentences”: where the matching/comparison of the user input question [a speech-to-text-generated text input question in the Guo/Petrov combination] to frequently asked question[s] in the local cache is more specifically a semantic similarity matching/comparison [i.e. “semantically matching”] that matches/compares, for each/every one of the frequently asked questions in the local cache [“cached sample sentences”], the user input question to the frequently asked question in the local cache by determining the semantic similarity between the user input question and the frequently asked question, and by determining whether the semantic similarity between the user input question and the frequently asked question is sufficient for the user input question and the frequently asked question to be considered a match) Therefore, it would have been obvious to one of ordinary skill in the art at the time of effective filing to perform a simple substitution of one type of user-question-to-stored-question comparison with another because the prior art teaches the claimed invention except for the substitution of user-question-to-stored-question comparison which does not necessarily determine, for each of a plurality of stored questions, how semantically similar the user question and the stored question are with user-question-to-stored-question comparison which does. Ducatel teaches that user-question-to-stored-question comparison which determines, for each of a plurality of stored questions, how semantically similar the user question and the stored question are was known in the art. One of ordinary skill in the art could have substituted one type of user-question-to-stored-question comparison with another to obtain the predictable results of a system including a cloud server and a Q&A user computer device which receives a input question from a user, determines if there is a matching frequently asked question that matches the input question in a local cache containing frequently-asked-question/answer pairs, and if there is a matching frequently asked question in the local cache, provides an answer associated with the matching frequently asked question to the user, and if there is no matching frequently asked question in the local cache, sends the input question to the cloud server, receives an answer to the input question from the cloud server, and provides the answer received from the cloud server to the user (as per Guo) where the input question is an input text question that is generated by performing speech-to-text processing on a question spoken by the user at the Q&A user computer device (as per Petrov) where the determining of if there is a matching frequently asked question that matches the input question in a local cache containing frequently-asked-question/answer pairs determines, for each of the frequently asked questions in the local cache, if there is sufficient semantic similarity between the input question and the frequently asked question in the local cache (as suggested by Ducatel). Guo, in view of Petrov and Ducatel, do not, but Tseretopoulos suggests the server, wherein the server is configured to: receive the at least one of the input sentence or the collected voice signals from the terminal device, and perform semantic understanding of the at least one of the input sentence or the collected voice signals according to a knowledge base stored on the server to acquire the transmitted response content of the at least one of the input sentence or the collected voice signals, and send the transmitted response content of the at least one of the input sentence or the collected voice signals to the terminal device (Figure 1; paragraphs 38, 40, 49, 52, 76-77; Tseretopoulos describes where a “server-side” system [conversational analysis system 102 in Figure 1] receives an input from a client device input which can be text or audio is analyzed to determine an intent [e.g., a particular question to which a response may be generated] and generating a response based on the determined intent [paragraph 40], where a result of intent determination can be a set of lexical semantics of a received input which can be provided to a response analysis engine [paragraph 49], where a determined intent of the received input can be provided to a response analysis engine which determines an ideal response, where the response analysis engine determines a responsive answer to a question when the determined intent is a question, where the answer is provided to an NLG engine for generation of a sentence [at least suggested to be an answer sentence], and generates a response to the determined intent [paragraph 52] and where a conversational input which is analyzed to determine an intent can be text or audio [paragraphs 76-77] and where intent can be used to determine a meaning of a particular request and may be used to determine a particular response to be generated [paragraph 77]. Paragraphs 49 and 52 suggest where the determined intent includes semantic information of the received input because paragraph 52 describes where the determined intent is provided to a response analysis engine and paragraph 49 describes where “the result of the intent determination can be a set of lexical semantics of the received input” [semantic information] “can then be provided to the response analysis engine” [i.e. both the set of lexical semantics which are the result of the intent determination and the determined intent are provided to the same response analysis engine, which suggests that the determined intent is “the result of the intent determination which includes the set of lexical semantics]. Paragraph 77 further suggests where the determined intent [which is suggested to be semantic information as per paragraphs 49 and 52] reflects/indicates the meaning of the input because it states that “The intent associated with the input… can be used to determine meaning of a particular request associated with the conversational input” [the meaning of the input logically can’t be determined from the intent if the intent contained no information about the meaning of the input]. The determined intent is thus suggested to be a form of “understanding” of the “semantics” of the input which is determined. Paragraph 46 describes where operations of the conversational analysis system are performed by a processor included in the conversational analysis system executing instructions, and paragraph 47 describes where software instructions can be stored in a non-transitory medium, and paragraph 24 describes where a system comprising “at least one process” [at least suggested to be “at least one processor”] “and a memory operatively coupled to the at least one processor where the memory stores instructions that when executed cause the at least one processor to perform the operations”. These portions suggest where the “server-side” conversational analysis system includes/”stores” a “base” of “knowledge” containing all data for all functions/operations, and where the functions of the conversational analysis system [including the intent determination] are performed “according to” the conversational analysis system’s “knowledge base” [each function is performed according to respective software-specific software instructions stored in the “knowledge base”]. Additionally/alternatively, all data that exists in the conversational analysis system can be interpreted as a “base” of “knowledge” that is “stored” in the conversational analysis system. Additionally/alternatively, the specific portion of the memory storing software instructions used particularly to perform the intent determination can be interpreted as the claimed “knowledge base” Tseretopoulos suggests “the server, wherein the server is configured to: receive the at least one of the input sentence or the collected voice signals from the terminal device, and perform semantic understanding of the at least one of the input sentence or the collected voice signals according to a knowledge base stored on the server to acquire the transmitted response content of the at least one of the input sentence or the collected voice signals, and send the transmitted response content of the at least one of the input sentence or the collected voice signals to the terminal device”: where, in the Guo/Petrov/Ducatel combination, the cloud server which receives the speech-to-text-generated text user’s question [“the input sentence”] which is not matched to [i.e. not determined to be sufficiently semantically similar to] a frequently asked question in the local cache and provides an answer to the user’s question to the Q&A device more specifically determines the intent of the speech-to-text-generated text user’s question [i.e. the intent of “the input sentence”, where the intent includes semantic information indicating the meaning of the particular question that the user has asked], and determines/generates/”acquires”, based on the determined intent, the answer received by the Q&A device from the cloud server [i.e. “the response content of the input sentence” if there is no matching frequently asked question in the local cache]. The determining of the intent of the speech-to-text-generated text user’s question is suggested to be a form of “semantic understanding of the input sentence” in the sense that the determined intent [suggested to be semantic information that indicates the meaning of input text, as discussed above] reflects what the server “understood” the meaning/“semantics” of the speech-to-text-generated user’s question to be [such that the process that determines the intent is suggested to be a process the semantically understands the speech-to-text-generated user’s question]. The determining of the intent of the speech-to-text-generated text user’s question is also suggested to be “according to a knowledge base stored on the server” [as discussed above, performed according to intent determination software instructions stored in a particular portion of the server’s memory, where any of the particular potion, the entire memory storing all data for all functions, and all data in the cloud server can be interpreted as the claimed “knowledge base stored on the server”]) Therefore, it would have been obvious to one of ordinary skill in the art at the time of effective filing to perform a simple substitution of one type of server-side question answering with another because the prior art teaches the claimed invention except for the substitution of server-side question answering which does not necessarily determine an answer to a question based on semantic understanding of the question with server-side question answering which does. Tseretopoulos suggests that server-side question answering which determines an answer to a question based on semantic understanding of the question was known in the art. One of ordinary skill in the art could have substituted one type of server-side question answering with another to obtain the predictable results of a system including a cloud server and a Q&A user computer device which receives a input question from a user, determines if there is a matching frequently asked question that matches the input question in a local cache containing frequently-asked-question/answer pairs, and if there is a matching frequently asked question in the local cache, provides an answer associated with the matching frequently asked question to the user, and if there is no matching frequently asked question in the local cache, sends the input question to the cloud server, receives an answer to the input question from the cloud server, and provides the answer received from the cloud server to the user (as per Guo) where the input question is an input text question that is generated by performing speech-to-text processing on a question spoken by the user at the Q&A user computer device (as per Petrov) where the determining of if there is a matching frequently asked question that matches the input question in a local cache containing frequently-asked-question/answer pairs determines, for each of the frequently asked questions in the local cache, if there is sufficient semantic similarity between the input question and the frequently asked question in the local cache (as suggested by Ducatel) where the cloud server determines intent of the input question and determines the provided answer received from the cloud server based on the determined intent (as per Tseretopoulos). As per Claim 11, Guo suggests wherein at least one of the cached sample sentences or the response content of the cached sample sentences is pre-configured and/or updated based on at least one of: a manual configuration by a user of the terminal device; an initial system configuration; configurations based on a user profile at the terminal device or at a server; or big data obtained by the server (Figures 3 and 5; paragraphs 21, 41, 43, 59-65, 67, 73-78; At a minimum, the local cache and its question/answer pairs [see e.g. paragraph 21, where the questions of the question/answer pairs can be interpreted as “the cached sample sentences” and the answers of the question/answer pairs can be interpreted as “the response content of the cached sample sentences”] can be interpreted as being “pre-configured… based on… an initial system configuration” in the sense that the local cache’s questions and answers are, prior/”pre-“ to being used, “configured” to be what they are “based on” an “initial system configuration” [e.g. based on the system’s programming] which defines/”configures” the “system” to include the questions and answers of the local cache, where the “system configuration” can be interpreted as an “initial” “system configuration” in the sense that it defines how the system is configured prior to being used [when the system is “initiated” “initially” so that it can be used]). Claim(s) 12 is/are rejected under 35 U.S.C. 103 as being unpatentable over Guo, in view of Petrov, Ducatel, and Tseretopoulos, as applied to Claim 1, above, and further in view of Busey et al. (US 6,377,944), hereafter Busey. As per Claim 12, Guo, in view of Petrov, Ducatel, and Tseretopoulos, do not, but Busey suggests updating at least one of the cached sample sentences or response content of the cached sample sentences according to the input sentence and the first response content of the input sentence (col. 9, line 41 – col. 10, line 22; The combination [thus far] is as discussed in the rejection of claim 1 including where, if a speech-to-text-generated text user’s question does not match any of the most frequently asked questions in the local cache, the speech-to-text-generated text user’s question is sent by the Q&A user device to a cloud server and an answer to the user’s question is received by the Q&A user device from the cloud server. Busey [col. 9, line 41 – col. 10, line 22] similarly describes where a customer [at least suggested to be a form of “user”] enters a query [at least suggested to be a question since col. 9, lines 41-55 describes where a web page prompts for a question, and col. 9, lines 55-64 describe “a preferred embodiment” that is at least suggested to be a preferred embodiment of a query “uses natural language-based query syntax” and “free-form query”, and col. 10, lines 1-6 describes where “an answer” is “obtained for the query” and where “it is determined that the answer satisfactorily resolves the customer’s question”] and where an answer to the customer’s query is displayed [at least suggested to be displayed/presented to the user, see col. 9, line 67 – col. 10, line 6]. Busey [col. 9, line 41 – col. 10, line 22] further describes where “the question and answer pair are flagged for possible entry into the FAQ [a “Frequently Asked Questions” database], is used to update the FAQ [or other database], and is entered into the FAQ if the customer indicates that their question was adequately addressed. Busey thus at least suggests where a customer’s question and an answer obtained for the customer’s question are used to update a frequently asked questions database [e.g., if the customer indicates that his/her question was adequately addressed]. Busey thus suggests “updating at least one of the cached sample sentences or response content of the cached sample sentences according to the input sentence and the first response content of the input sentence”: interpreting “the cached sample sentences” as the collective set of the plurality of frequently asked questions in the question-and-answer pairs in the local cache of the Q&A device, and interpreting “response content of the cached sample sentences” as the corresponding answers to the plurality of frequently asked questions in the question-and-answer pairs in the local cache, the pair of the speech-to-text-generated text user’s question [i.e. “the input sentence”] and the answer to the user’s question received from the cloud server [i.e. “the response content of the input sentence” when there is no matching frequently asked question in the local cache] are added-to/entered-into the local cache containing frequently-asked-question-and-corresponding-answer pairs, thereby “updating” the local cache to include the question-and-answer pair of the speech-to-text-generated text user’s question and the answer to the user’s question received from the cloud server [thereby “updating… the cached sample sentences” and “response content of the cached sample sentences” collectively “according to the input sentence and the response content of the input sentence”, where “the cached sample sentences” are updated because the speech-to-text-generated text user’s question “input sentence” is added to the plurality of frequently asked questions in the local cache and where “response content of the cached sample sentences” is updated because the answer to the user’s question received from the cloud server “response content of the input sentence” is added to the corresponding answers to the plurality of frequently asked questions in the question-and-answer pairs in the local cache]) Therefore, it would have been obvious to one of ordinary skill in the art at the time of effective filing to combine prior art elements according to known methods because the prior art included each element claimed, although not necessarily in a single prior art reference, with the only difference between the claimed invention and the prior art being the lack of actual combination of the elements in a single prior art reference (Guo suggests a local cache containing frequently-asked-question-and-answer pairs and where a user’s question is received and where an answer to the user’s question is received from a cloud server, and Busey suggests where a customer’s question and an answer to the user’s question, as a question and answer pair, is used to update a frequently asked questions database). One of ordinary skill in the art could have combined the elements as claimed by known methods (by adding the function in Busey which adds, as a question and answer pair, the pair of a customer question and the answer to the customer’s question to an FAQ database to the set of functions performed on the user’s question, the answer to the user’s question obtained from the cloud server, and the local cache [a form of FAQ database] in Guo), and that in combination, each element merely performs the same function as it does separately (the adding of the question-and-answer pair of the user’s question and the answer to the user’s question to the FAQ database is a separate process relative to the receiving of the user’s question, the matching of the user’s question to the local cache frequently asked questions, the sending of the user’s question to the cloud server, the receiving of an answer to the user’s question from the cloud server, and the providing of an answer to the user’s question to the user). The combination is the predictable results of a Q&A user computer device which receives a input question from a user, determines if there is a matching frequently asked question that matches the input question in a local cache containing frequently-asked-question/answer pairs, and if there is a matching frequently asked question in the local cache, provides an answer associated with the matching frequently asked question to the user, and if there is no matching frequently asked question in the local cache, sends the input question to a cloud server, receives an answer to the input question from the cloud server, and provides the answer received from the cloud server to the user (as per Guo) where the input question is an input text question that is generated by performing speech-to-text processing on a question spoken by the user at the Q&A user computer device (as per Petrov) where the determining of if there is a matching frequently asked question that matches the input question in a local cache containing frequently-asked-question/answer pairs determines, for each of the frequently asked questions in the local cache, if there is sufficient semantic similarity between the input question and the frequently asked question in the local cache (as suggested by Ducatel) where the cloud server determines intent of the input question and determines the provided answer received from the cloud server based on the determined intent (as per Tseretopoulos) where the input question and the answer to the input question received from the cloud server are used to form a question-and-answer pair that is added to the local cache containing frequently-asked-question/answer pairs (as suggested by Busey). Claim 13 is/are rejected under 35 U.S.C. 103 as being unpatentable over Guo, in view of Petrov, Ducatel, Tseretopoulos, and Busey, as applied to claim 12, above, and further in view of Gray (US 9,990,176). As per Claim 13, Guo, in view of Petrov, Ducatel, and Tseretopoulos, do not, but Busey suggests wherein the updating at least one of the cached sample sentences and the response content of the cached sample sentences according to the input sentence and the first response content of the input sentence further comprises:… updating at least one of the cached sample sentences and the response content of the cached sample sentences according to the input sentence and the first response content of the input sentence (col. 9, line 41 – col. 10, line 22; Same combination as discussed in the rejection of claim 12, above) Therefore, it would have been obvious to one of ordinary skill in the art at the time of effective filing to combine prior art elements according to known methods because the prior art included each element claimed, although not necessarily in a single prior art reference, with the only difference between the claimed invention and the prior art being the lack of actual combination of the elements in a single prior art reference (Guo suggests a local cache containing frequently-asked-question-and-answer pairs and where a user’s question is received and where an answer to the user’s question is received from a cloud server, and Busey suggests where a customer’s question and an answer to the user’s question, as a question and answer pair, is used to update a frequently asked questions database). One of ordinary skill in the art could have combined the elements as claimed by known methods (by adding the function in Busey which adds, as a question and answer pair, the pair of a customer question and the answer to the customer’s question to an FAQ database to the set of functions performed on the user’s question, the answer to the user’s question obtained from the cloud server, and the local cache [a form of FAQ database] in Guo), and that in combination, each element merely performs the same function as it does separately (the adding of the question-and-answer pair of the user’s question and the answer to the user’s question to the FAQ database is a separate process relative to the receiving of the user’s question, the matching of the user’s question to the local cache frequently asked questions, the sending of the user’s question to the cloud server, the receiving of an answer to the user’s question from the cloud server, and the providing of an answer to the user’s question to the user). The combination is the predictable results of a Q&A user computer device which receives a input question from a user, determines if there is a matching frequently asked question that matches the input question in a local cache containing frequently-asked-question/answer pairs, and if there is a matching frequently asked question in the local cache, provides an answer associated with the matching frequently asked question to the user, and if there is no matching frequently asked question in the local cache, sends the input question to a cloud server, receives an answer to the input question from the cloud server, and provides the answer received from the cloud server to the user (as per Guo) where the input question is an input text question that is generated by performing speech-to-text processing on a question spoken by the user at the Q&A user computer device (as per Petrov) where the determining of if there is a matching frequently asked question that matches the input question in a local cache containing frequently-asked-question/answer pairs determines, for each of the frequently asked questions in the local cache, if there is sufficient semantic similarity between the input question and the frequently asked question in the local cache (as suggested by Ducatel) where the cloud server determines intent of the input question and determines the provided answer received from the cloud server based on the determined intent (as per Tseretopoulos) where the input question and the answer to the input question received from the cloud server are used to form a question-and-answer pair that is added to the local cache containing frequently-asked-question/answer pairs (as suggested by Busey). Guo, in view of Petrov, Ducatel, Tseretopoulos, and Busey, do not, but Gray suggests wherein the updating at least one of the cached sample sentences and the response content of the cached sample sentences according to the input sentence and the first response content of the input sentence further comprises: determining an acquisition frequency of the input sentence; comparing the acquisition frequency of the input sentence to a first preset threshold; and in response to determining that the acquisition frequency of the input sentence is greater than the first preset threshold, updating at least one of the cached sample sentences and the response content of the cached sample sentences according to the input sentence and the first response content of the input sentence (col. 36, line 33 - col. 37, line 12; col. 37, lines 32-52; Figures 1 and 5; The combination [thus far] is as discussed in the rejection of claim 12, above, including where Busey suggests where the speech-to-text-generated text user’s question and the answer to the user’s question received from the cloud server are, as a question-and-answer pair, used to update the local cache [by adding the question-and-answer pair of the speech-to-text-generated text user’s question and the answer to the user’s question received from the cloud server to the set of question-and-answer pairs in the local cache]. Gray describes [col. 37, lines 32-52 and Figures 1 and 5] where response data for a frequently asked question is, prior to receiving an utterance of a frequently asked question, stored locally on an electronic device [where the electronic device is suggested to be a user device as per Figure 1 which depicts electronic device 10 as a phone], which is similar to how Guo’s system stores answer data for frequently asked questions in the local cache. Gray [col. 37, lines 32-52] describes where a criterion for a question to be considered “frequently asked” [thereby leading to response data for the “frequently asked” utterance data being locally stored] is that the question is asked “in excess of a threshold value”. Gray thus suggests “wherein the updating at least one of the cached sample sentences and the response content of the cached sample sentences according to the input sentence and the first response content of the input sentence further comprises: determining an acquisition frequency of the input sentence; comparing the acquisition frequency of the input sentence to a first preset threshold; and in response to determining that the acquisition frequency of the input sentence is greater than the first preset threshold, updating at least one of the cached sample sentences and the response content of the cached sample sentences according to the input sentence and the first response content of the input sentence”: where a condition for the question-and-answer pair of the speech-to-text-generated text user’s question [“the input sentence”] and the answer to the user’s question received from the cloud server [“the response content of the input sentence” if there is no matching frequently asked question in the local cache] being added to the local cache [thereby “updating the cached sample sentences and the response content of the cached sample sentences according to the input sentence and the response content of the input sentence”] is determining that a number-of-times/”frequency” that the user’s question has been asked exceeds/”is greater than” a threshold value [a “first preset threshold”]. Determining that the number of times that the user’s question has been asked exceeds a threshold value is at least suggested to include determining the number of times that the user’s question has been asked and comparing that number of times to the threshold value [because the determination that the number of times is greater/less-than/equal-to the threshold value logically cannot be made without knowing what the number of times is, and without comparing the number of times to the threshold value]. The number-of-times/”frequency” that the user’s question has been asked is suggested to be “an acquisition frequency of the input sentence” because a device logically cannot determine a question has been asked or keep count of how many times a question has been asked if it did not acquire the question from the user, and the number of times that the user’s question has been asked can be interpreted as a frequency/number-of-times “of the input sentence” [i.e. of the speech-to-text-generated text user’s question because the user’s question is the same question represented by the speech-to-text-generated user’s question].) Therefore, it would have been obvious to one of ordinary skill in the art at the time of effective filing to perform a simple substitution of one type of frequently asked question data with another because the prior art teaches the claimed invention except for the substitution of frequently asked question data which is not necessarily stored if a corresponding question been asked a number of times that exceeds a threshold value with frequently asked question data which is. Gray suggests that frequently asked question data which is stored if a corresponding question has been asked a number of times that exceeds a threshold value was known in the art. One of ordinary skill in the art could have substituted one type of frequently asked question data with another to obtain the predictable results of a Q&A user computer device which receives a input question from a user, determines if there is a matching frequently asked question that matches the input question in a local cache containing frequently-asked-question/answer pairs, and if there is a matching frequently asked question in the local cache, provides an answer associated with the matching frequently asked question to the user, and if there is no matching frequently asked question in the local cache, sends the input question to a cloud server, receives an answer to the input question from the cloud server, and provides the answer received from the cloud server to the user (as per Guo) where the input question is an input text question that is generated by performing speech-to-text processing on a question spoken by the user at the Q&A user computer device (as per Petrov) where the determining of if there is a matching frequently asked question that matches the input question in a local cache containing frequently-asked-question/answer pairs determines, for each of the frequently asked questions in the local cache, if there is sufficient semantic similarity between the input question and the frequently asked question in the local cache (as suggested by Ducatel) where the cloud server determines intent of the input question and determines the provided answer received from the cloud server based on the determined intent (as per Tseretopoulos) where the input question and the answer to the input question received from the cloud server are used to form a question-and-answer pair that is added to the local cache containing frequently-asked-question/answer pairs (as suggested by Busey) where the question-and-answer pair formed by the input question and the answer to the input question is added to the local cache if the input question has been asked a number of times that exceeds a threshold value (as suggested by Gray). Claim 4 is/are rejected under 35 U.S.C. 103 as being unpatentable over Guo, in view of Petrov, Ducatel, and Tseretopoulos, as applied to Claim 1, above, and further in view of Kennewick et al. (US 2004/0044516), hereafter Kennewick. As per Claim 4, Guo, in view of Petrov, Ducatel, and Tseretopoulos, do not, but Kennewick suggests wherein: the first response content of the input sentence further comprises voice response content; and the responding to the input sentence according to the first response content of the input sentence includes carrying out a voice broadcast on the voice response content (paragraphs 86, 89, 155, 173; Figures 1-2; As discussed in the rejection of claim 1 Guo [paragraph 21] suggests where a response to a question can be output via audio output. Kennewick [paragraph 86, 89, 155, 173, and Figures 1 and 2] describes where a speech unit receives utterances of a user and where speech received from a main unit is annunciated by the speaker [paragraph 86, at least suggesting that outputting speech by the speech unit outputs the speech within hearing range of the user] and where a response from a server-side element [agents 106 in paragraph 89 are in the main unit of Figure 1] is sent to a “client-side” device [i.e. sent to the speech unit which receives the user utterances] in the form of speech [i.e. “voice content” generated by processing a response string with text to speech processing, see paragraphs 89 and 173] and where “responses to users” are “results of questions” [paragraph 89] and where a domain agent among agents 106 creates a satisfactory response to a question which is formatted into a format used by text to speech engine 124 [Figure 2, paragraph 173, at least suggesting an embodiment where the response string in paragraph 89 is a response to a question, see also paragraph 155 which describes which describes where a question is asked by a user] Kennewick thus suggests “wherein: the first response content of the input sentence further comprises voice response content; and the responding to the input sentence according to the first response content of the input sentence includes carrying out a voice broadcast on the voice response content”: the answer received by the Q&A device from the cloud server [“the response content of the input sentence”] is, instead, in the form of speech data [speech-answer/“voice response” data “content”] which is sent by the cloud server to the Q&A device and which is to be output audibly by the Q&A device [audibly-outputting-the-speech-answer-received-from-the-cloud-server/“carrying out a voice broadcast on the voice response content” as part of providing-the-answer-to-the-user’s-question-received-from-the-cloud-server-to-the-user/“responding to the input sentence according to the response content of the input sentence”]) Therefore, it would have been obvious to one of ordinary skill in the art at the time of effective filing to perform a simple substitution of one type of answer to a user question received by a user device from a remote device over a network with another because the prior art teaches the claimed invention except for the substitution of an answer to a user question received by a user device from a remote device over a network which is not necessarily speech data to be output by a speaker of the user device with an answer to a user question received by a user device from a remote device over a network which is. Kennewick teaches that an answer to a user question received by a user device from a remote device over a network which is speech data to be output by a speaker of the user device was known in the art. One of ordinary skill in the art could have substituted one type of answer to a user question received by a user device from a remote device over a network with another to obtain the predictable results of a Q&A user computer device which receives a input question from a user, determines if there is a matching frequently asked question that matches the input question in a local cache containing frequently-asked-question/answer pairs, and if there is a matching frequently asked question in the local cache, provides an answer associated with the matching frequently asked question to the user, and if there is no matching frequently asked question in the local cache, sends the input question to a cloud server, receives an answer to the input question from the cloud server, and provides the answer received from the cloud server to the user (as per Guo) where the input question is an input text question that is generated by performing speech-to-text processing on a question spoken by the user at the Q&A user computer device (as per Petrov) where the determining of if there is a matching frequently asked question that matches the input question in a local cache containing frequently-asked-question/answer pairs determines, for each of the frequently asked questions in the local cache, if there is sufficient semantic similarity between the input question and the frequently asked question in the local cache (as suggested by Ducatel) where the cloud server determines intent of the input question and determines the provided answer received from the cloud server based on the determined intent (as per Tseretopoulos) where the answer to the question received from the cloud server is in the form of speech data to be output by a speaker of the Q&A user computer device (as per Kennewick). Allowable Subject Matter The following is a statement of reasons for the indication of allowable subject matter: As per Claim(s) 2 (and similarly claim[s] 15, and consequently claim[s] 6-10 and 18 which depend on claim[s] 2 and 15), the prior art of record does not teach or suggest the combination of all limitations in claim(s) 1 and 2 together, including (i.e. in combination with the remaining limitations in claim[s] 1 and 2) wherein the first response content of the input sentence further comprises a control instruction and the responding to the input sentence according to the first response content of the input sentence includes performing a corresponding action according to the control instruction. Lee et al. (US 2014/0200896) teaches “the image processing apparatus may store a list of the simple sentence voice commands and the corresponding operations, and the transmitting the first voice command to the server comprises transmitting the first voice command if the first voice command is not retrieved from the list” (paragraph 15). Figure 3 depicts spoken commands corresponding to operations. Paragraph 59 describes where the voice commands include short simple sentences, and where, if an input voice command is not a simple sentence but a descriptive sentence, the voice command is not retrieved from the list and the audio processor may not determine a corresponding operation. Paragraph 60 describes where a descriptive sentence voice command is one that is not retrieved from the list, and where a descriptive sentence is transmitted to a server which transmits a control signal so that a display apparatus conducts the operation according to the control signal. Paragraph 63 describes where a display apparatus (“terminal device”) can perform functions of STT server 20 and interactive server 30 (where interactive server 30 “analyze[s] a voice command of a descriptive sentence”, see also descriptions of Figures 8 and 9 which describe “client-side” embodiments which perform STT conversion in the display apparatus [in Figures 8-9] and which perform descriptive sentence command processing in the audio processor [Figure 9]). Paragraph 74 describes where the list 210 may include both simple sentences and descriptive sentences (suggesting where the list may include more complex sentences). This reference does not appear to describe where a voice command is semantically matched to a plurality of cached sample sentences. Henmi et al. (US 2016/0147873) teaches “The automatic candidate Q providing module 423 of the suggest control module 420 acquires a list of input sentences semantically close to the input sentence entered by the user 10 with reference to the suggestion data 510 of the suggest DB 500” (paragraph 205) and “At step S16, the response determination module 212 determines whether the candidate list includes any reference text (Q). If the candidate list include some reference text (Q) (YES at step S16), the response determination module 212 determines the reference text (Q) having the highest score (or determined to be semantically closest to the input sentence) in the candidate list, acquires the response text (A) associated with the reference text (Q), and determines the response text (A) to be the response (step S17).” (paragraph 251). Paragraph 247 describes where a received input sentence is compared to all reference texts stored in knowledge data. Paragraphs 247-251 at least suggest where an input sentence is semantically compared to a plurality of stored/”cached” knowledge data. Paragraph 9 suggests that an input text is a question entered by a user. Paragraph 224 describes various embodiments of a reference text including a question, a word, an affirmative sentence, a negative sentence, a greeting, and the like. This reference does not appear to associate instructions/controls with reference texts. CN 112185370 A (publication date follows effective date of this application, international filing date precedes effective date of this application, does not appear to designate US) teaches “Specifically, in this step, when matching the text recognition result in the command word list, the following manner may be adopted: calculating semantic similarity between the text recognition result and each command word in the command word list; and taking the command word with the semantic similarity meeting the preset matching condition as the command word matched with the text recognition result. The preset matching condition may be a command word with the highest semantic similarity calculation result” and “In addition, the command word list in this step may include instructions corresponding to the command words, in addition to the command words. Therefore, in this step, after the command word matched with the text recognition result is obtained, the instruction corresponding to the matched command word is obtained from the command word list. For example, if the command word list includes 4 command words, i.e., "return", "back", "go back", and "back", the command word list also includes control instructions "return" corresponding to the four command words, i.e., the instructions corresponding to the four command words are "return" no matter the user utters "return" or "back"” (see Google Translation). This reference does clearly describe where semantic similarity between a text recognition result and each of a plurality of sentences is calculated (even though the examples of command words include “simple sentences”, similar to Lee). This reference also does not clearly qualify as prior art. 12093250 (62/821326 has earlier filing date than this application but does not clearly support the cited passage) teaches “the analysis system saves the questions in question store 265 to build a knowledge base of questions and their corresponding instructions that can be executed for answering the question. The knowledge base is used as a training data set for machine learning based models. The knowledge based is also used for recommending questions that can be asked using a given set of data sources. A user can search within the question store 265 for relevant questions. The user can select an appropriate question and the analysis system accesses the instructions for answering the questions and executes them”. It is not clear if this passage of this reference qualifies as prior art (Paragraphs 17 and 55 of the provisional application 62/821,326 appears to describe natural language questions being translated into instructions for processing the question but not where a knowledge base is built). I. semantically matching sentences/questions 2004/0030556 teaches “FIG. 21 illustrates a method for computing the closest semantic match between user articulated questions and stored semantic variants of the same” (paragraph 125). 2016/0217129 teaches “When a first sentence and a second sentence are sentences to be matched, a semantic matching degree between the first sentence and the second sentence can be obtained using the matching module” (paragraph 40) 2022/0114824 teaches “In particular, in a Japanese environment, it is difficult to calculate the similarity that accurately indicates how much the input sentence and each sentence of the plurality of sentences are semantically similar due to, for example, the large number of vocabularies and ambiguous sentence expressions. As a result, the probability of succeeding in specifying the sentence semantically similar to the input sentence from among the plurality of sentences may be 70%, or 80% or less” (paragraph 35). II. pairs of queries/requests and functions/executable-actions 5682542 teaches “The function execution management facility, which executes functions corresponding to requests, by making pairs of requests and responses” 2018/0373758 teaches “The method 300 includes collecting workload metrics for each of multiple pairs constituting a corresponding query and associated executable action set (act 301)” (paragraph 44). Paragraphs 41-42 describe where a query is processed by finding and executing an associated executable action set, where “The executable action set 220 is a set of computer-interpretable actions that may be executed (along with associated dependencies and orders for execution) that define how the query is to be executed against the data store”. This reference does not appear to describe where a query is a natural language sentence. 2018/0349256 teaches ““Open the browser to the banking homepage and login using a regular account”—and “Login” may be test action within testing framework. “Login” would be a test script containing code that executes to perform the test action of logging into the application, contained within a “keyword”. The natural language classifier would machine-interpret these keywords as classifications for the natural language descriptions. The set of data required to train the natural language classifier could be pairs of natural language descriptions and their respective test actions as seen in FIG. 1” (paragraph 76). As per Claim(s) 3 (and similarly claim[s] 16), the prior art of record does not teach or suggest the combination of all limitations in claim(s) 1 and 3 together, including (i.e. in combination with the remaining limitations in claim[s] 1 and 3) preconfiguring a sample sentence and a control instruction corresponding to the sample sentence; and updating the cached sample sentences and the response content of the cached sample sentences with the preconfigured sample sentence and the control instruction corresponding to the sample sentence As per Claim(s) 5 (and similarly claim[s] 17), the prior art of record does not teach or suggest the combination of all limitations in claim(s) 1 and 5 together, including (i.e. in combination with the remaining limitations in claim[s] 1 and 5) preconfiguring a sample sentence and a voice response content corresponding to the sample sentence; and updating the cached sample sentences and the response content of the cached sample sentences with the preconfigured sample sentence and the voice response content corresponding to the sample sentence. Claims 3, 5, 16, and 17, are interpreted as having an effective date of 08/29/2019. 2013/0339032 teaches “Thus, the storage unit 210 may have pre-stored the control command corresponding to the user's utterance intention. For example, when the user's utterance intention is to change the channel, the storage unit 210 may match and store the control command for changing the channel of the display apparatus 110. When the user's utterance intention is to schedule a recording, the storage unit 210 may match and store the control command for executing the scheduled recording for a specific program in the display apparatus 100” (paragraph 127). This reference does not describe where the intention is semantically matched, and also does not specifically describe where the storage is a cache. 2019/0138330 teaches “The above example is an example of purchasing a subway ticket. For other scenarios, it is generally necessary to set question-and-answer pairs with respect to respective requirements of the scenarios. For example, if it is a machine that purchases a train ticket, not only a “destination” and a “number of ticket” need to be known, but a “place of departure”, a “departure time”, and a “seat type” are also needed to be known, in order to obtain complete condition information to trigger a ticketing process. Therefore, it is necessary to set not only question-and-answer pairs corresponding to the “destination” and the “number of tickets”, but also question-and answer-pairs corresponding to the “place of departure”, the “departure time” and the “seat type”” (paragraph 84) and “wherein the question and answer pair comprises necessary information corresponding to an execution of the predetermined task” (paragraph 189). This reference appears to describe question-and-answer pairs that request information needed to implement a task, not where the answer is a task to be executed. 2023/0176829 (LATE filing date) teaches “Algorithm QAPR (Question-Answer Pair Ranking) shown in FIG. 17 depicts a question-answer pair selection technique utilized in some embodiments. The primary inputs are (1) a corpus of question answer pairs, QA=(q0, a0), (q1, a1), . . . (qn, an), where each question qi is a natural language description and the answer ai is the corresponding program, and (2) a question q* that represents the task in hand. The procedure returns a sequence RelevantQA=(qi0, ai0), . . . , (qik, aik) of k question-answer pairs to be used in the prompt. The QAPR algorithm is parameterized by a relevance metric R 420 on questions. A greater R(q, q′) score indicates that question q (and its answer) is more relevant to (answering) q′. At a high level, Algorithm QAPR orders the available question-answer pairs in QA based on their relevance 418 to q* and identifies the highest ranked question-answer pair (line 3)” (paragraph 161). This reference does not qualify as prior art. Double Patenting The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969). A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on nonstatutory double patenting provided the reference application or patent either is shown to be commonly owned with the examined application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. See MPEP § 717.02 for applications subject to examination under the first inventor to file provisions of the AIA as explained in MPEP § 2159. See MPEP § 2146 et seq. for applications not subject to examination under the first inventor to file provisions of the AIA . A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b). The filing of a terminal disclaimer by itself is not a complete reply to a nonstatutory double patenting (NSDP) rejection. A complete reply requires that the terminal disclaimer be accompanied by a reply requesting reconsideration of the prior Office action. Even where the NSDP rejection is provisional the reply must be complete. See MPEP § 804, subsection I.B.1. For a reply to a non-final Office action, see 37 CFR 1.111(a). For a reply to final Office action, see 37 CFR 1.113(c). A request for reconsideration while not provided for in 37 CFR 1.113(c) may be filed after final for consideration. See MPEP §§ 706.07(e) and 714.13. The USPTO Internet website contains terminal disclaimer forms which may be used. Please visit www.uspto.gov/patent/patents-forms. The actual filing date of the application in which the form is filed determines what form (e.g., PTO/SB/25, PTO/SB/26, PTO/AIA /25, or PTO/AIA /26) should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto-processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to www.uspto.gov/patents/apply/applying-online/eterminal-disclaimer. Claims 1, 2, 4, 12, 13, 14, 15, 19, and 20, are rejected on the ground of nonstatutory double patenting as being unpatentable over claims 1, 2, 3, 7, 8, 14, 15, and 17, of U.S. Patent No. 11,373,642, hereafter Parent Patent 1. Although the claims at issue are not identical, they are not patentably distinct from each other because the claims of this application are rendered obvious by the claims of Parent Patent 1. Claims 1, 14, 19, and 20, are taught by Claims 1 (everything up to and including “responding to the input sentence according to the first response content of the input sentence”), 8 (everything up to and including “respond to the input sentence according to the first response content of the input sentence”), and 15 of Parent Patent 1 (everything up to and including “send the transmitted response content of the input sentence to the terminal device”), and 17, respectively (“the input sentence” in the “in response to determining that there is no…” limitations in the independent claims of Parent Patent 1 reads on the “input sentence” only embodiment of “at least one of the input sentence or the collected voice signals” in the “in response to determining that there is no…” limitations of the independent claims of Parent Patent 1). Claims 2 and 15 are taught by Claims 7 and 14 of Parent Patent 1 (specifically the “control instruction” and “performing a corresponding action” embodiment of Claims 7 and 14 of Parent Patent 1). Claim 4 is taught by Claim 7 of Parent Patent 1 (specifically the “voice response content” and “carrying out a voice broadcast” embodiment of Claim 7 of Parent Patent 1). Claim 12 is taught by Claim 2 of Parent Patent 1. Claim 13 is taught by Claim 3 of Parent Patent 1. Claims 1-20 are rejected on the ground of nonstatutory double patenting as being unpatentable over claims 1-11, 14, 19, and 20, of U.S. Patent No. 12,223,957, hereafter Parent Patent 2. Although the claims at issue are not identical, they are not patentably distinct from each other because the claims of this application are rendered obvious by the claims of Parent Patent 2. Claims 1, 14, 19, and 20, are broader than Claims 1, 10, 19, and 20, of Parent Patent 2, respectively (Claims 1, 14, and 19 do not include the “wherein the first response content of the input sentence further comprises a control instruction and the responding to the input sentence according to the first response content of the input sentence includes performing a corresponding action according to the control instruction” limitation of Claims 1, 10, and 19 of Parent Patent 2, but are otherwise identical) Claims 2 and 15 are taught by Claims 1 and 10 of Parent Patent 2, respectively (Claims 2 and 15 contain the “wherein” clause which is in Claims 1 and 10 of Parent Patent 2 and which was not included in claims 1 and 14) Claim 3 is suggested by Claim 1 of Parent Patent 2 (one of the embodiments of “the first response content of the input sentence” in Claim 1 of Parent Patent 2 is “cached response content corresponding to the sample sentence” which is “acquired” “in response to determining that there is a sample sentence having the same or similar semantics as the input sentence among the cached sample sentences”, and in that embodiment, there is logically data in the system which associates a cached sample sentence to a corresponding control instruction, and in order for such data to exist within the system, the data logically must have been generated sometime before it is stored within the system [i.e. “preconfigured”] and the data logically must have been stored in the system [such that system memory including cached sample sentences and response content of the cached sample sentences” is “updated” to include the data]) Claim 4 corresponds to claim 2 of Parent Patent 2. Claim 5 is suggested by Claim 2 of Parent Patent 2 (one of the embodiments of “the first response content of the input sentence” in Claim 1 of Parent Patent 2 is “cached response content corresponding to the sample sentence” which is “acquired” “in response to determining that there is a sample sentence having the same or similar semantics as the input sentence among the cached sample sentences”, and in that embodiment for claim 2 of Parent Patent 2, there is logically data in the system which associates a cached sample sentence to a corresponding voice response content, and in order for such data to exist within the system, the data logically must have been generated sometime before it is stored within the system [i.e. “preconfigured”] and the data logically must have been stored in the system [such that system memory including cached sample sentences and response content of the cached sample sentences” is “updated” to include the data]) Claim 6 corresponds to claim 3 of Parent Patent 2. Claim 7 corresponds to claim 4 of Parent Patent 2. Claim 8 is suggested by Claims 2 and 5 of Parent Patent 2 (Claim 5 of Parent Patent 2 further limits the plurality of execution instructions to “be[ing] executed in a preset execution sequence”, and logically instructions are either executed in a preset execution sequence or they are not, and so if Claim 5 of Parent Patent 2 further limits the plurality of execution instructions to “be[ing] executed in a preset execution sequence”, then the remaining claim scope of Claim 2 of Parent Patent 2 which is not included in the scope of claim 5 of Parent Patent 2 is “wherein: the control instruction comprises a plurality of execution instructions that are to be executed with no preset execution sequence” [i.e. claim 8 of this application]) Claim 9 corresponds to claim 5 of Parent Patent 2. Claim 10 corresponds to claim 6 of Parent Patent 2. Claim 11 corresponds to claim 7 of Parent Patent 2. Claim 12 corresponds to claim 8 of Parent Patent 2. Claim 13 corresponds to claim 9 of Parent Patent 2. Claim 16 is suggested by Claim 10 of Parent Patent 2 (one of the embodiments of “the first response content of the input sentence” in Claim 10 of Parent Patent 2 is “cached response content corresponding to the sample sentence” which is “acquired” “in response to determining that there is a sample sentence having the same or similar semantics as the input sentence among the cached sample sentences”, and in that embodiment, there is logically data in the system which associates a cached sample sentence to a corresponding control instruction, and in order for such data to exist within the system, the data logically must have been generated sometime before it is stored within the system [i.e. “preconfigured”] and the data logically must have been stored in the system [such that system memory including cached sample sentences and response content of the cached sample sentences” is “updated” to include the data]) Claim 17 is suggested by Claim 11 of Parent Patent 2 (one of the embodiments of “the first response content of the input sentence” in Claim 10 of Parent Patent 2 is “cached response content corresponding to the sample sentence” which is “acquired” “in response to determining that there is a sample sentence having the same or similar semantics as the input sentence among the cached sample sentences”, and in that embodiment for claim 11 of Parent Patent 2, there is logically data in the system which associates a cached sample sentence to a corresponding voice response content, and in order for such data to exist within the system, the data logically must have been generated sometime before it is stored within the system [i.e. “preconfigured”] and the data logically must have been stored in the system [such that system memory including cached sample sentences and response content of the cached sample sentences” is “updated” to include the data]) Claim 18 is suggested by Claims 11 and 14 of Parent Patent 2 (Claim 14 of Parent Patent 2 further limits the plurality of execution instructions to “be[ing] executed in a preset execution sequence”, and logically instructions are either executed in a preset execution sequence or they are not, and so if Claim 14 of Parent Patent 2 further limits the plurality of execution instructions to “be[ing] executed in a preset execution sequence”, then the remaining claim scope of Claim 11 of Parent Patent 2 which is not included in the scope of claim 14 of Parent Patent 2 is “wherein: the control instruction comprises a plurality of execution instructions that are to be executed with no preset execution sequence” [i.e. claim 18 of this application]) Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to ERIC YEN whose telephone number is (571)272-4249. The examiner can normally be reached M-F 12:00PM -8:30PM EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, RICHEMOND DORVIL can be reached at (571)272-7602. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. EY 7/22/2026 /ERIC YEN/ Primary Examiner, Art Unit 2658
Read full office action

Prosecution Timeline

Dec 13, 2024
Application Filed
Jul 27, 2026
Non-Final Rejection mailed — §103, §112, §DP (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12699718
REASONING SYSTEM FOR PERFORMING QUESTIONING AND ANSWERING
2y 5m to grant Granted Aug 04, 2026
Patent 12699693
FRAMEWORK FOR LANGUAGE MODEL COPILOT DEVELOPMENT
2y 2m to grant Granted Aug 04, 2026
Patent 12694223
Meta-Tagging Based Configuration Transformation for Heterogeneous Systems
2y 3m to grant Granted Jul 28, 2026
Patent 12694211
SYSTEM AND METHOD FOR INTELLIGENT EVALUATION OF ARTIFICIAL INTELLIGENCE GENERATED TEXTS
2y 3m to grant Granted Jul 28, 2026
Patent 12657388
SYSTEMS AND METHODS FOR DETECTING EMERGING EVENTS
2y 3m to grant Granted Jun 16, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
85%
Grant Probability
97%
With Interview (+11.7%)
2y 9m (~1y 0m remaining)
Median Time to Grant
Low
PTA Risk
Based on 777 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month