DETAILED ACTION
Notice of AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Priority
Regarding Russian Patent App. No. RU2023113361 (filed 5/23/2023), receipt is acknowledged of certified copies of papers required by 37 CFR 1.55.
Information Disclosure Statement
The information disclosure statement submitted on 5/21/2024 has been considered.
The listing of references in the specification is not a proper information disclosure statement. 37 CFR 1.98(b) requires a list of all patents, publications, or other information submitted for consideration by the Office, and MPEP § 609.04(a) states, "the list may not be incorporated into the specification but must be submitted in a separate paper." Therefore, unless the references have been cited by the examiner on form PTO-892, they have not been considered. In particular, the Chang reference (see para. 00116), and the Sennrich reference (see para. 00119) have not been considered.
Drawings
The drawings are objected to as failing to comply with 37 CFR 1.84(p)(5) because they do not include the following reference sign(s) mentioned in the description: reference sign 714, referenced in at least paras. 00095-96, 00125-126, and 00152-53 does not appear in Fig. 7 or any other figure.
Corrected drawing sheets in compliance with 37 CFR 1.121(d) are required in reply to the Office action to avoid abandonment of the application. Any amended replacement drawing sheet should include all of the figures appearing on the immediate prior version of the sheet, even if only one figure is being amended. Each drawing sheet submitted after the filing date of an application must be labeled in the top margin as either “Replacement Sheet” or “New Sheet” pursuant to 37 CFR 1.121(d). If the changes are not accepted by the examiner, the applicant will be notified and informed of any required corrective action in the next Office action. The objection to the drawings will not be held in abeyance.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
Claims 3-6 are rejected under 35 U.S.C. 112(b) as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor regards as the invention.
Claim 3 recites “wherein the identifying further comprises” in line 1. However, it’s unclear if this “identifying” is meant to refer to the “identifying, using the sematic similarity ML model, from the fact data, … the textual representation of a respective fact” as set forth in claim 1, or the “identifying the respective fact relevant to the at least one of the given human request…” as set forth in claim 2. For purposes of compact prosecution, this will be interpreted as referring to the “identifying the respective fact relevant to the at least one of the given human request…” as set forth in claim 2.
Claim 4 recites “wherein the identifying further comprises” in line 1. However, it’s unclear if this “identifying” is meant to refer to the “identifying, using the sematic similarity ML model, from the fact data, … the textual representation of a respective fact” as set forth in claim 1, or the “identifying the respective fact relevant to the at least one of the given human request…” as set forth in claim 2. For purposes of compact prosecution, this will be interpreted as referring to the “identifying the respective fact relevant to the at least one of the given human request…” as set forth in claim 2.
Claims 5-6 depend from claim 4, do not remedy the deficiencies of claim 4, and are rejected for the same reasons explained above with respect to claim 4.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 1, 7-11, 13-14, and 17-19 are rejected under 35 U.S.C. 103 as being unpatentable over US 20230147096 A1, hereinafter referenced as GETSELEVICH, in view of Nie, Yixin, et al. "Combining fact extraction and verification with neural semantic matching networks." Proceedings of the AAAI conference on artificial intelligence. Vol. 33. No. 01. 2019, hereinafter referenced as NIE, and further in view of US 20230297603 A1, hereinafter referenced as M’HAMDI.
Regarding Claim 1
GETSELEVICH teaches:
А computer-implemented method of training a chatbot system to generate machine-generated answers to users’ requests of users of the chatbot system, the chatbot system including: (GETSELEVICH, para. 0015: “In at least one embodiment, systems and methods are used with chat bots or conversational artificial intelligence (AI) systems in order to store and retrieval data that may be stored with differed storage schemas and/or without a structured storage schema response to a query.”;
GETSELEVICH, para. 0020: “An extractive QA model may be trained and then used to provide a response to the information based query, such as by searching through the unstructured text to identify an answer to the input query.”
GETSELEVICH, para. 0091: “Accordingly, in at least one embodiment, computer systems are configured to implement one or more services that singly or collectively perform operations of processes described herein and such computer systems are configured with applicable hardware and/or software that enable performance of operations.”)
(i) a … machine-learning (ML) model configured to identify respective facts relevant to the users’ requests; and (GETSELEVICH, para. 0029: “By way of example, the extractive question answer model 312 may be a trained neural network that is utilized to extract one or more portions of an input sequence to answer a natural language question associated with such a sequence. As noted above, for an input such as “what colors can I paint the car” unstructured text may be evaluated to identify potential colors for the car, where those colors may then be presented to the user. For example, if unstructured text included natural language information such as “car colors are white, black, red, yellow, and gray” then the response to the question would be “white, black, red, yellow, and gray.”;
GETSELEVICH, para. 0030: “In various embodiments, training data 314 may be utilized to train the model 312, where the data includes a corpus of information, such as the multiQA dataset. As a result, the model 312 may be capable of extracting relevant facts directly from a corpus of unstructured text 316, which corresponds to information provided for the conversational system 304.”;
Examiner’s Note: Model 312 identifies facts (e.g., potential colors for a car) relevant to a user’s query)
(ii) a generative ML model to be trained to generate textual representations of the machine-generated answers to the users’ requests based on the users’ request and the respective facts; (GETSELEVICH, para. 0035: “In at least one embodiment, a generative response 408 may be enabled such that the answer 406 is provided in a sentence structure to the user. By way of example, a generative neural network may be utilized to receive, as an input, the answer 406 and then to determine an appropriate response incorporating the answer 406. In this example, the generative response 408 provides the answer 406 in a sentence format to the user. As will be appreciated, using the generative response 408 may provide an improved interaction experience for the user, where the user may feel as if they are engaging in a conversation with the system, as opposed to receiving only the information. Accordingly, the user may be encouraged to use the system for more purposes.”)
the method comprising:
acquiring dialogue data including (i) textual representations of human requests of dialogues of the users in a natural language; and (ii) textual representations of respective human answers of the dialogues, responsive to the human requests; (GETSELEVICH, para. 0030: “In various embodiments, training data 314 may be utilized to train the model 312, where the data includes a corpus of information, such as the multiQA dataset.”;
Examiner’s Note: Pursuant to MPEP 2131.01 II, to explain the meaning of the “multiQA dataset” term referenced in GETSELEVICH, the examiner further cites to Talmor, Alon, et al. "MultiQA: An empirical investigation of generalization and transfer in reading comprehension." Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019., which explains on pp. 4912-13 in section 3 that MultiQA uses 10 different question-context-answer datasets, and that these datasets asked crowdsourced workers to author questions and answers from Wikipedia articles, for example)
acquiring fact data including textual representations of facts; (GETSELEVICH, para. 0030: “By way of example, the corpus 316 may include information presented as natural language, such as sentences, paragraphs, CSV data, and the like. Furthermore, the corpus 316 may further include one or more structure datasets.”)
identifying, using the … ML model, from the fact data, …, the textual representation of a respective fact, which is relevant to at least one of the given human request and the respective human answer; (GETSELEVICH, para. 0030: “In various embodiments, training data 314 may be utilized to train the model 312, where the data includes a corpus of information, such as the multiQA dataset. As a result, the model 312 may be capable of extracting relevant facts directly from a corpus of unstructured text 316, which corresponds to information provided for the conversational system 304. By way of example, the corpus 316 may include information presented as natural language, such as sentences, paragraphs, CSV data, and the like. Furthermore, the corpus 316 may further include one or more structure datasets.”)
generating a training set of data including a plurality of training digital objects, a given one of which includes: (i) the textual representation of the given human request; (ii) the textual representation of the respective fact; and (iii) a respective label being the textual representation of the respective human answer responsive to the given human request; (GETSELEVICH, para. 0030: “In various embodiments, training data 314 may be utilized to train the model 312, where the data includes a corpus of information, such as the multiQA dataset.”;
Examiner’s Note: Pursuant to MPEP 2131.01 II, to explain the meaning of the “multiQA dataset” term referenced in GETSELEVICH, the examiner further cites to Talmor, Alon, et al. "MultiQA: An empirical investigation of generalization and transfer in reading comprehension." Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019., which explains on pp. 4912-13 in section 3 that MultiQA uses 10 different question-context-answer datasets, where the question corresponds to “(i) the textual representation of the given human request”, the context corresponds to “(ii) the textual representation of the respective fact” and the answer corresponds to “(iii) a respective label being the textual representation of the respective human answer responsive to the given human request”)
feeding the training set of data to the chatbot system, the feeding including: (GETSELEVICH, para. 0030: “In various embodiments, training data 314 may be utilized to train the model 312, where the data includes a corpus of information, such as the multiQA dataset.”)
However, GETSELEVICH fails to explicitly teach:
… semantic similarity…
… semantic similarity… for a given dialogue pair including textual representations of a given human request and a respective human answer responsive thereto from the dialogue data
for the given training digital object of the plurality of training digital objects, feeding a concatenation of the textual representation of the given human request and the textual representation of the respective fact to the generative ML model to generate the textual representation of a respective machine-generated answer to the given human request given a context of the respective fact; and
optimizing a difference between the respective machine-generated answer and the respective human answer to the given human request, thereby training the chatbot system to generate the machine-generated answers to the users’ requests.
However, in a related field of endeavor (fact extraction and verification with respect to natural language, see p. 6859, section 1), NIE teaches and makes obvious:
a semantic similarity machine-learning (ML) model configured to identify respective facts relevant to the users’ requests (NIE, p. 6861, section 4: “we first describe the architecture of our Neural Semantic Matching Network (NMSM), and then elaborate on the three subtasks of document retrieval, sentence selection, and claim verification”;
NIE, p. 6282, section 4.2.2: “Sentence selection is the extraction of evidential sentences from the retrieved documents regarding a claim”;
Examiner’s Note: NIE teaches a semantic matching neural network (corresponding to recited “semantic similarity ML model” that is configured to extract sentences as evidence from documentation in order to verify a claim; the GETSELEVICH-NIE combination now modifies the model 312 of GETSELEVICH to use semantic matching as in NIE)
identifying, using the semantic similarity ML model, from the fact data, for a given dialogue pair including textual representations of a given human request and a respective human answer responsive thereto from the dialogue data, the textual representation of a respective fact, which is relevant to at least one of the given human request and the respective human answer (NIE, p. 6861, section 4: “we first describe the architecture of our Neural Semantic Matching Network (NMSM), and then elaborate on the three subtasks of document retrieval, sentence selection, and claim verification”;
NIE, p. 6282, section 4.2.2: “Sentence selection is the extraction of evidential sentences from the retrieved documents regarding a claim”;
Examiner’s Note: As shown in Fig. 2, the input to the NMSN is a corpus of data and a claim to be verified, and the output is the evidence to verify the claim; the GETSELEVICH-NIE combination now modifies the model 312 of GETSELEVICH (which operates on an input sequence such as a user query, see para. 0029) to use semantic matching as in NIE, and to further consider the claim of NIE (corresponding to recited “human answer responsive thereto”), and using the teachings of NIE to verify the claim, which is the answer to the query of NIE)
Before the effective filing date of the present application, it would have been obvious to one of ordinary skill to combine the teachings of GETSELEVICH with NIE as explained above. As disclosed by NIE, one of ordinary skill would have been motivated to do so in order to provide “large-scale fact checking.” (p. 6860, section 1).
However, GETSELEVICH and NIE fail to explicitly teach:
for the given training digital object of the plurality of training digital objects, feeding a concatenation of the textual representation of the given human request and the textual representation of the respective fact to the generative ML model to generate the textual representation of a respective machine-generated answer to the given human request given a context of the respective fact; and
optimizing a difference between the respective machine-generated answer and the respective human answer to the given human request, thereby training the chatbot system to generate the machine-generated answers to the users’ requests.
However, in a related field of endeavor (natural language processing, see para. 0003, including question answering, see paras. 0030-0034), M’HAMDI teaches and makes obvious:
for the given training digital object of the plurality of training digital objects, feeding a concatenation of the textual representation of the given human request and the textual representation of the respective fact to the generative ML model to generate the textual representation of a respective machine-generated answer to the given human request given a context of the respective fact; and (M’HAMDI, para. 0072: “According to some embodiments, the input question (after prepending the input question with a [CLS] token) and the context are concatenated as a single packed sequence separated by a [SEP] token. That is, input to multi-lingual transformer network 600 includes a [CLS] token, tokens corresponding to a question, a [SEP] token, and tokens corresponding to context, in this order. Next, the embeddings of the context are input to linear layer 605.”;
Examiner’s Note: M'HAMDI teaches concatenating an input question and the context (which includes text of the “respective fact”); the GETSELEVICH-NIE-M’HAMDI combination now concatenates the input question and the extracted fact and inputs it into the generative neural network of GETSELEVICH to generate the “textual representation of a respective machine-generated answer to the human request” in view of the context provided by the multiQA training set of GETSELEVICH)
optimizing a difference between the respective machine-generated answer and the respective human answer to the given human request, thereby training the chatbot system to generate the machine-generated answers to the users’ requests. (M’HAMDI, para. 0090: “Accordingly, during the training process, the parameters and weights of the machine learning model are adjusted to increase the accuracy of the result (i.e., by minimizing a loss function which corresponds in some way to the difference between the current result and the target result).”;
M’HAMDI, para. 0091: “The term loss function refers to a function that impacts how a machine learning model is trained in a supervised learning model. Specifically, during each training iteration, the output of the model is compared to the known annotation information in the training data. The loss function provides a value for how close the predicted annotation data is to the actual annotation data. After computing the loss function, the parameters of the model are updated accordingly and a new set of predictions are made during the next iteration.”;
Examiner’s Note: the GETSELEVICH-NIE-M’HAMDI combination now compares the answer generated by the generative neural network of GETSELEVICH with the known answer from the multiQA database, and minimizes a loss function for how close the generated and known answer are as in M’HAMDI, where minimizing the loss function corresponds to recited “optimizing a difference between the respective machine-generated answer and the respective human answer to the given human request”)
Before the effective filing date of the present application, it would have been obvious to one of ordinary skill to combine the teachings of GETSELEVICH with NIE and M’HAMDI as explained above. As disclosed by M’HAMDI, one of ordinary skill would have been motivated to do so in order to extend a model to be able to perform a task in a second language. (para. 0003). As further disclosed by M’HAMDI, one of ordinary skill would further have been motivated to do so in order to implement the meta-learning techniques to improve multiple downstream learning tasks. (see para. 0093).
Regarding Claim 7
GETSELEVICH, NIE, and M’HAMDI teach the method of claim 1 as explained above. GETSELEVICH further teaches:
wherein the respective fact is for providing a context to the given dialogue pair. (GETSELEVICH, para. 0030: “In various embodiments, training data 314 may be utilized to train the model 312, where the data includes a corpus of information, such as the multiQA dataset. As a result, the model 312 may be capable of extracting relevant facts directly from a corpus of unstructured text 316, which corresponds to information provided for the conversational system 304.”;
Examiner’s Note: Pursuant to MPEP 2131.01 II, to explain the meaning of the “multiQA dataset” term referenced in GETSELEVICH, the examiner further cites to Talmor, Alon, et al. "MultiQA: An empirical investigation of generalization and transfer in reading comprehension." Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019., which explains on pp. 4912-13 in section 3 that MultiQA uses 10 different question-context-answer datasets, and that these datasets asked crowdsourced workers to author questions and answers from Wikipedia articles, for example)
Regarding Claim 8
GETSELEVICH, NIE, and M’HAMDI teach the method of claim 1 as explained above. However, GETSELEVICH and NIE fail to explicitly teach:
wherein the concatenation of the textual representation of the given human request and the textual representation of the respective fact comprises a concatenation of respective vector embeddings of the given human request and of the respective fact in a given vector space.
However, in a related field of endeavor (natural language processing, see para. 0003, including question answering, see paras. 0030-0034), M’HAMDI teaches and makes obvious:
wherein the concatenation of the textual representation of the given human request and the textual representation of the respective fact comprises a concatenation of respective vector embeddings of the given human request and of the respective fact in a given vector space. (M’HAMDI, para. 0061: “ In some examples, machine learning model 425 receives a query and context text. Machine learning model 425 combines the query and the context text to obtain an input text. Machine learning model 425 generates a word embedding corresponding to each word of the input text.”;
M’HAMDI, para. 0072: “According to some embodiments, the input question (after prepending the input question with a [CLS] token) and the context are concatenated as a single packed sequence separated by a [SEP] token. That is, input to multi-lingual transformer network 600 includes a [CLS] token, tokens corresponding to a question, a [SEP] token, and tokens corresponding to context, in this order. Next, the embeddings of the context are input to linear layer 605.”;
Examiner’s Note: M'HAMDI teaches concatenating an input question and the context (which includes text of the “respective fact”); the GETSELEVICH-NIE-M’HAMDI combination now concatenates the input question and the extracted fact in word embedding formats)
Before the effective filing date of the present application, it would have been obvious to one of ordinary skill to combine the teachings of GETSELEVICH with NIE and M’HAMDI as explained above. As disclosed by M’HAMDI, one of ordinary skill would have been motivated to do so in order to extend a model to be able to perform a task in a second language. (para. 0003). As further disclosed by M’HAMDI, one of ordinary skill would further have been motivated to do so in order to implement the meta-learning techniques to improve multiple downstream learning tasks. (see para. 0093).
Regarding Claim 9
GETSELEVICH, NIE, and M’HAMDI teach the method of claim 1 as explained above. However, GETSELEVICH and NIE fail to explicitly teach:
wherein the optimizing the difference comprises optimizing a loss function representative of the difference between the respective machine-generated answer and the respective human answer.
However, in a related field of endeavor (natural language processing, see para. 0003, including question answering, see paras. 0030-0034), M’HAMDI teaches and makes obvious:
wherein the optimizing the difference comprises optimizing a loss function representative of the difference between the respective machine-generated answer and the respective human answer. (M’HAMDI, para. 0090: “Accordingly, during the training process, the parameters and weights of the machine learning model are adjusted to increase the accuracy of the result (i.e., by minimizing a loss function which corresponds in some way to the difference between the current result and the target result).”;
M’HAMDI, para. 0091: “The term loss function refers to a function that impacts how a machine learning model is trained in a supervised learning model. Specifically, during each training iteration, the output of the model is compared to the known annotation information in the training data. The loss function provides a value for how close the predicted annotation data is to the actual annotation data. After computing the loss function, the parameters of the model are updated accordingly and a new set of predictions are made during the next iteration.”;
Examiner’s Note: the GETSELEVICH-NIE-M’HAMDI combination now compares the answer generated by the generative neural network of GETSELEVICH with the known answer from the multiQA database, and minimizes a loss function for how close the generated and known answer are as in M’HAMDI, where minimizing the loss function corresponds to recited “optimizing a difference between the respective machine-generated answer and the respective human answer to the given human request”)
Before the effective filing date of the present application, it would have been obvious to one of ordinary skill to combine the teachings of GETSELEVICH with NIE and M’HAMDI as explained above. As disclosed by M’HAMDI, one of ordinary skill would have been motivated to do so in order to extend a model to be able to perform a task in a second language. (para. 0003). As further disclosed by M’HAMDI, one of ordinary skill would further have been motivated to do so in order to implement the meta-learning techniques to improve multiple downstream learning tasks. (see para. 0093).
Regarding Claim 10
GETSELEVICH, NIE, and M’HAMDI teach the method of claim 1 as explained above. GETSELEVICH further teaches:
further comprising using the chatbot system for generating the machine-generated answers to the users’ requests, the using comprising: (GETSELEVICH, para. 0015: “In at least one embodiment, systems and methods are used with chat bots or conversational artificial intelligence (AI) systems in order to store and retrieval data that may be stored with differed storage schemas and/or without a structured storage schema response to a query.”;
GETSELEVICH, para. 0020: “An extractive QA model may be trained and then used to provide a response to the information based query, such as by searching through the unstructured text to identify an answer to the input query.”)
receiving the textual representation of an in-use human request of a given user; (GETSELEVICH, para. 0027: “In this example, an input processor 308 receives the input 308 and may perform one or more pre- or post-processing steps. For example, input processor 308 may include one or more NLP systems that evaluate an auditory input to extract one or more features from the input, among other options. Furthermore, in embodiments, input processor 308 may include a text processing system for preprocessing (e.g., tokenization, removal of punctuation, removal of stop words, stemming, lemmatization, etc.), feature extraction, and the like.”)
identifying, using the … ML model, from the fact data, the respective fact relevant to the in-use human request; (GETSELEVICH, para. 0029: “By way of example, the extractive question answer model 312 may be a trained neural network that is utilized to extract one or more portions of an input sequence to answer a natural language question associated with such a sequence. As noted above, for an input such as “what colors can I paint the car” unstructured text may be evaluated to identify potential colors for the car, where those colors may then be presented to the user. For example, if unstructured text included natural language information such as “car colors are white, black, red, yellow, and gray” then the response to the question would be “white, black, red, yellow, and gray.”;
GETSELEVICH, para. 0030: “In various embodiments, training data 314 may be utilized to train the model 312, where the data includes a corpus of information, such as the multiQA dataset. As a result, the model 312 may be capable of extracting relevant facts directly from a corpus of unstructured text 316, which corresponds to information provided for the conversational system 304.”;
Examiner’s Note: Model 312 identifies facts (e.g., potential colors for a car) relevant to a user’s query)
… thereby causing the generative ML model to generate the textual representation of a respective in-use machine-generated answer responsive to the in-use human request given the context of the respective fact. (GETSELEVICH, para. 0035: “In at least one embodiment, a generative response 408 may be enabled such that the answer 406 is provided in a sentence structure to the user. By way of example, a generative neural network may be utilized to receive, as an input, the answer 406 and then to determine an appropriate response incorporating the answer 406. In this example, the generative response 408 provides the answer 406 in a sentence format to the user. As will be appreciated, using the generative response 408 may provide an improved interaction experience for the user, where the user may feel as if they are engaging in a conversation with the system, as opposed to receiving only the information. Accordingly, the user may be encouraged to use the system for more purposes.”)
However, GETSELEVICH fails to explicitly teach:
… semantic similarity …
feeding a concatenation of the textual representations of the in-use human request and the respective fact to the generative ML model,
However, in a related field of endeavor (fact extraction and verification with respect to natural language, see p. 6859, section 1), NIE teaches and makes obvious:
a semantic similarity ML model (NIE, p. 6861, section 4: “we first describe the architecture of our Neural Semantic Matching Network (NMSM), and then elaborate on the three subtasks of document retrieval, sentence selection, and claim verification”;
NIE, p. 6282, section 4.2.2: “Sentence selection is the extraction of evidential sentences from the retrieved documents regarding a claim”;
Examiner’s Note: NIE teaches a semantic matching neural network (corresponding to recited “semantic similarity ML model” that is configured to extract sentences as evidence from documentation in order to verify a claim; the GETSELEVICH-NIE combination now modifies the model 312 of GETSELEVICH to use semantic matching as in NIE)
Before the effective filing date of the present application, it would have been obvious to one of ordinary skill to combine the teachings of GETSELEVICH with NIE as explained above. As disclosed by NIE, one of ordinary skill would have been motivated to do so in order to provide “large-scale fact checking.” (p. 6860, section 1).
However, GETSELEVICH and NIE fail to explicitly teach:
feeding a concatenation of the textual representations of the in-use human request and the respective fact to the generative ML model,
However, in a related field of endeavor (natural language processing, see para. 0003, including question answering, see paras. 0030-0034), M’HAMDI teaches and makes obvious:
feeding a concatenation of the textual representations of the in-use human request and the respective fact to the generative ML model, (M’HAMDI, para. 0072: “According to some embodiments, the input question (after prepending the input question with a [CLS] token) and the context are concatenated as a single packed sequence separated by a [SEP] token. That is, input to multi-lingual transformer network 600 includes a [CLS] token, tokens corresponding to a question, a [SEP] token, and tokens corresponding to context, in this order. Next, the embeddings of the context are input to linear layer 605.”;
Examiner’s Note: M'HAMDI teaches concatenating an input question and the context (which includes text of the “respective fact”); the GETSELEVICH-NIE-M’HAMDI combination now concatenates the input question and the extracted fact and inputs it into the generative neural network of GETSELEVICH)
Before the effective filing date of the present application, it would have been obvious to one of ordinary skill to combine the teachings of GETSELEVICH with NIE and M’HAMDI as explained above. As disclosed by M’HAMDI, one of ordinary skill would have been motivated to do so in order to extend a model to be able to perform a task in a second language. (para. 0003). As further disclosed by M’HAMDI, one of ordinary skill would further have been motivated to do so in order to implement the meta-learning techniques to improve multiple downstream learning tasks. (see para. 0093).
Regarding Claim 11
GETSELEVICH, NIE, and M’HAMDI teach the method of claim 1 as explained above. However, GETSELEVICH and NIE fail to explicitly teach:
wherein each one of the semantic similarity ML model and the generative ML model is a Transformer-based ML model.
However, in a related field of endeavor (natural language processing, see para. 0003, including question answering, see paras. 0030-0034), M’HAMDI teaches and makes obvious:
wherein each one of the semantic similarity ML model and the generative ML model is a Transformer-based ML model. (M’HAMDI, para. 0069: “In some examples, multi-lingual transformer network 500 comprises Bidirectional Encoder Representations from Transformers (BERT).”;
Examiner’s Note: the GETSELEVICH-NIE-M’HAMDI combination now modifies these models of GETSELEVICH to use the BERT pre-trained language model as in M’HAMDI)
Before the effective filing date of the present application, it would have been obvious to one of ordinary skill to combine the teachings of GETSELEVICH with NIE and M’HAMDI as explained above. As disclosed by M’HAMDI, one of ordinary skill would have been motivated to do so in order to extend a model to be able to perform a task in a second language. (para. 0003). As further disclosed by M’HAMDI, one of ordinary skill would further have been motivated to do so in order to implement the meta-learning techniques to improve multiple downstream learning tasks. (see para. 0093). One of ordinary skill would further understand the benefit of using the well-known, peer reviewed BERT language model as a starting point rather than building a neural network from scratch.
Regarding Claim 13
GETSELEVICH, NIE, and M’HAMDI teach the method of claim 1 as explained above. GETSELEVICH further teaches:
wherein the chatbot system is one of (i) a text-to-text chatbot system; (ii) text-to-speech chatbot system; (iii) a speech-to-text chatbot system; and (iv) a speech-to-speech chatbot system. (GETSELEVICH, para. 0015: “In at least one embodiment, systems and methods are used with chat bots or conversational artificial intelligence (AI) systems in order to store and retrieval data that may be stored with differed storage schemas and/or without a structured storage schema response to a query.”;
GETSELEVICH, para. 0017: “Various embodiments may be utilized to provide a response to a user input, which may be in the form of an auditory input, a textual input, a selective input (e.g., selecting a content element), or an instructional input, such as a data file that executes one or more operations within the interaction environment. Systems and methods may not only store relevant information as natural text in unstructured memory and answer flexible questions based on the information, but moreover, may retrieve pieces of information to use in commands. For example, a result may be associated with a textural or voice response to a user, as well as or additionally, fulfillment of one or more actions connected to the result.”)
Regarding Claim 14
GETSELEVICH teaches:
А server … the server comprising: (i) at least one processor and (ii) at least one non-transitory computer-readable medium comprising executable instructions that, when executed by the at least one processor, cause the system to: (GETSELEVICH, para. 0056: “In at least one embodiment, computer system 800 is a single processor desktop or server system, but in another embodiment computer system 800 may be a multiprocessor system.”;
GETSELEVICH, para. 0074: “ In at least one embodiment, memory device 1020 can be a dynamic random access memory (DRAM) device, a static random access memory (SRAM) device, flash memory device, phase-change memory device, or some other memory device having suitable performance to serve as process memory. In at least one embodiment memory device 1020 can operate as system memory for system 1000, to store data 1022 and instructions 1021 for use when one or more processors 1002 executes an application or process.”)
The remaining limitations correspond to the method of claim 1, and therefore this claim is rejected for the same reasons explained above with respect to claim 1.
Claim 17 depends from claim 14 and claims a server that corresponds to the method of claim 9, and is therefore rejected for the same reasons explained above with respect to claims 9 and 14.
Claim 18 depends from claim 14 and claims a server that corresponds to the method of claim 10, and is therefore rejected for the same reasons explained above with respect to claims 10 and 14.
Claim 19 depends from claim 14 and claims a server that corresponds to the method of claim 11, and is therefore rejected for the same reasons explained above with respect to claims 11 and 14.
Claims 2 and 15 are rejected under 35 U.S.C. 103 as being unpatentable over GETSELEVICH, NIE, and M’HAMDI and further in view of US 20240311348 A1, hereinafter referenced as LUTZ.
Regarding Claim 2
GETSELEVICH, NIE, and M’HAMDI teach the method of claim 1 as explained above. However, GETSELEVICH and NIE fail to explicitly teach:
feeding each one of: (i) the textual representations of the human requests; (ii) the textual representations of the human answers; and (iii) the textual representations of the facts, to the semantic similarity ML model, to generate respective vector embeddings of each one of (i) the textual representations of the human requests; (ii) the textual representations of the human answers; and (iii) the textual representations of the facts;
mapping the respective vector embeddings to a vector space; and
identifying the respective fact relevant to the at least one of the given human request and the respective human answer as being a fact associated with the respective vector embedding that is closest to the respective vector embedding of the at least one of the given human request and the respective human answer in the vector space.
However, in a related field of endeavor (natural language processing, see para. 0003, including question answering, see paras. 0030-0034), M’HAMDI teaches and makes obvious:
feeding each one of: (i) the textual representations of the human requests; (ii) the textual representations of the human answers; and (iii) the textual representations of the facts, to the semantic similarity ML model, to generate respective vector embeddings of each one of (i) the textual representations of the human requests; (ii) the textual representations of the human answers; and (iii) the textual representations of the facts; (M’HAMDI, para. 0061: “ In some examples, machine learning model 425 receives a query and context text. Machine learning model 425 combines the query and the context text to obtain an input text. Machine learning model 425 generates a word embedding corresponding to each word of the input text.”;
M'HAMDI, para. 0095: “QA is not considered a standard classification task with fixed classes. QA is not directly amenable to class distribution balancing across pseudo-task query and support sets. The following procedure is used to construct pseudo-tasks for QA from the (i.e., question, context, answer) span triplet data. A task T=(S, Q), is drawn by first randomly drawing q triplets, forming Q. The k/q most similar triplets to t are drawn from the remaining available data for each triplet t in Q, thus forming S. k is constrained to be a multiple of q. Similarity is calculated as cos(f(t.sub.1), f(t.sub.2)) for two triplets t.sub.1, t.sub.2, where f(.) is a representation of the concatenation of the triplet elements delimited by a space. In some cases, a cross-lingual extension to SBERT's pre-trained model is used.”
Examiner’s Note: M’HAMDI discloses concatenating the question-context-answer triplet and further discloses creating word embeddings for each word in the input to the ML model; the GETSELEVICH-NIE-M’HAMDI combination now uses the teachings of M’HAMDI to generate word embeddings for each of the words of the requests, answers, and facts (context))
mapping the respective vector embeddings to a vector space; and (M’HAMDI, para. 0095: “Similarity is calculated as cos(f(t.sub.1), f(t.sub.2)) for two triplets t.sub.1, t.sub.2, where f(.) is a representation of the concatenation of the triplet elements delimited by a space.”;
Examiner’s Note: the GETSELEVICH-NIE-M’HAMDI combination now uses the teachings of M’HAMDI to generate word embeddings mapped to a particular space as in M’HAMDI)
Before the effective filing date of the present application, it would have been obvious to one of ordinary skill to combine the teachings of GETSELEVICH with NIE and M’HAMDI as explained above. As disclosed by M’HAMDI, one of ordinary skill would have been motivated to do so in order to extend a model to be able to perform a task in a second language. (para. 0003). As further disclosed by M’HAMDI, one of ordinary skill would further have been motivated to do so in order to implement the meta-learning techniques to improve multiple downstream learning tasks. (see para. 0093).
However, GETSELEVICH, NIE, M’HAMDI fail to explicitly teach:
identifying the respective fact relevant to the at least one of the given human request and the respective human answer as being a fact associated with the respective vector embedding that is closest to the respective vector embedding of the at least one of the given human request and the respective human answer in the vector space.
However, in a related field of endeavor (large language models, see para. 0007), LUTZ teaches and makes obvious:
identifying the respective fact relevant to the at least one of the given human request and the respective human answer as being a fact associated with the respective vector embedding that is closest to the respective vector embedding of the at least one of the given human request and the respective human answer in the vector space. (LUTZ, para. 0089: “The application management component 130 then searches the structured database 106 to identify the category vector(s) and fact vector(s) that are closest to the query vector in vector space. In some implementations, the application management component 130 computes the distance between two vectors using the cosine similarity metric.”;
Examiner’s Note: the GETSELEVICH-NIE-M’HAMDI-LUTZ combination now compares a fact vector in a vector space to a query vector in a vector space to find the closest march as in LUTZ)
Before the effective filing date of the present application, it would have been obvious to one of ordinary skill to combine the teachings of GETSELEVICH with NIE, M’HAMDI, and LUTZ as explained above. As disclosed by LUTZ, one of ordinary skill would have been motivated to do so in order to utilize pattern-completion technologies to act as an “expert interrogator” to progressively resolve questions asked by users. (para. 0093).
Claim 15 depends from claim 14 and claims a server that corresponds to the method of claim 2, and is therefore rejected for the same reasons explained above with respect to claims 2 and 14.
Claims 3-6 and 16 are rejected under 35 U.S.C. 103 as being unpatentable over GETSELEVICH, NIE, M’HAMDI, and LUTZ, and further in view of US 20200372294 A1, hereinafter referenced as KOVAL.
Regarding Claim 3
GETSELEVICH, NIE, M’HAMDI, and LUTZ teach the method of claim 2 as explained above. However, GETSELEVICH, NIE, M’HAMDI, and LUTZ fail to explicitly teach:
wherein the identifying further comprises applying a k-nearest neighbors algorithm.
However, in a related field of endeavor (processing user queries for content, see para. 0024), KOVAL teaches and makes obvious:
wherein the identifying further comprises applying a k-nearest neighbors algorithm. (KOVAL, para. 0045: “In one embodiment, a subset of all local features may clustered using the k-nearest neighbors (KNN) algorithm and centers of those clusters may be found. Then each cluster may be subclustered to have a specified number of features in each cluster and again find centers of the new clusters.”;
Examiner’s Note: the GETSELEVICH-NIE-M’HAMDI-LUTZ-KOVAL combination now utilizes the kNN search algorithm of KOVAL when searching for facts to extract as in GETSELEVICH)
Before the effective filing date of the present application, it would have been obvious to one of ordinary skill to combine the teachings of GETSELEVICH with NIE, M’HAMDI, LUTZ, and KOVAL as explained above. As disclosed by KOVAL, one of ordinary skill would have been motivated to do so in order to use clustering algorithms to determine facts that are nearest to user query. (para. 0045). As disclosed by KOVAL, one of ordinary skill would have been motivated to do so because “organizing the feature data … into clusters may speed up the search, and thus increase the efficiency of the system 200 itself.” (para. 0045).
Regarding Claim 4
GETSELEVICH, NIE, M’HAMDI, and LUTZ teach the method of claim 2 as explained above. However, GETSELEVICH, NIE, M’HAMDI, and LUTZ fail to explicitly teach:
wherein the identifying further comprises applying a heuristic algorithm.
However, in a related field of endeavor (processing user queries for content, see para. 0024), KOVAL teaches and makes obvious:
wherein the identifying further comprises applying a heuristic algorithm. (KOVAL, para. 0054: “The selected candidates may be further ranked in order of relevance. For example, a comparison of features may be performed using the Okapi BM25 formula, which is a ranking function for ranking matching documents according to their relevance to a given search query.”;
Examiner’s Note: the GETSELEVICH-NIE-M’HAMDI-LUTZ-KOVAL combination now utilizes the BM25 formula of KOVAL when ranking potential facts of GETSELEVICH)
Before the effective filing date of the present application, it would have been obvious to one of ordinary skill to combine the teachings of GETSELEVICH with NIE, M’HAMDI, LUTZ, and KOVAL as explained above. As disclosed by KOVAL, one of ordinary skill would have been motivated to do so because “organizing the feature data … into clusters may speed up the search, and thus increase the efficiency of the system 200 itself.” (para. 0045). As further disclosed by KOVAL, one of ordinary skill would have been motivated to do so in order to rank the relevance of particular facts to concentrate only on the potential facts that are most relevant. (para. 0054).
Regarding Claim 5
GETSELEVICH, NIE, M’HAMDI, LUTZ, and KOVAL teach the method of claim 4 as explained above. However, GETSELEVICH, NIE, M’HAMDI, and LUTZ fail to explicitly teach:
wherein the heuristic algorithm comprises a ranking function.
However, in a related field of endeavor (processing user queries for content, see para. 0024), KOVAL teaches and makes obvious:
wherein the heuristic algorithm comprises a ranking function. (KOVAL, para. 0054: “The selected candidates may be further ranked in order of relevance. For example, a comparison of features may be performed using the Okapi BM25 formula, which is a ranking function for ranking matching documents according to their relevance to a given search query.”;
Examiner’s Note: the GETSELEVICH-NIE-M’HAMDI-LUTZ-KOVAL combination now utilizes the BM25 formula of KOVAL when ranking potential facts of GETSELEVICH)
Before the effective filing date of the present application, it would have been obvious to one of ordinary skill to combine the teachings of GETSELEVICH with NIE, M’HAMDI, LUTZ, and KOVAL as explained above. As disclosed by KOVAL, one of ordinary skill would have been motivated to do so because “organizing the feature data … into clusters may speed up the search, and thus increase the efficiency of the system 200 itself.” (para. 0045). As further disclosed by KOVAL, one of ordinary skill would have been motivated to do so in order to rank the relevance of particular facts to concentrate only on the potential facts that are most relevant. (para. 0054).
Regarding Claim 6
GETSELEVICH, NIE, M’HAMDI, LUTZ, and KOVAL teach the method of claim 5 as explained above. However, GETSELEVICH, NIE, M’HAMDI, and LUTZ fail to explicitly teach:
wherein the ranking function is a BM25 ranking function.
However, in a related field of endeavor (processing user queries for content, see para. 0024), KOVAL teaches and makes obvious:
wherein the ranking function is a BM25 ranking function. (KOVAL, para. 0054: “The selected candidates may be further ranked in order of relevance. For example, a comparison of features may be performed using the Okapi BM25 formula, which is a ranking function for ranking matching documents according to their relevance to a given search query.”;
Examiner’s Note: the GETSELEVICH-NIE-M’HAMDI-LUTZ-KOVAL combination now utilizes the BM25 formula of KOVAL when ranking potential facts of GETSELEVICH)
Before the effective filing date of the present application, it would have been obvious to one of ordinary skill to combine the teachings of GETSELEVICH with NIE, M’HAMDI, LUTZ, and KOVAL as explained above. As disclosed by KOVAL, one of ordinary skill would have been motivated to do so because “organizing the feature data … into clusters may speed up the search, and thus increase the efficiency of the system 200 itself.” (para. 0045). As further disclosed by KOVAL, one of ordinary skill would have been motivated to do so in order to rank the relevance of particular facts to concentrate only on the potential facts that are most relevant. (para. 0054).
Claim 16 depends from claim 14 and claims a server that corresponds to the method of claim 3, and is therefore rejected for the same reasons explained above with respect to claims 3 and 14.
Claims 12 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over GETSELEVICH, NIE, and M’HAMDI and further in view of US 20230334074 A1, hereinafter referenced as KULKARNI.
Regarding Claim 12
GETSELEVICH, NIE, and M’HAMDI teach the method of claim 11 as explained above. However, GETSELEVICH, NIE, and M’HAMDI fail to explicitly teach:
wherein the generative ML model is devoid of an encoder portion of the Transformer-based ML model.
However, in a related field of endeavor (large language models, see para. 0054), KULKARNI teaches and makes obvious:
wherein the generative ML model is devoid of an encoder portion of the Transformer-based ML model. (KULKARNI, para. 0054: “Specifically, the system 110 may use either encode only models (e.g., Sentence-Bidirectional Encoder Representations from Transformers (BERT)), the decoder only models (e.g., Generative Pre-trained Transformer (GPT)) or encoder-decoder models (e.g., Text-To-Text Transfer Transformer (T5), Bidirectional Auto-encoder Representations from Transformers (BART)) to obtain the vector representation of the query.”;
Examiner’s Note: the GETSELEVICH-NIE-M’HAMDI-KULKARNI combination now modifies these models of GETSELEVICH to use the GPT decoder-only transformer model as in KULKARNI)
Before the effective filing date of the present application, it would have been obvious to one of ordinary skill to combine the teachings of GETSELEVICH with NIE, M’HAMDI, and KULKARNI as explained above. One of ordinary skill would further understand the benefit of using the well-known, peer reviewed GPT language model as a starting point rather than building a neural network from scratch.
Claim 20 depends from claim 19 and claims a server that corresponds to the method of claim 12, and is therefore rejected for the same reasons explained above with respect to claims 12 and 19.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
US 20240362286 A1 (He). “Embodiments may implement a generative AI to provide an abstractive summary of search results relevant to a given search request or search query.” (para. 0039).
US 20180174020 A1 (Wu). “In summary, the disclosure generally relates to systems and methods for emotionally intelligent automated chatting. The systems and methods as described herein provide emotionally intelligent automated (or artificial intelligence (AI)) chatting by determining a context and an emotion of a conversation with a user.” (para. 0003).
Any inquiry concerning this communication or earlier communications from the examiner should be directed to MICHAEL C LEE whose telephone number is (571)272-4933. The examiner can normally be reached M-F 12:00 pm - 8:00 pm ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Omar Fernandez Rivas can be reached at 571-272-2589. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/MICHAEL C. LEE/Examiner, Art Unit 2128