Prosecution Insights
Last updated: October 02, 2026
Application No. 18/738,673

CONTEXT RECOMMENDATION FOR RETRIEVAL AUGMENTED GENERATION ARCHITECTURES

Non-Final OA §101§103
Filed
Jun 10, 2024
Examiner
ADMASU, MAHLIET TASEW
Art Unit
Tech Center
Assignee
Dell Products L.P.
OA Round
1 (Non-Final)
0%
Grant Probability
At Risk
1-2
OA Rounds
1y 1m
Est. Remaining
0%
With Interview

Examiner Intelligence

Grants only 0% of cases
0%
Career Allowance Rate
0 granted / 1 resolved
-60.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
3y 5m
Avg Prosecution
15 currently pending
Career history
12
Total Applications
across all art units

Statute-Specific Performance

§101
29.0%
-11.0% vs TC avg
§103
62.0%
+22.0% vs TC avg
§112
7.0%
-33.0% vs TC avg
Black line = Tech Center average estimate • Based on career data from 1 resolved cases

Office Action

§101 §103
DETAILED ACTION This communication is in response to the Application No. 18/738, 673 filed on June 10, 2024 in which Claims 1 – 20 are presented for examination. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claim 1-20 are rejected under 35 U.S.C. 101 because these claimed inventions are directed to an abstract idea without significantly more. Regarding Claim 1: Step 1: Claim 1 is a method type claim. Therefore, Claims 1-15 fall within one of the four statutory categories (i.e., process, machine, manufacture, or composition of matter). 2A Prong 1: If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation by mathematical calculation but for the recitation of generic computer components, then it falls within the “Mathematical Concepts” grouping of abstract ideas. analyzing the large language model request […] (mental process – analyzing the large language model request may be performed mentally or using pen and paper by a user observing/reading the received request, considering its subject matter and requirements, and accordingly using judgment/evaluation to characterize what the request is asking for) predicting, based at least in part on the analyzing: (i) a large language model of a plurality of large language models to process and to respond to the large language model request (mental process – predicting which one of a plurality of large language models should respond to the request may be performed mentally by a user analyzing the request, considering the known strengths of each available model, and using judgment/evaluation to select an appropriate model) and (ii) at least one database from which data is to be used to generate a prompt for the large language model (mental process - predicting which database should supply the data for generating the prompt may be performed mentally by a user analyzing the request, considering which available data source is pertinent, and using judgment/evaluation to select an appropriate database) Step 2A Prong 2: This judicial exception is not integrated into a practical application. receiving a large language model request (Adding insignificant extra-solution activity to the judicial exception - see MPEP 2106.05(g)) […] using one or more machine learning algorithms (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of using machine learning algorithms without significantly more) interfacing with the large language model and the at least one database to enable the large language model to process and to respond to the large language model request (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of using a large language model and database to perform an abstract idea without significantly more) a processing device operatively coupled to a memory (recited at a high-level of generality (i.e., a generic processor, computer-readable storage medium, a communication interface, a processing device and memory) such that it amounts to no more than mere instructions to apply the exception using generic computer components) Step 2B: The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. receiving a large language model request (MPEP 2106.05(d)(II) indicates that merely “Receiving or transmitting data over a network” is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer) […] using one or more machine learning algorithms (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of using machine learning algorithms without significantly more) interfacing with the large language model and the at least one database to enable the large language model to process and to respond to the large language model request (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of using a large language model and database to perform an abstract idea without significantly more) a processing device operatively coupled to a memory (recited at a high-level of generality (i.e., a generic processor, computer-readable storage medium, a communication interface, a processing device and memory) such that it amounts to no more than mere instructions to apply the exception using generic computer components) For the reasons above, Claim 1 is rejected as being directed to an abstract idea without significantly more. This rejection applies equally to dependent claims 1 - 15. The additional limitations of the dependent claims are addressed below. Regarding Claim 2: Step 2A Prong 1: See the rejection of Claim 1 above, which Claim 2 depends on. generating a vector of the large language model request (mental process - generating a vector of the request may be performed mentally or using pen and paper by a user reading the request and writing down numerical values representing its content) Step 2A Prong 2 & Step 2B: Accordingly, under Step 2A Prong 2 and Step 2B, there are no additional elements that integrate the abstract idea into practical application. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. Regarding Claim 3: Step 2A Prong 1: See the rejection of Claim 2 above, which Claim 3 depends on. executing a hash function on the vector to create a unique identifier for the large language model request (mathematical concept / mental process – executing a hash function on the vector recites a mathematical concept, as it applies a mathematical function to the numerical vector to compute an identifier value, and may also be performed mentally or using pen and paper by a user applying the hash calculation to the vector values to derive a unique identifier) Step 2A Prong 2 & Step 2B: Accordingly, under Step 2A Prong 2 and Step 2B, there are no additional elements that integrate the abstract idea into practical application. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. Regarding Claim 4: Step 2A Prong 1: See the rejection of Claim 3 above, which Claim 4 depends on. Step 2A Prong 2 & Step 2B: receiving a response to the large language model request (Adding insignificant extra-solution activity to the judicial exception - see MPEP 2106.05(g) & MPEP 2106.05(d)(II) indicates that merely “Receiving or transmitting data over a network” is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer) storing the response to the large language model request in correspondence with the unique identifier for the large language model request (Adding insignificant extra-solution activity to the judicial exception – see MPEP 2106.05(g). Further, MPEP 2106.05(d)(II) indicates that "Storing and retrieving information in memory" is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer) Accordingly, under Step 2A Prong 2 and Step 2B, this additional element does not integrate the abstract idea into practical application because it does not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 3. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. Regarding Claim 5: Step 2A Prong 1: See the rejection of Claim 1 above, which Claim 5 depends on. determining whether the large language model request matches a previous large language model request of a plurality of previous large language model requests (mental process – determining whether the request matches a previous request may be performed mentally or using pen and paper by a user observing/analyzing the request, comparing it against previously received requests, and accordingly using judgment/evaluation to decide whether it matches any of the previous requests) performing the analyzing, predicting and interfacing in response to determining that the large language model request differs from the plurality of previous large language model requests (mental process - performing the analyzing, predicting and interfacing in response to determining that the request differs from the previous requests may be performed mentally or using pen and paper by a user analyzing the request, judging/evaluating that it does not match any previous request, and accordingly proceeding to carry out the analyzing, predicting, and interfacing steps recited above) Step 2A Prong 2 & Step 2B: Accordingly, under Step 2A Prong 2 and Step 2B, there are no additional elements that integrate the abstract idea into practical application. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. Regarding Claim 6: Step 2A Prong 1: See the rejection of Claim 5 above, which Claim 6 depends on. creating a unique identifier for the large language model request (mental process – creating a unique identifier for the request may be performed mentally or using pen and paper by a user analyzing the request and assigning a corresponding identifier to it) comparing the unique identifier for the large language model request to a plurality of stored unique identifiers corresponding to respective ones of the plurality of previous large language model requests to determine whether the unique identifier for the large language model request matches a stored unique identifier of the plurality of stored unique identifiers (mental process - comparing the unique identifier to a plurality of stored unique identifiers may be performed mentally or using pen and paper by a user analyzing the identifier, comparing it against each of the stored identifiers, and accordingly using judgment/evaluation to determine whether it matches any stored identifier) Step 2A Prong 2 & Step 2B: Accordingly, under Step 2A Prong 2 and Step 2B, there are no additional elements that integrate the abstract idea into practical application. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. Regarding Claim 7: Step 2A Prong 1: See the rejection of Claim 1 above, which Claim 7 depends on. creating a unique identifier for the additional large language model request (mental process – creating a unique identifier for the request may be performed mentally or using pen and paper by a user analyzing the request and assigning a corresponding identifier to it) comparing the unique identifier for the additional large language model request to a plurality of stored unique identifiers corresponding to respective ones of a plurality of previous large language model requests to determine whether the unique identifier for the additional large language model request matches a stored unique identifier of the plurality of stored unique identifiers (mental process - comparing the unique identifier to a plurality of stored unique identifiers may be performed mentally or using pen and paper by a user analyzing the identifier, comparing it against each of the stored identifiers, and accordingly using judgment/evaluation to determine whether it matches any stored identifier) Step 2A Prong 2 & Step 2B: receiving an additional large language model request (Adding insignificant extra-solution activity to the judicial exception - see MPEP 2106.05(g) & MPEP 2106.05(d)(II) indicates that merely “Receiving or transmitting data over a network” is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer) Accordingly, under Step 2A Prong 2 and Step 2B, this additional element does not integrate the abstract idea into practical application because it does not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 1. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. Regarding Claim 8: Step 2A Prong 1: See the rejection of Claim 7 above, which Claim 8 depends on. Step 2A Prong 2 & Step 2B: retrieving a stored large language model response corresponding to the stored unique identifier in response to determining that the unique identifier for the additional large language model request matches the stored unique identifier (Adding insignificant extra-solution activity to the judicial exception – see MPEP 2106.05(g). Further, MPEP 2106.05(d)(II) indicates that "Storing and retrieving information in memory" is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer) Accordingly, under Step 2A Prong 2 and Step 2B, this additional element does not integrate the abstract idea into practical application because it does not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 7. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. Regarding Claim 9: Step 2A Prong 1: See the rejection of Claim 1 above, which Claim 9 depends on. Step 2A Prong 2 & Step 2B: the one or more machine learning algorithms comprise a neural network configured to predict a plurality of targets (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of using a neural network to perform an abstract idea without significantly more) the plurality of targets comprise the large language model and the at least one database (Field of Use – limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that the plurality of targets comprise the large language model and the at least one database does not integrate the exception into a practical application nor amount to significantly more – See MPEP 2106.05(h)) the neural network includes a plurality of parallel networks respectively corresponding to the plurality of targets (Field of Use – limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that the neural network includes a plurality of parallel networks respectively corresponding to the plurality of targets does not integrate the exception into a practical application nor amount to significantly more – See MPEP 2106.05(h)) Accordingly, under Step 2A Prong 2 and Step 2B, this additional element does not integrate the abstract idea into practical application because it does not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 1. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. Regarding Claim 10: Step 2A Prong 1: See the rejection of Claim 9 above, which Claim 10 depends on. Step 2A Prong 2 & Step 2B: wherein a first parallel network of the plurality of parallel networks corresponding to the large language model comprises a multi-class classifier and a second parallel network of the plurality of parallel networks corresponding to the at least one database comprises a multi-label classifier (Field of Use – limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that a first parallel network of the plurality of parallel networks corresponding to the large language model comprises a multi-class classifier and a second parallel network of the plurality of parallel networks corresponding to the at least one database comprises a multi-label classifier does not integrate the exception into a practical application nor amount to significantly more – See MPEP 2106.05(h)) Accordingly, under Step 2A Prong 2 and Step 2B, this additional element does not integrate the abstract idea into practical application because it does not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 9. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. Regarding Claim 11: Step 2A Prong 1: See the rejection of Claim 1 above, which Claim 11 depends on. Step 2A Prong 2 & Step 2B: wherein the one or more machine learning algorithms are trained with historical data of a plurality of large language model requests (Field of Use – limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that the one or more machine learning algorithms are trained with historical data of a plurality of large language model requests does not integrate the exception into a practical application nor amount to significantly more – See MPEP 2106.05(h)) Accordingly, under Step 2A Prong 2 and Step 2B, this additional element does not integrate the abstract idea into practical application because it does not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 1. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. Regarding Claim 12: Step 2A Prong 1: See the rejection of Claim 11 above, which Claim 12 depends on. Step 2A Prong 2 & Step 2B: wherein the historical data specifies for respective ones of the plurality of large language model requests at least one of: (i) a request vector; (ii) a domain; (iii) usefulness of a response to a corresponding request; (iv) a database used in connection with generating a large language model prompt; and (v) a large language model used to generate the response to the corresponding request (Field of Use – limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that the historical data specifies for respective ones of the plurality of large language model requests at least one of: (i) a request vector; (ii) a domain; (iii) usefulness of a response to a corresponding request; (iv) a database used in connection with generating a large language model prompt; and (v) a large language model used to generate the response to the corresponding request does not integrate the exception into a practical application nor amount to significantly more – See MPEP 2106.05(h)) Accordingly, under Step 2A Prong 2 and Step 2B, this additional element does not integrate the abstract idea into practical application because it does not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 11. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. Regarding Claim 13: Step 2A Prong 1: See the rejection of Claim 11 above, which Claim 13 depends on. updating training of the one or more machine learning algorithms based on the collected feedback data (mental process – updating the training of the machine learning algorithms based on the collected feedback data may be performed mentally or using pen and paper by a user analyzing the collected feedback data and using judgment/evaluation to adjust the algorithm accordingly) Step 2A Prong 2 & Step 2B: collecting feedback data regarding quality of a response to the large language model request (Adding insignificant extra-solution activity to the judicial exception – see MPEP 2106.05(g). Further, MPEP 2106.05(d)(II) indicates that gathering data is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer) Accordingly, under Step 2A Prong 2 and Step 2B, this additional element does not integrate the abstract idea into practical application because it does not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 11. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. Regarding Claim 14: Step 2A Prong 1: See the rejection of Claim 1 above, which Claim 14 depends on. Step 2A Prong 2 & Step 2B: wherein the at least one database comprises a vector store (Field of Use – limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that the at least one database comprises a vector store does not integrate the exception into a practical application nor amount to significantly more – See MPEP 2106.05(h)) Accordingly, under Step 2A Prong 2 and Step 2B, this additional element does not integrate the abstract idea into practical application because it does not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 1. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. Regarding Claim 15: Step 2A Prong 1: See the rejection of Claim 1 above, which Claim 15 depends on. Step 2A Prong 2 & Step 2B: wherein the interfacing comprises generating one or more application programming interface calls to at least one of query the at least one database for the data to be used to generate the prompt, send the prompt to the large language model and receive a response to the large language model request (Adding the words “apply it” (or an equivalent) with the judicial exception, or, merely uses a computer as a tool to implement the judicial exception by reciting, at a high level of generality, API calls for querying a database, transmitting a prompt to an LLM, and receiving a response, without reciting an improvement to the operation of the API, database, LLM, or computer itself. Accordingly, the limitation amounts to mere instructions to implement the abstract idea using computer technology. See MPEP § 2106.05(f)) Accordingly, under Step 2A Prong 2 and Step 2B, this additional element does not integrate the abstract idea into practical application because it does not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 1. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. Regarding Claim 16: Step 1: Claim 16 is an apparatus type claim. Therefore, Claims 16-18 fall within one of the four statutory categories (i.e., process, machine, manufacture, or composition of matter). 2A Prong 1: If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation by mathematical calculation but for the recitation of generic computer components, then it falls within the “Mathematical Concepts” grouping of abstract ideas. to analyze the large language model request […] (mental process – analyzing the large language model request may be performed mentally or using pen and paper by a user observing/reading the received request, considering its subject matter and requirements, and accordingly using judgment/evaluation to characterize what the request is asking for) to predict, based at least in part on the analyzing: (i) a large language model of a plurality of large language models to process and to respond to the large language model request (mental process – predicting which one of a plurality of large language models should respond to the request may be performed mentally by a user analyzing the request, considering the known strengths of each available model, and using judgment/evaluation to select an appropriate model) and (ii) at least one database from which data is to be used to generate a prompt for the large language model (mental process - predicting which database should supply the data for generating the prompt may be performed mentally by a user analyzing the request, considering which available data source is pertinent, and using judgment/evaluation to select an appropriate database) Step 2A Prong 2: This judicial exception is not integrated into a practical application. to receive a large language model request (Adding insignificant extra-solution activity to the judicial exception - see MPEP 2106.05(g)) […] using one or more machine learning algorithms (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of using machine learning algorithms without significantly more) to interface with the large language model and the at least one database to enable the large language model to process and to respond to the large language model request (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of using a large language model and database to perform an abstract idea without significantly more) a processing device operatively coupled to a memory (recited at a high-level of generality (i.e., a generic processor, computer-readable storage medium, a communication interface, a processing device and memory) such that it amounts to no more than mere instructions to apply the exception using generic computer components) Step 2B: The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. to receive a large language model request (MPEP 2106.05(d)(II) indicates that merely “Receiving or transmitting data over a network” is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer) […] using one or more machine learning algorithms (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of using machine learning algorithms without significantly more) to interface with the large language model and the at least one database to enable the large language model to process and to respond to the large language model request (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of using a large language model and database to perform an abstract idea without significantly more) a processing device operatively coupled to a memory (recited at a high-level of generality (i.e., a generic processor, computer-readable storage medium, a communication interface, a processing device and memory) such that it amounts to no more than mere instructions to apply the exception using generic computer components) For the reasons above, Claim 16 is rejected as being directed to an abstract idea without significantly more. This rejection applies equally to dependent claims 16 - 18. The additional limitations of the dependent claims are addressed below. Regarding Claim 17: Step 2A Prong 1: See the rejection of Claim 16 above, which Claim 17 depends on. to determine whether the large language model request matches a previous large language model request of a plurality of previous large language model requests (mental process – determining whether the request matches a previous request may be performed mentally or using pen and paper by a user observing/analyzing the request, comparing it against previously received requests, and accordingly using judgment/evaluation to decide whether it matches any of the previous requests) to perform the analyzing, predicting and interfacing in response to determining that the large language model request differs from the plurality of previous large language model requests (mental process - performing the analyzing, predicting and interfacing in response to determining that the request differs from the previous requests may be performed mentally or using pen and paper by a user analyzing the request, judging/evaluating that it does not match any previous request, and accordingly proceeding to carry out the analyzing, predicting, and interfacing steps recited above) Step 2A Prong 2 & Step 2B: Accordingly, under Step 2A Prong 2 and Step 2B, there are no additional elements that integrate the abstract idea into practical application. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. Regarding Claim 18: Step 2A Prong 1: See the rejection of Claim 17 above, which Claim 18 depends on. to creating a unique identifier for the large language model request (mental process – creating a unique identifier for the request may be performed mentally or using pen and paper by a user analyzing the request and assigning a corresponding identifier to it) to compare the unique identifier for the large language model request to a plurality of stored unique identifiers corresponding to respective ones of the plurality of previous large language model requests to determine whether the unique identifier for the large language model request matches a stored unique identifier of the plurality of stored unique identifiers (mental process - comparing the unique identifier to a plurality of stored unique identifiers may be performed mentally or using pen and paper by a user analyzing the identifier, comparing it against each of the stored identifiers, and accordingly using judgment/evaluation to determine whether it matches any stored identifier) Step 2A Prong 2 & Step 2B: Accordingly, under Step 2A Prong 2 and Step 2B, there are no additional elements that integrate the abstract idea into practical application. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. Regarding Claim 19: Step 1: Claim 19 is an article of manufacture type claim. Therefore, Claims 19-20 fall within one of the four statutory categories (i.e., process, machine, manufacture, or composition of matter). 2A Prong 1: If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation by mathematical calculation but for the recitation of generic computer components, then it falls within the “Mathematical Concepts” grouping of abstract ideas. analyzing the large language model request […] (mental process – analyzing the large language model request may be performed mentally or using pen and paper by a user observing/reading the received request, considering its subject matter and requirements, and accordingly using judgment/evaluation to characterize what the request is asking for) predicting, based at least in part on the analyzing: (i) a large language model of a plurality of large language models to process and to respond to the large language model request (mental process – predicting which one of a plurality of large language models should respond to the request may be performed mentally by a user analyzing the request, considering the known strengths of each available model, and using judgment/evaluation to select an appropriate model) and (ii) at least one database from which data is to be used to generate a prompt for the large language model (mental process - predicting which database should supply the data for generating the prompt may be performed mentally by a user analyzing the request, considering which available data source is pertinent, and using judgment/evaluation to select an appropriate database) Step 2A Prong 2: This judicial exception is not integrated into a practical application. receiving a large language model request (Adding insignificant extra-solution activity to the judicial exception - see MPEP 2106.05(g)) […] using one or more machine learning algorithms (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of using machine learning algorithms without significantly more) interfacing with the large language model and the at least one database to enable the large language model to process and to respond to the large language model request (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of using a large language model and database to perform an abstract idea without significantly more) processing device (recited at a high-level of generality (i.e., a generic processor, computer-readable storage medium, a communication interface, a processing device and memory) such that it amounts to no more than mere instructions to apply the exception using generic computer components) Step 2B: The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. receiving a large language model request (MPEP 2106.05(d)(II) indicates that merely “Receiving or transmitting data over a network” is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer) […] using one or more machine learning algorithms (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of using machine learning algorithms without significantly more) interfacing with the large language model and the at least one database to enable the large language model to process and to respond to the large language model request (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of using a large language model and database to perform an abstract idea without significantly more) processing device (recited at a high-level of generality (i.e., a generic processor, computer-readable storage medium, a communication interface, a processing device and memory) such that it amounts to no more than mere instructions to apply the exception using generic computer components) For the reasons above, Claim 19 is rejected as being directed to an abstract idea without significantly more. This rejection applies equally to dependent claim 20. The additional limitations of the dependent claim are addressed below. Regarding Claim 20: Step 2A Prong 1: See the rejection of Claim 19 above, which Claim 20 depends on. determining whether the large language model request matches a previous large language model request of a plurality of previous large language model requests (mental process – determining whether the request matches a previous request may be performed mentally or using pen and paper by a user observing/analyzing the request, comparing it against previously received requests, and accordingly using judgment/evaluation to decide whether it matches any of the previous requests) performing the analyzing, predicting and interfacing in response to determining that the large language model request differs from the plurality of previous large language model requests (mental process - performing the analyzing, predicting and interfacing in response to determining that the request differs from the previous requests may be performed mentally or using pen and paper by a user analyzing the request, judging/evaluating that it does not match any previous request, and accordingly proceeding to carry out the analyzing, predicting, and interfacing steps recited above) Step 2A Prong 2 & Step 2B: Accordingly, under Step 2A Prong 2 and Step 2B, there are no additional elements that integrate the abstract idea into practical application. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-2, 5, 11-17 and 19-20 are rejected under 35 U.S.C. 103 as being unpatentable over Chen et al. (hereinafter Chen, a non-patent literature reference titled “FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance”) in view of Jiang et al. (hereinafter Jiang) (CN 117573834). Regarding Claim 1, Chen teaches: receiving a large language model request (Chen, Page 3 – Section 2, “We consider answering queries via the LLM market, which comprises K different LLM APIs, denoted by {fi(·)}K i=1. Each fi(·) : P → A is a function that, given a prompt p from the prompt space P, generates an answer from the answer distribution A”, & Page 5 – Section 3, “The LLM router selects mLLM APIs to include in the list. Let L ∈ [K]m denote the indexes of the m APIs selected by the router. Given a new query, it iteratively invokes the ith API in the list to obtain an answer fLi (q)”, thus receiving a large language model request is disclosed, because Chen teaches receiving a new query for processing through an LLM marketplace comprising multiple LLM APIs, where the LLM router processes the new query by invoking one or more selected LLM APIs to obtain an answer. Chen’s new query corresponds to the large language model request, and the LLM APIs correspond to large language models configured to process the request and generate an answer) predicting, based at least in part on the analyzing: (i) a large language model of a plurality of large language models to process and to respond to the large language model request (Chen, Page 2 – Section 1, “LLM cascade focuses on how to adaptively choose which LLM APIs to use for different queries. To illustrate the potential of these ideas, we implement and evaluate a simple version of FrugalGPT using LLM cascade. On each dataset and task, FrugalGPT learns how to adaptively triage different queries in the dataset to different combinations of LLMs, including ChatGPT [Cha], GPT-3 [BMR+20] and GPT-4”, & Page 5 – Section 3, “The LLM router selects mLLM APIs to include in the list. Let L ∈ [K]m denote the indexes of the m APIs selected by the router. Given a new query, it iteratively invokes the ith API in the list to obtain an answer fLi (q)”, thus predicting, based at least in part on the analyzing: (i) a large language model of a plurality of large language models to process and to respond to the large language model request is disclosed, because Chen teaches adaptively choosing which LLM APIs to use for different queries and learning to triage different queries to different combinations of LLMs, including ChatGPT, GPT-3, and GPT-4. Chen further teaches that an LLM router selects LLM APIs and invokes a selected API for a new query to obtain an answer. Chen’s new query corresponds to the large language model request, the available LLM APIs correspond to the plurality of large language models, and the selected LLM API corresponds to the large language model predicted to process and respond to the large language model request) Chen does not explicitly teach analyzing […the large language model request…] using one or more machine learning algorithms, and (ii) at least one database from which data is to be used to generate a prompt for […the large language model…], interfacing with […the large language model…] and the at least one database to enable […the large language model…] to process and to respond to […the large language model request…], and wherein the steps of the method are executed by a processing device operatively coupled to a memory. However, Jiang teaches: analyzing […the large language model request…] using one or more machine learning algorithms (Jiang, Page 11, “according to the type of the dialog content, taking the dialog content as a training set, establishing a deep learning network through deep learning, respectively establishing a picture description model, a voice conversion model, a text multi-round dialog conversion model, a text semantic understanding model, a text subject identification model, a text entity identification model and a text language identification model”, & Page 12, “The solution realizes the comprehensive understanding and analysis of the user conversation content by establishing multiple professional models and connectors; Through the deep learning network, it can convert the dialog content of picture, voice, text and so on into text form, at the same time, identify the intention, subject and entity information of the user, and language identification”, thus analyzing […the large language model request…] using one or more machine learning algorithms is disclosed, because Jiang teaches analyzing user conversation content using a deep learning network having multiple models, including a text semantic understanding model, text subject identification model, text entity identification model, and text language identification model. Jiang further teaches that the deep learning network analyzes the conversation content to identify the user’s intention, subject, entity information, and language. Jiang’s user conversation content corresponds to […the large language model request…], and the deep learning network and associated models correspond to the one or more machine learning algorithms used to analyze […the large language model request…]) and (ii) at least one database from which data is to be used to generate a prompt for […the large language model…] (Jiang, Page 13, “using the keyword matching and semantic recall to find the related knowledge from the enterprise knowledge base; S305: sorting the recalled knowledge, training the sorting model or transferring the large language model LLM to score the correlation degree, and screening out several pieces of knowledge with the highest score; constructing LLM prompt based on knowledge with highest score, user context, pre processed language, entity, subject and single-cycle writing result”, & “storing the vector in the vector database, wherein the vector database uses the data structure to match with the index acceleration similarity; in each session, according to the problem and context of the user, using the semantic embedding model to convert the user problem into the vector expression; finding the most relevant example of the user problem by calculating the similarity between the example vector and the user problem vector; S3053: according to the matched most relevant example, extracting the answer part in the example as the prompt of the large language model LLM”, thus and (ii) at least one database from which data is to be used to generate a prompt for […the large language model…] is disclosed, because Jiang teaches retrieving related knowledge from an enterprise knowledge base and screening the retrieved knowledge to identify knowledge having the highest score, which is then used to construct an LLM prompt. Jiang further teaches storing vectors in a vector database, identifying the most relevant example based on similarity between the user problem vector and stored example vectors, and extracting the answer portion of the identified example as the prompt for the large language model LLM. Jiang’s enterprise knowledge base and vector database correspond to the at least one database, and the retrieved knowledge and identified example correspond to the data used to generate a prompt for […the large language model…]) interfacing with […the large language model…] and the at least one database to enable […the large language model…] to process and to respond to […the large language model request…] (Jiang, Page 13, “using the keyword matching and semantic recall to find the related knowledge from the enterprise knowledge base; secondly, sorting the recalled knowledge, training the sorting model or transferring the large language model LLM to score the correlation degree, screening out several pieces of knowledge with the highest score; constructing LLM prompt based on knowledge with highest score, user context, pre-processed language, entity, subject and single-cycle writing result; the large language model LLM comprises large language models such as ChatGPT, Wenxi dialect or star fire; finally, inputting the prompt to the large language model LLM to obtain the answer, and replying to the user”, thus interfacing with […the large language model…] and the at least one database to enable […the large language model…] to process and to respond to […the large language model request…] is disclosed, because Jiang teaches retrieving related knowledge from an enterprise knowledge base, sorting and screening the retrieved knowledge, constructing an LLM prompt based on the selected knowledge, and inputting the prompt to the large language model LLM to obtain an answer and reply to the user. Jiang’s enterprise knowledge base corresponds to the at least one database, and the retrieval of knowledge from the database and submission of the resulting prompt to the LLM corresponds to interfacing with […the large language model…] and the at least one database to enable […the large language model…] to process and to respond to […the large language model request…]) wherein the steps of the method are executed by a processing device operatively coupled to a memory (Jiang, Page 6, “The invention provides a multi-robot dialog system for software-oriented service platform, comprising: a conversation management module for obtaining the conversation content input by the user, performing the conversation management including the complexity and the problem type, judging the corresponding relation between the complexity and the triggering flow robot and the problem type and the similar question robot, and confirming the triggering flow robot or the similar question robot according to the corresponding relation; a pre-processing module for establishing a picture description model, a voice conversion model, a text multi-turn conversation conversion model, a text semantic understanding model, a text subject identification model, a text entity identification model and a text language identification model by taking the conversation content as a training set; the dialog content is input to the model for picture description, voice conversion, core information acquisition, multi-round dialog conversion, text semantic understanding, text subject identification, pre-processing the text entity identification and the text language identification”, & “a trigger judging module for judging whether the corresponding service flow is triggered according to the user intention and triggering the corresponding service flow”, & Page 12, “judging whether the read mapping relation contains new mapping relation, if not, writing the new mapping relation in the memory; if so, combining the new mapping relation and the read mapping relation, and updating the memory full index”, thus wherein the steps of the method are executed by a processing device operatively coupled to a memory is disclosed, because Jiang teaches a computerized multi-robot dialog system having processing modules that execute operations including obtaining and processing conversation content, establishing and applying models, and triggering corresponding service flows. Jiang further teaches that the system reads, writes, combines, and updates information stored in memory. Jiang’s computerized system and processing modules correspond to the processing device, and Jiang’s memory corresponds to the memory operatively coupled to the processing device) It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to combine Chen’s teaching of receiving a large language model request and selecting, based on the request, one of a plurality of large language models to process and respond to the request, with Jiang’s teaching of analyzing the request using machine learning, retrieving relevant data from at least one database, and using the retrieved data to generate a prompt for the selected large language model. Therefore, a POSITA would have been motivated to incorporate Jiang’s request analysis, database retrieval, and prompt generation teachings into Chen’s large language model selection system so that, after selecting an appropriate large language model for the received large language model request, relevant data could be retrieved from at least one database and used to generate a prompt for the selected large language model, thereby enabling the selected large language model to process and respond to the large language model request with more accurate and informative information (Jiang, Page 2, “RAG is a natural language processing model combined with two methods of retrieval and generation for generating a natural language text reply. It provides more accurate and informative answers by first searching the retrieved related text and then generating a reply”, & Page 14, “The solution extracts the related knowledge from the enterprise knowledge base by combining the preset screening condition, the keyword matching, the semantic battery recall and the sequencing model, and uses the large language model LLM to generate the proper answer to provide the accurate and personalized service for the user; The invention can improve the answer quality and efficiency of the intelligent customer service system”) Regarding Claim 2, Chen combined with Jiang teaches all the limitations of claim 1 as cited above and Jiang further teaches: generating a vector of […the large language model request…] (Jiang, Page 6, “in each session, according to the problem and context of the user, using the semantic embedding model to convert the user problem into the vector expression; finding the most relevant example of the user problem by calculating the similarity between the example vector and the user problem vector”, thus generating a vector of […the large language model request…] is disclosed, because Jiang teaches using a semantic embedding model to convert the user problem into a vector expression. Jiang’s user problem corresponds to […the large language model request…], and converting the user problem into the vector expression corresponds to generating a vector) Regarding Claim 5, Chen combined with Jiang teaches all the limitations of claim 1 as cited above and Chen further teaches: determining whether the large language model request matches a previous large language model request of a plurality of previous large language model requests (Chen, Page 5 – Section 3, “To process a new query, we first verify if a similar query has been previously answered. If so, the response is retrieved from the cache. An LLM API is invoked only if no similar query is discovered in the cache. The completion cache provides substantial cost savings when similar queries are frequently posed”, thus determining whether the large language model request matches a previous large language model request of a plurality of previous large language model requests is disclosed, because Chen teaches checking a new query against previously answered queries to determine whether a similar query is already present in the cache. Chen’s new query corresponds to the large language model request, and the previously answered similar queries correspond to the plurality of previous large language model requests) performing […the analyzing, predicting and interfacing…] in response to determining that the large language model request differs from the plurality of previous large language model requests (Chen, Page 5 – Section 3, “An LLM API is invoked only if no similar query is discovered in the cache. The completion cache provides substantial cost savings when similar queries are frequently posed”, & Page 5 – Section 3, “The LLM router selects mLLM APIs to include in the list. Let L ∈ [K]m denote the indexes of the m APIs selected by the router. Given a new query, it iteratively invokes the ith API in the list to obtain an answer fLi (q)”, thus performing […the analyzing, predicting and interfacing…] in response to determining that the large language model request differs from the plurality of previous large language model requests is disclosed, because Chen teaches that when no similar previously answered query is found in the cache, the system proceeds with processing the new query, including selecting an LLM API and invoking the selected LLM API to obtain an answer) Regarding Claim 11, Chen combined with Jiang teaches all the limitations of claim 1 as cited above and Chen further teaches: wherein […the one or more machine learning algorithms…] are trained with historical data of a plurality of large language model requests (Chen, Page 6 - Section 4, “Each dataset is randomly split into a training set to learn the LLM cascade and a test set for evaluation”, & Page 2- Section 1, “On each dataset and task, FrugalGPT learns how to adaptively triage different queries in the dataset to different combinations of LLMs, including ChatGPT, GPT-3 and GPT-4”, thus […the one or more machine learning algorithms…] are trained with historical data of a plurality of large language model requests is disclosed, because Chen teaches training the LLM cascade using a training set containing a plurality of previously collected queries, where the learned cascade determines how the different queries are directed to different LLMs. The previously collected queries used to train the LLM cascade correspond to historical data of a plurality of large language model requests) Regarding Claim 12, Chen combined with Jiang teaches all the limitations of claim 11 as cited above and Chen further teaches: wherein the historical data specifies for respective ones of the plurality of large language model requests at least one of: (i) a request vector; (ii) a domain; (iii) usefulness of a response to a corresponding request; (iv) a database used in connection with generating a large language model prompt; and (v) a large language model used to generate the response to the corresponding request (Chen, Page 5 - Section 3, “The scoring function can be obtained by training a simple regression model that learns whether a generation is correct from the query and a generated answer”, & “the objective measures the quality of the generation \(f_{L_z}(q)\) for a query \(q\) compared to the true answer \(a\)”, thus wherein the historical data specifies for respective ones of the plurality of large language model requests at least one of: […] (iii) usefulness of a response to a corresponding request is disclosed, because Chen teaches training data associating a query with a generated answer and information indicating whether the generated answer is correct or its quality relative to the true answer. The correctness or quality information associated with the generated answer corresponds to usefulness of a response to the corresponding large language model request) Regarding Claim 13, Chen combined with Jiang teaches all the limitations of claim 11 as cited above and Jiang further teaches: collecting feedback data regarding quality of a response to […the large language model request…] (Jiang, Page 6, “inputting the prompt to the large language model LLM to obtain the answer, and replying to the user; evaluating the result of the reply of the large language model LLM, judging whether it is correct, whether it is necessary to apologize and whether it is necessary to transfer to manual customer service and so on”, thus collecting feedback data regarding quality of a response to […the large language model request…] is disclosed, because Jiang teaches evaluating an LLM generated response after the response is produced and determining whether the response is correct, wherein the result of that evaluation constitutes feedback data regarding the quality of the response) Chen further teaches: updating training of […the one or more machine learning algorithms…] based on the collected feedback data (Chen, Page 5 - Section 3, “The scoring function can be obtained by training a simple regression model that learns whether a generation is correct from the query and a generated answer”, thus updating training of […the one or more machine learning algorithms…] based on the collected feedback data is disclosed, because Chen teaches training a machine learning model using correctness information associated with a generated response) Regarding Claim 14, Chen combined with Jiang teaches all the limitations of claim 1 as cited above and Jiang further teaches: wherein the at least one database comprises a vector store (Jiang, Page 6, “storing the vector in the vector database, wherein the vector database uses the data structure to match with the index acceleration similarity; in each session, according to the problem and context of the user, using the semantic embedding model to convert the user problem into the vector expression; finding the most relevant example of the user problem by calculating the similarity between the example vector and the user problem vector”, thus wherein the at least one database comprises a vector store is disclosed, because Jiang teaches a vector database that stores vector representations and retrieves relevant information by performing similarity matching between a stored vector and a vector representation of the user problem, wherein such a vector database corresponds to a vector store) Regarding Claim 15, Chen combined with Jiang teaches all the limitations of claim 1 as cited above and Chen further teaches: wherein […the interfacing…] comprises generating one or more application programming interface calls to at least one of query […the at least one database…] for the data to be used to generate the prompt, send the prompt to the large language model and receive a response to the large language model request (Chen, Page 3 - Section 2, “Each \(f_i(\cdot): P \rightarrow A\) is a function that, given a prompt \(p\) from the prompt space \(P\), generates an answer from the answer distribution \(A\). Note that to use LLM APIs, one has to convert each query \(q\) to some corresponding prompt first”, & Page 5 - Section 3, “Given a new query, it iteratively invokes the ith API in the list to obtain an answer”, thus wherein […the interfacing…] comprises generating one or more application programming interface calls to at least one of […] send the prompt to the large language model and receive a response to the large language model request is disclosed, because Chen teaches converting a query into a corresponding prompt, invoking an LLM API with the query/prompt, and obtaining an answer generated by the invoked LLM API in response to the request) Regarding Claim 16, Chen teaches: to receive a large language model request (Chen, Page 3 – Section 2, “We consider answering queries via the LLM market, which comprises K different LLM APIs, denoted by {fi(·)}K i=1. Each fi(·) : P → A is a function that, given a prompt p from the prompt space P, generates an answer from the answer distribution A”, & Page 5 – Section 3, “The LLM router selects mLLM APIs to include in the list. Let L ∈ [K]m denote the indexes of the m APIs selected by the router. Given a new query, it iteratively invokes the ith API in the list to obtain an answer fLi (q)”, thus receiving a large language model request is disclosed, because Chen teaches receiving a new query for processing through an LLM marketplace comprising multiple LLM APIs, where the LLM router processes the new query by invoking one or more selected LLM APIs to obtain an answer. Chen’s new query corresponds to the large language model request, and the LLM APIs correspond to large language models configured to process the request and generate an answer) to predict, based at least in part on the analyzing: (i) a large language model of a plurality of large language models to process and to respond to the large language model request (Chen, Page 2 – Section 1, “LLM cascade focuses on how to adaptively choose which LLM APIs to use for different queries. To illustrate the potential of these ideas, we implement and evaluate a simple version of FrugalGPT using LLM cascade. On each dataset and task, FrugalGPT learns how to adaptively triage different queries in the dataset to different combinations of LLMs, including ChatGPT [Cha], GPT-3 [BMR+20] and GPT-4”, & Page 5 – Section 3, “The LLM router selects mLLM APIs to include in the list. Let L ∈ [K]m denote the indexes of the m APIs selected by the router. Given a new query, it iteratively invokes the ith API in the list to obtain an answer fLi (q)”, thus predicting, based at least in part on the analyzing: (i) a large language model of a plurality of large language models to process and to respond to the large language model request is disclosed, because Chen teaches adaptively choosing which LLM APIs to use for different queries and learning to triage different queries to different combinations of LLMs, including ChatGPT, GPT-3, and GPT-4. Chen further teaches that an LLM router selects LLM APIs and invokes a selected API for a new query to obtain an answer. Chen’s new query corresponds to the large language model request, the available LLM APIs correspond to the plurality of large language models, and the selected LLM API corresponds to the large language model predicted to process and respond to the large language model request) Chen does not explicitly teach to analyze […the large language model request…] using one or more machine learning algorithms, and (ii) at least one database from which data is to be used to generate a prompt for […the large language model…], and to interface with […the large language model…] and the at least one database to enable […the large language model…] to process and to respond to […the large language model request…]. However, Jiang teaches: to analyze […the large language model request…] using one or more machine learning algorithms (Jiang, Page 11, “according to the type of the dialog content, taking the dialog content as a training set, establishing a deep learning network through deep learning, respectively establishing a picture description model, a voice conversion model, a text multi-round dialog conversion model, a text semantic understanding model, a text subject identification model, a text entity identification model and a text language identification model”, & Page 12, “The solution realizes the comprehensive understanding and analysis of the user conversation content by establishing multiple professional models and connectors; Through the deep learning network, it can convert the dialog content of picture, voice, text and so on into text form, at the same time, identify the intention, subject and entity information of the user, and language identification”, thus analyzing […the large language model request…] using one or more machine learning algorithms is disclosed, because Jiang teaches analyzing user conversation content using a deep learning network having multiple models, including a text semantic understanding model, text subject identification model, text entity identification model, and text language identification model. Jiang further teaches that the deep learning network analyzes the conversation content to identify the user’s intention, subject, entity information, and language. Jiang’s user conversation content corresponds to […the large language model request…], and the deep learning network and associated models correspond to the one or more machine learning algorithms used to analyze […the large language model request…]) and (ii) at least one database from which data is to be used to generate a prompt for […the large language model…] (Jiang, Page 13, “using the keyword matching and semantic recall to find the related knowledge from the enterprise knowledge base; S305: sorting the recalled knowledge, training the sorting model or transferring the large language model LLM to score the correlation degree, and screening out several pieces of knowledge with the highest score; constructing LLM prompt based on knowledge with highest score, user context, pre processed language, entity, subject and single-cycle writing result”, & “storing the vector in the vector database, wherein the vector database uses the data structure to match with the index acceleration similarity; in each session, according to the problem and context of the user, using the semantic embedding model to convert the user problem into the vector expression; finding the most relevant example of the user problem by calculating the similarity between the example vector and the user problem vector; S3053: according to the matched most relevant example, extracting the answer part in the example as the prompt of the large language model LLM”, thus and (ii) at least one database from which data is to be used to generate a prompt for […the large language model…] is disclosed, because Jiang teaches retrieving related knowledge from an enterprise knowledge base and screening the retrieved knowledge to identify knowledge having the highest score, which is then used to construct an LLM prompt. Jiang further teaches storing vectors in a vector database, identifying the most relevant example based on similarity between the user problem vector and stored example vectors, and extracting the answer portion of the identified example as the prompt for the large language model LLM. Jiang’s enterprise knowledge base and vector database correspond to the at least one database, and the retrieved knowledge and identified example correspond to the data used to generate a prompt for […the large language model…]) to interface with […the large language model…] and the at least one database to enable […the large language model…] to process and to respond to […the large language model request…] (Jiang, Page 13, “using the keyword matching and semantic recall to find the related knowledge from the enterprise knowledge base; secondly, sorting the recalled knowledge, training the sorting model or transferring the large language model LLM to score the correlation degree, screening out several pieces of knowledge with the highest score; constructing LLM prompt based on knowledge with highest score, user context, pre-processed language, entity, subject and single-cycle writing result; the large language model LLM comprises large language models such as ChatGPT, Wenxi dialect or star fire; finally, inputting the prompt to the large language model LLM to obtain the answer, and replying to the user”, thus interfacing with […the large language model…] and the at least one database to enable […the large language model…] to process and to respond to […the large language model request…] is disclosed, because Jiang teaches retrieving related knowledge from an enterprise knowledge base, sorting and screening the retrieved knowledge, constructing an LLM prompt based on the selected knowledge, and inputting the prompt to the large language model LLM to obtain an answer and reply to the user. Jiang’s enterprise knowledge base corresponds to the at least one database, and the retrieval of knowledge from the database and submission of the resulting prompt to the LLM corresponds to interfacing with […the large language model…] and the at least one database to enable […the large language model…] to process and to respond to […the large language model request…]) It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to combine Chen’s teaching of receiving a large language model request and selecting, based on the request, one of a plurality of large language models to process and respond to the request, with Jiang’s teaching of analyzing the request using machine learning, retrieving relevant data from at least one database, and using the retrieved data to generate a prompt for the selected large language model. Therefore, a POSITA would have been motivated to incorporate Jiang’s request analysis, database retrieval, and prompt generation teachings into Chen’s large language model selection system so that, after selecting an appropriate large language model for the received large language model request, relevant data could be retrieved from at least one database and used to generate a prompt for the selected large language model, thereby enabling the selected large language model to process and respond to the large language model request with more accurate and informative information (Jiang, Page 2, “RAG is a natural language processing model combined with two methods of retrieval and generation for generating a natural language text reply. It provides more accurate and informative answers by first searching the retrieved related text and then generating a reply”, & Page 14, “The solution extracts the related knowledge from the enterprise knowledge base by combining the preset screening condition, the keyword matching, the semantic battery recall and the sequencing model, and uses the large language model LLM to generate the proper answer to provide the accurate and personalized service for the user; The invention can improve the answer quality and efficiency of the intelligent customer service system”) Regarding Claim 17, Chen combined with Jiang teaches all the limitations of claim 16 as cited above and Chen further teaches: to determine whether the large language model request matches a previous large language model request of a plurality of previous large language model requests (Chen, Page 5 – Section 3, “To process a new query, we first verify if a similar query has been previously answered. If so, the response is retrieved from the cache. An LLM API is invoked only if no similar query is discovered in the cache. The completion cache provides substantial cost savings when similar queries are frequently posed”, thus determining whether the large language model request matches a previous large language model request of a plurality of previous large language model requests is disclosed, because Chen teaches checking a new query against previously answered queries to determine whether a similar query is already present in the cache. Chen’s new query corresponds to the large language model request, and the previously answered similar queries correspond to the plurality of previous large language model requests) to perform […the analyzing, predicting and interfacing…] in response to determining that the large language model request differs from the plurality of previous large language model requests (Chen, Page 5 – Section 3, “An LLM API is invoked only if no similar query is discovered in the cache. The completion cache provides substantial cost savings when similar queries are frequently posed”, & Page 5 – Section 3, “The LLM router selects mLLM APIs to include in the list. Let L ∈ [K]m denote the indexes of the m APIs selected by the router. Given a new query, it iteratively invokes the ith API in the list to obtain an answer fLi (q)”, thus performing […the analyzing, predicting and interfacing…] in response to determining that the large language model request differs from the plurality of previous large language model requests is disclosed, because Chen teaches that when no similar previously answered query is found in the cache, the system proceeds with processing the new query, including selecting an LLM API and invoking the selected LLM API to obtain an answer) Regarding Claim 19, Chen teaches: receiving a large language model request (Chen, Page 3 – Section 2, “We consider answering queries via the LLM market, which comprises K different LLM APIs, denoted by {fi(·)}K i=1. Each fi(·) : P → A is a function that, given a prompt p from the prompt space P, generates an answer from the answer distribution A”, & Page 5 – Section 3, “The LLM router selects mLLM APIs to include in the list. Let L ∈ [K]m denote the indexes of the m APIs selected by the router. Given a new query, it iteratively invokes the ith API in the list to obtain an answer fLi (q)”, thus receiving a large language model request is disclosed, because Chen teaches receiving a new query for processing through an LLM marketplace comprising multiple LLM APIs, where the LLM router processes the new query by invoking one or more selected LLM APIs to obtain an answer. Chen’s new query corresponds to the large language model request, and the LLM APIs correspond to large language models configured to process the request and generate an answer) predicting, based at least in part on the analyzing: (i) a large language model of a plurality of large language models to process and to respond to the large language model request (Chen, Page 2 – Section 1, “LLM cascade focuses on how to adaptively choose which LLM APIs to use for different queries. To illustrate the potential of these ideas, we implement and evaluate a simple version of FrugalGPT using LLM cascade. On each dataset and task, FrugalGPT learns how to adaptively triage different queries in the dataset to different combinations of LLMs, including ChatGPT [Cha], GPT-3 [BMR+20] and GPT-4”, & Page 5 – Section 3, “The LLM router selects mLLM APIs to include in the list. Let L ∈ [K]m denote the indexes of the m APIs selected by the router. Given a new query, it iteratively invokes the ith API in the list to obtain an answer fLi (q)”, thus predicting, based at least in part on the analyzing: (i) a large language model of a plurality of large language models to process and to respond to the large language model request is disclosed, because Chen teaches adaptively choosing which LLM APIs to use for different queries and learning to triage different queries to different combinations of LLMs, including ChatGPT, GPT-3, and GPT-4. Chen further teaches that an LLM router selects LLM APIs and invokes a selected API for a new query to obtain an answer. Chen’s new query corresponds to the large language model request, the available LLM APIs correspond to the plurality of large language models, and the selected LLM API corresponds to the large language model predicted to process and respond to the large language model request) Chen does not explicitly teach analyzing […the large language model request…] using one or more machine learning algorithms, and (ii) at least one database from which data is to be used to generate a prompt for […the large language model…], and interfacing with […the large language model…] and the at least one database to enable […the large language model…] to process and to respond to […the large language model request…] However, Jiang teaches: analyzing […the large language model request…] using one or more machine learning algorithms (Jiang, Page 11, “according to the type of the dialog content, taking the dialog content as a training set, establishing a deep learning network through deep learning, respectively establishing a picture description model, a voice conversion model, a text multi-round dialog conversion model, a text semantic understanding model, a text subject identification model, a text entity identification model and a text language identification model”, & Page 12, “The solution realizes the comprehensive understanding and analysis of the user conversation content by establishing multiple professional models and connectors; Through the deep learning network, it can convert the dialog content of picture, voice, text and so on into text form, at the same time, identify the intention, subject and entity information of the user, and language identification”, thus analyzing […the large language model request…] using one or more machine learning algorithms is disclosed, because Jiang teaches analyzing user conversation content using a deep learning network having multiple models, including a text semantic understanding model, text subject identification model, text entity identification model, and text language identification model. Jiang further teaches that the deep learning network analyzes the conversation content to identify the user’s intention, subject, entity information, and language. Jiang’s user conversation content corresponds to […the large language model request…], and the deep learning network and associated models correspond to the one or more machine learning algorithms used to analyze […the large language model request…]) and (ii) at least one database from which data is to be used to generate a prompt for […the large language model…] (Jiang, Page 13, “using the keyword matching and semantic recall to find the related knowledge from the enterprise knowledge base; S305: sorting the recalled knowledge, training the sorting model or transferring the large language model LLM to score the correlation degree, and screening out several pieces of knowledge with the highest score; constructing LLM prompt based on knowledge with highest score, user context, pre processed language, entity, subject and single-cycle writing result”, & “storing the vector in the vector database, wherein the vector database uses the data structure to match with the index acceleration similarity; in each session, according to the problem and context of the user, using the semantic embedding model to convert the user problem into the vector expression; finding the most relevant example of the user problem by calculating the similarity between the example vector and the user problem vector; S3053: according to the matched most relevant example, extracting the answer part in the example as the prompt of the large language model LLM”, thus and (ii) at least one database from which data is to be used to generate a prompt for […the large language model…] is disclosed, because Jiang teaches retrieving related knowledge from an enterprise knowledge base and screening the retrieved knowledge to identify knowledge having the highest score, which is then used to construct an LLM prompt. Jiang further teaches storing vectors in a vector database, identifying the most relevant example based on similarity between the user problem vector and stored example vectors, and extracting the answer portion of the identified example as the prompt for the large language model LLM. Jiang’s enterprise knowledge base and vector database correspond to the at least one database, and the retrieved knowledge and identified example correspond to the data used to generate a prompt for […the large language model…]) interfacing with […the large language model…] and the at least one database to enable […the large language model…] to process and to respond to […the large language model request…] (Jiang, Page 13, “using the keyword matching and semantic recall to find the related knowledge from the enterprise knowledge base; secondly, sorting the recalled knowledge, training the sorting model or transferring the large language model LLM to score the correlation degree, screening out several pieces of knowledge with the highest score; constructing LLM prompt based on knowledge with highest score, user context, pre-processed language, entity, subject and single-cycle writing result; the large language model LLM comprises large language models such as ChatGPT, Wenxi dialect or star fire; finally, inputting the prompt to the large language model LLM to obtain the answer, and replying to the user”, thus interfacing with […the large language model…] and the at least one database to enable […the large language model…] to process and to respond to […the large language model request…] is disclosed, because Jiang teaches retrieving related knowledge from an enterprise knowledge base, sorting and screening the retrieved knowledge, constructing an LLM prompt based on the selected knowledge, and inputting the prompt to the large language model LLM to obtain an answer and reply to the user. Jiang’s enterprise knowledge base corresponds to the at least one database, and the retrieval of knowledge from the database and submission of the resulting prompt to the LLM corresponds to interfacing with […the large language model…] and the at least one database to enable […the large language model…] to process and to respond to […the large language model request…]) It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to combine Chen’s teaching of receiving a large language model request and selecting, based on the request, one of a plurality of large language models to process and respond to the request, with Jiang’s teaching of analyzing the request using machine learning, retrieving relevant data from at least one database, and using the retrieved data to generate a prompt for the selected large language model. Therefore, a POSITA would have been motivated to incorporate Jiang’s request analysis, database retrieval, and prompt generation teachings into Chen’s large language model selection system so that, after selecting an appropriate large language model for the received large language model request, relevant data could be retrieved from at least one database and used to generate a prompt for the selected large language model, thereby enabling the selected large language model to process and respond to the large language model request with more accurate and informative information (Jiang, Page 2, “RAG is a natural language processing model combined with two methods of retrieval and generation for generating a natural language text reply. It provides more accurate and informative answers by first searching the retrieved related text and then generating a reply”, & Page 14, “The solution extracts the related knowledge from the enterprise knowledge base by combining the preset screening condition, the keyword matching, the semantic battery recall and the sequencing model, and uses the large language model LLM to generate the proper answer to provide the accurate and personalized service for the user; The invention can improve the answer quality and efficiency of the intelligent customer service system”) Regarding Claim 20, Chen combined with Jiang teaches all the limitations of claim 19 as cited above and Chen further teaches: determining whether the large language model request matches a previous large language model request of a plurality of previous large language model requests (Chen, Page 5 – Section 3, “To process a new query, we first verify if a similar query has been previously answered. If so, the response is retrieved from the cache. An LLM API is invoked only if no similar query is discovered in the cache. The completion cache provides substantial cost savings when similar queries are frequently posed”, thus determining whether the large language model request matches a previous large language model request of a plurality of previous large language model requests is disclosed, because Chen teaches checking a new query against previously answered queries to determine whether a similar query is already present in the cache. Chen’s new query corresponds to the large language model request, and the previously answered similar queries correspond to the plurality of previous large language model requests) performing […the analyzing, predicting and interfacing…] in response to determining that the large language model request differs from the plurality of previous large language model requests (Chen, Page 5 – Section 3, “An LLM API is invoked only if no similar query is discovered in the cache. The completion cache provides substantial cost savings when similar queries are frequently posed”, & Page 5 – Section 3, “The LLM router selects mLLM APIs to include in the list. Let L ∈ [K]m denote the indexes of the m APIs selected by the router. Given a new query, it iteratively invokes the ith API in the list to obtain an answer fLi (q)”, thus performing […the analyzing, predicting and interfacing…] in response to determining that the large language model request differs from the plurality of previous large language model requests is disclosed, because Chen teaches that when no similar previously answered query is found in the cache, the system proceeds with processing the new query, including selecting an LLM API and invoking the selected LLM API to obtain an answer) Claims 3-4, 6-8, and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Chen et al. (hereinafter Chen, a non-patent literature reference titled “FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance”) in view of Jiang et al. (hereinafter Jiang) (CN 117573834), and in further view of Walker et al. (hereinafter Walker) (US 20190251069). Regarding Claim 3, Chen combined with Jiang teaches all the limitations of claim 2 as cited above. Chen combined with Jiang does not explicitly teach executing a hash function on […the vector…] to create a unique identifier for […the large language model request…]. However, Walker teaches: executing a hash function on […the vector…] to create a unique identifier for […the large language model request…] (Walker, Par. [0003], “each data vector having: i) a sequence of elements, wherein each element in each data vector can be configured to store a payload of data; and ii) a unique identifier based on a cryptographic hash of the sequence of elements”, & Par. [0019], “comparing a cryptographic hash of the proposed new data vector with the cryptographic hash of each data vector already stored in the at least one memory. When the at least one memory is not already storing proposed new data vector, the method can comprise storing the proposed new data vector in the at least one memory, and storing a unique identifier associated with the cryptographic hash of the proposed new data vector”, thus executing a hash function on […the vector…] to create a unique identifier for […the large language model request…] is disclosed, because Walker teaches generating a cryptographic hash from a data vector and associating a unique identifier with the cryptographic hash of that vector. Walker’s data vector corresponds to […the vector…], and the unique identifier associated with the cryptographic hash corresponds to the unique identifier for […the large language model request…] represented by the vector) It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to further combine Chen and Jiang with Walker’s teaching of executing a hash function on a vector to create a unique identifier. Jiang teaches generating a vector representation of the user problem, while Walker teaches deriving a unique identifier from the contents of a vector using a cryptographic hash. Therefore, a POSITA would have been motivated to apply Walker’s cryptographic hashing technique to the vector generated by Jiang so that the vector representing the large language model request could be uniquely identified and efficiently compared and managed, thereby improving the efficiency of processing and storing the request vectors (Walker, Par. [0125], “rather than assign an abstract identifier (VID) to a vector, a unique identifier can be derived based on the ordered elements of the vector. Such a unique identifier can be a content-based hash, such that the hash has an infinitesimal chance of collision for two different vectors. This identifier, a vector hash (or “Vhash”) can be a cryptographic hash”, & Par. [0146], “the Vhash of the new vector can be calculated and compared to the index of VHashs to know if the new vector is already present, without actually checking the contents of matching vectors, as shown, for example, in FIG. 17. This is in contrast to the case when VIDs are used: the contents of a vector with a matching has must be checked with the new vector (see FIG. 5). If the new vector is not in the vector store, its Vhash is indexed into the Vhash index; there is no need to assign an arbitrary VID. This increases the efficiency of data processing and data storage”) Regarding Claim 4, Chen and Jiang combined with Walker teaches all the limitations of claim 3 as cited above and Chen further teaches: receiving a response to the large language model request (Chen, Page 5 – Section 3, “Given a new query, it iteratively invokes the ith API in the list to obtain an answer fLi (q)”, thus eceiving a response to the large language model request is disclosed, because Chen teaches invoking a selected LLM API for a new query and obtaining an answer from the LLM API. Chen’s new query corresponds to the large language model request, and the obtained answer corresponds to the response to the large language model request) storing the response to the large language model request […] (Chen, Page 5 – Section 3, “the fundamental idea involves storing the response locally in a cache (e.g., a database) when submitting a query to an LLM API. To process a new query, we first verify if a similar query has been previously answered. If so, the response is retrieved from the cache”, thus storing the response to the large language model request […] is disclosed, because Chen teaches storing the response generated for a query submitted to an LLM API in a cache for later retrieval when the same or a similar query is received) Walker further teaches: […] in correspondence with the unique identifier for […the large language model request…] (Walker, Par. [0013], “each data vector having: i) a sequence of elements, wherein each element in each data vector can be configured to store a payload of data; and ii) a unique identifier based on a cryptographic hash of the sequence of elements”, thus […] in correspondence with the unique identifier for […the large language model request…] is disclosed, because Walker teaches associating a unique identifier with a data vector based on a cryptographic hash of that vector. When Walker’s unique identifier is applied to the vector representing […the large language model request…], the unique identifier corresponds to and identifies that large language model request) Regarding Claim 6, Chen combined with Jiang teaches all the limitations of claim 5 as cited above. Chen combined with Jiang does not explicitly teach creating a unique identifier for […the large language model request…], and comparing the unique identifier for […the large language model request…] to a plurality of stored unique identifiers corresponding to respective ones of […the plurality of previous large language model requests…] to determine whether the unique identifier for […the large language model request…] matches a stored unique identifier of the plurality of stored unique identifiers. However, Walker teaches: creating a unique identifier for […the large language model request…] (Walker, Par. [0013], “each data vector having: i) a sequence of elements, wherein each element in each data vector can be configured to store a payload of data; and ii) a unique identifier based on a cryptographic hash of the sequence of elements”, thus creating a unique identifier for […the large language model request…] is disclosed, because Walker teaches creating a unique identifier for a data vector based on a cryptographic hash of the vector. When applied to the vector representing the large language model request, Walker’s unique identifier corresponds to a unique identifier for the large language model request) comparing the unique identifier for […the large language model request…] to a plurality of stored unique identifiers corresponding to respective ones of […the plurality of previous large language model requests…] to determine whether the unique identifier for […the large language model request…] matches a stored unique identifier of the plurality of stored unique identifiers (Walker, Par. [0019], “determining whether the at least one memory is already storing by comparing a cryptographic hash of the proposed new data vector with the cryptographic hash of each data vector already stored in the at least one memory”, thus comparing the unique identifier for […the large language model request…] to a plurality of stored unique identifiers […] to determine whether the unique identifier for […the large language model request…] matches a stored unique identifier is disclosed, because Walker teaches comparing the cryptographic hash of a new data vector with the cryptographic hashes of previously stored data vectors to determine whether a match exists. When applied to the vectors representing […the plurality of previous large language model requests…], the stored cryptographic hashes correspond to the plurality of stored unique identifiers associated with the respective previous large language model requests) It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to further combine Chen and Jiang with Walker’s teaching of executing a hash function on a vector to create a unique identifier. Jiang teaches generating a vector representation of the user problem, while Walker teaches deriving a unique identifier from the contents of a vector using a cryptographic hash. Therefore, a POSITA would have been motivated to apply Walker’s cryptographic hashing technique to the vector generated by Jiang so that the vector representing the large language model request could be uniquely identified and efficiently compared and managed, thereby improving the efficiency of processing and storing the request vectors (Walker, Par. [0125], “rather than assign an abstract identifier (VID) to a vector, a unique identifier can be derived based on the ordered elements of the vector. Such a unique identifier can be a content-based hash, such that the hash has an infinitesimal chance of collision for two different vectors. This identifier, a vector hash (or “Vhash”) can be a cryptographic hash”, & Par. [0146], “the Vhash of the new vector can be calculated and compared to the index of VHashs to know if the new vector is already present, without actually checking the contents of matching vectors, as shown, for example, in FIG. 17. This is in contrast to the case when VIDs are used: the contents of a vector with a matching has must be checked with the new vector (see FIG. 5). If the new vector is not in the vector store, its Vhash is indexed into the Vhash index; there is no need to assign an arbitrary VID. This increases the efficiency of data processing and data storage”) Regarding Claim 7, Chen combined with Jiang teaches all the limitations of claim 1 as cited above and Chen further teaches: receiving an additional large language model request (Chen, Page 5 – Section 3, “Given a new query, it iteratively invokes the ith API in the list to obtain an answer fLi(q)”, thus receiving an additional large language model request is disclosed, because Chen teaches receiving and processing a new query after prior queries have been processed. Chen’s new query corresponds to the additional large language model request) Chen combined with Jiang does not explicitly teach creating a unique identifier for […the additional large language model request…], and comparing the unique identifier for […the additional large language model request…] to a plurality of stored unique identifiers corresponding to respective ones of a plurality of previous large language model requests to determine whether the unique identifier for […the additional large language model request…] matches a stored unique identifier of the plurality of stored unique identifiers. However, Walker teaches: creating a unique identifier for […the additional large language model request…] (Walker, Par. [0013], “each data vector having: i) a sequence of elements, wherein each element in each data vector can be configured to store a payload of data; and ii) a unique identifier based on a cryptographic hash of the sequence of elements”, thus creating a unique identifier for […the additional large language model request…] is disclosed, because Walker teaches generating a unique identifier for a data vector based on a cryptographic hash of the vector. When Walker’s hashing technique is applied to the vector representing […the additional large language model request…], the resulting unique identifier corresponds to the unique identifier for […the additional large language model request…]) comparing the unique identifier for […the additional large language model request…] to a plurality of stored unique identifiers corresponding to respective ones of a plurality of previous large language model requests to determine whether the unique identifier for […the additional large language model request…] matches a stored unique identifier of the plurality of stored unique identifiers (Walker, Par. [0019], “determining whether the at least one memory is already storing by comparing a cryptographic hash of the proposed new data vector with the cryptographic hash of each data vector already stored in the at least one memory”, thus comparing the unique identifier for […the additional large language model request…] to a plurality of stored unique identifiers corresponding to respective ones of a plurality of previous large language model requests to determine whether the unique identifier for […the additional large language model request…] matches a stored unique identifier of the plurality of stored unique identifiers is disclosed, because Walker teaches comparing the cryptographic hash of a new data vector with the cryptographic hashes of previously stored data vectors to determine whether a matching stored vector exists. When applied to vectors representing the plurality of previous large language model requests, the stored cryptographic hashes correspond to the stored unique identifiers for the respective previous large language model requests) It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to further combine Chen and Jiang with Walker’s teaching of executing a hash function on a vector to create a unique identifier. Jiang teaches generating a vector representation of the user problem, while Walker teaches deriving a unique identifier from the contents of a vector using a cryptographic hash. Therefore, a POSITA would have been motivated to apply Walker’s cryptographic hashing technique to the vector generated by Jiang so that the vector representing the large language model request could be uniquely identified and efficiently compared and managed, thereby improving the efficiency of processing and storing the request vectors (Walker, Par. [0125], “rather than assign an abstract identifier (VID) to a vector, a unique identifier can be derived based on the ordered elements of the vector. Such a unique identifier can be a content-based hash, such that the hash has an infinitesimal chance of collision for two different vectors. This identifier, a vector hash (or “Vhash”) can be a cryptographic hash”, & Par. [0146], “the Vhash of the new vector can be calculated and compared to the index of VHashs to know if the new vector is already present, without actually checking the contents of matching vectors, as shown, for example, in FIG. 17. This is in contrast to the case when VIDs are used: the contents of a vector with a matching has must be checked with the new vector (see FIG. 5). If the new vector is not in the vector store, its Vhash is indexed into the Vhash index; there is no need to assign an arbitrary VID. This increases the efficiency of data processing and data storage”) Regarding Claim 8, Chen and Jiang combined with Walker teaches all the limitations of claim 7 as cited above and Chen further teaches: retrieving a stored large language model response corresponding to […the stored unique identifier…] […] (Chen, Page 5 – Section 3, “To process a new query, we first verify if a similar query has been previously answered. If so, the response is retrieved from the cache”, thus Chen teaches retrieving a stored large language model response when a new large language model request is determined to match a previously answered large language model request) Walker further teaches: […] in response to determining that the unique identifier for […the additional large language model request…] matches the stored unique identifier (Walker, Par. [0009], “determine whether the at least one memory is already storing the proposed new data vector by comparing a cryptographic hash of the proposed new data vector with the cryptographic hash of each data vector already stored in the at least one memory”, thus determining that the unique identifier for […the additional large language model request…] matches the stored unique identifier is disclosed, because Walker teaches comparing the cryptographic hash of a new data vector with the cryptographic hashes of previously stored data vectors to determine whether a matching stored vector exists. When applied to the vector representing […the additional large language model request…], the cryptographic hash corresponds to the unique identifier for […the additional large language model request…], and the cryptographic hash of the matching stored vector corresponds to the stored unique identifier) Regarding Claim 18, Chen combined with Jiang teaches all the limitations of claim 17 as cited above. Chen combined with Jiang does not explicitly teach to creating a unique identifier for […the large language model request…], and to compare the unique identifier for […the large language model request…] to a plurality of stored unique identifiers corresponding to respective ones of […the plurality of previous large language model requests…] to determine whether the unique identifier for […the large language model request…] matches a stored unique identifier of the plurality of stored unique identifiers. However, Walker teaches: to creating a unique identifier for […the large language model request…] (Walker, Par. [0013], “each data vector having: i) a sequence of elements, wherein each element in each data vector can be configured to store a payload of data; and ii) a unique identifier based on a cryptographic hash of the sequence of elements”, thus creating a unique identifier for […the large language model request…] is disclosed, because Walker teaches creating a unique identifier for a data vector based on a cryptographic hash of the vector. When applied to the vector representing the large language model request, Walker’s unique identifier corresponds to a unique identifier for the large language model request) to compare the unique identifier for […the large language model request…] to a plurality of stored unique identifiers corresponding to respective ones of […the plurality of previous large language model requests…] to determine whether the unique identifier for […the large language model request…] matches a stored unique identifier of the plurality of stored unique identifiers (Walker, Par. [0019], “determining whether the at least one memory is already storing by comparing a cryptographic hash of the proposed new data vector with the cryptographic hash of each data vector already stored in the at least one memory”, thus comparing the unique identifier for […the large language model request…] to a plurality of stored unique identifiers […] to determine whether the unique identifier for […the large language model request…] matches a stored unique identifier is disclosed, because Walker teaches comparing the cryptographic hash of a new data vector with the cryptographic hashes of previously stored data vectors to determine whether a match exists. When applied to the vectors representing […the plurality of previous large language model requests…], the stored cryptographic hashes correspond to the plurality of stored unique identifiers associated with the respective previous large language model requests) It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to further combine Chen and Jiang with Walker’s teaching of executing a hash function on a vector to create a unique identifier. Jiang teaches generating a vector representation of the user problem, while Walker teaches deriving a unique identifier from the contents of a vector using a cryptographic hash. Therefore, a POSITA would have been motivated to apply Walker’s cryptographic hashing technique to the vector generated by Jiang so that the vector representing the large language model request could be uniquely identified and efficiently compared and managed, thereby improving the efficiency of processing and storing the request vectors (Walker, Par. [0125], “rather than assign an abstract identifier (VID) to a vector, a unique identifier can be derived based on the ordered elements of the vector. Such a unique identifier can be a content-based hash, such that the hash has an infinitesimal chance of collision for two different vectors. This identifier, a vector hash (or “Vhash”) can be a cryptographic hash”, & Par. [0146], “the Vhash of the new vector can be calculated and compared to the index of VHashs to know if the new vector is already present, without actually checking the contents of matching vectors, as shown, for example, in FIG. 17. This is in contrast to the case when VIDs are used: the contents of a vector with a matching has must be checked with the new vector (see FIG. 5). If the new vector is not in the vector store, its Vhash is indexed into the Vhash index; there is no need to assign an arbitrary VID. This increases the efficiency of data processing and data storage”) Claims 9 and 10 are rejected under 35 U.S.C. 103 as being unpatentable over Chen et al. (hereinafter Chen, a non-patent literature reference titled “FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance”) in view of Jiang et al. (hereinafter Jiang) (CN 117573834), and in further view of Wu et al. (hereinafter Wu, a non-patent literature reference titled “Multi-task learning based Encoder-Decoder: A comprehensive detection and diagnosis system for multi-sensor data”). Regarding Claim 9, Chen combined with Jiang teaches all the limitations of claim 1 as cited above and Chen further teaches: […the plurality of targets…] comprise the large language model […](Chen, Page 2, “LLM cascade focuses on how to adaptively choose which LLM APIs to use for different queries”, & “FrugalGPT learns how to adaptively triage different queries in the dataset to different combinations of LLMs, including ChatGPT, GPT-3 and GPT-4”, thus Chen teaches the large language model as a prediction/selection target because the system determines which LLM from a plurality of LLMs is appropriate for a particular query) Jiang further teaches: […] and the at least one database (Jiang, Page 5, “according to what enterprise the user dialog comes from, the user's intention to invoke the retrieval enhancement of the dialog to generate the model RAG flow, comprising the following steps: obtaining the preset screening condition of the current user intention, comprising the hot selling goods under the goods recommendation intention, the goods with stock and the price requirement; combining the preset screening condition, using the keyword matching and semantic recall to find the related knowledge from the enterprise knowledge base”, & Page 6, “storing the vector in the vector database, wherein the vector database uses the data structure to match with the index acceleration similarity; in each session, according to the problem and context of the user, using the semantic embedding model to convert the user problem into the vector expression; finding the most relevant example of the user problem”, thus Jiang teaches a database as another selection/retrieval target because the user request is used to identify relevant information from an enterprise knowledge base or vector database for constructing the LLM prompt) Chen combined with Jiang does not explicitly teach […the one or more machine learning algorithms…] comprise a neural network configured to predict a plurality of targets, and the neural network includes a plurality of parallel networks respectively corresponding to the plurality of targets. However, Wu teaches: […the one or more machine learning algorithms…] comprise a neural network configured to predict a plurality of targets (Wu, Page 5, “MTLED uses a shared encoder to obtain the feature matrix, based on which several task-related decoders are used to generate their own output, including results of anomaly detection, anomaly diagnosis and event detection”, &, “The overall framework of MTLED is a Multi-Task Learning based Encoder-Decoder”, thus […the one or more machine learning algorithms…] comprise a neural network configured to predict a plurality of targets is disclosed, because Wu teaches a multi-task neural network including a shared encoder and several task-related decoders that generate different outputs for anomaly detection, anomaly diagnosis, and event detection, wherein the different task outputs correspond to a plurality of prediction targets) the neural network includes a plurality of parallel networks respectively corresponding to the plurality of targets (Wu, Page 5, “five decoders are designed: one Decoder for Anomaly DEtection (Dec_ADE), three Decoders for Anomaly DIagnosis (Dec_ADI), which provide anomaly types, anomaly states and anomaly channels respectively, and one Decoder for Event Detection (Dec_ED), & Page 6 -Figure 2, showing the shared feature matrix connected to separate decoder branches for known event, anomaly channel, anomaly state, anomaly type, and anomaly point, thus the neural network includes a plurality of parallel networks respectively corresponding to the plurality of targets is disclosed, because Wu teaches multiple task specific decoder networks operating from a shared feature representation, with the separate decoder networks corresponding to respective prediction outputs or targets) It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to further combine Chen and Jiang with Wu’s teaching of a multi-task neural network having a shared encoder and a plurality of task related decoder networks corresponding to different prediction targets. Chen and Jiang teach determining, based on a user request, different outputs including an appropriate large language model and an appropriate database, while Wu teaches using a shared encoder to obtain a common feature representation and multiple task related decoders to generate respective prediction outputs. Therefore, a POSITA would have been motivated to apply Wu’s multi-task neural network architecture to the large language model and database prediction tasks taught by Chen and Jiang so that features derived from the same request could be shared while separate task specific networks generate the respective large language model and database predictions, thereby improving feature extraction and the performance of the respective prediction tasks while enabling multiple related tasks to be performed within a single deep learning model (Wu, Page 3, “the multi-task learning framework in MTLED can theoretically improve the ability of feature extraction and the performance of each task”, & Page 5, “MTLED uses a shared encoder to obtain the feature matrix, based on which several task-related decoders are used to generate their own output, including results of anomaly detection, anomaly diagnosis and event detection”, & Page 12, “MTLED provides an intuitive and effective way to realize various tasks related to system monitoring and early warning in one deep learning model”) Regarding Claim 10, Chen and Jiang combined with Wu teaches all the limitations of claim 9 as cited above and Wu further teaches: wherein a first parallel network of the plurality of parallel networks corresponding to […the large language model…] comprises a multi-class classifier and a second parallel network of the plurality of parallel networks corresponding to […the at least one database…]comprises a multi-label classifier (Wu, Page 6, “Since multiple anomaly types and anomaly channels are allowed in one sliding window, multi-label classifiers are applied for anomaly type and anomaly channel. On the contrary, only one anomaly state is defined for each sliding window, and thus a multi-class classifier is more suitable and adopted”, & Page 7, “Among the five parallel decoders, only the one for anomaly state is a multi-class classifier, the other four are multi-label classifiers or regression models”, thus Wu teaches configuring different parallel networks with a multi-class classifier and a multi-label classifier according to the respective prediction targets, where a multi-class classifier is used when a single value is predicted and a multi-label classifier is used when multiple values may be predicted) Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to MAHLIET ADMASU whose telephone number is (571)272-0034. The examiner can normally be reached Mon-Fri, 8am-5pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Alexey Shmatov can be reached at (571)270-3428. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /M.T.A./Examiner, Art Unit 2123 /ALEXEY SHMATOV/Supervisory Patent Examiner, Art Unit 2123
Read full office action

Prosecution Timeline

Jun 10, 2024
Application Filed
Sep 23, 2026
Non-Final Rejection mailed — §101, §103 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
0%
Grant Probability
0%
With Interview (+0.0%)
3y 5m (~1y 1m remaining)
Median Time to Grant
Low
PTA Risk
Based on 1 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month