DETAILED ACTION
In response to communication filed on 13 May 2026, claims 1, 7, 8, 10, and 15-17 are amended. Claims 1-20 are pending.
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Continued Examination Under 37 CFR 1.114
A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 13 May 2026 has been entered.
Response to Arguments
Applicant’s arguments, see “Rejections Under 35 U.S.C. § 103”, filed 13 May 2026, have been carefully considered. The arguments are related to newly added limitations and are addressed in the rejection below
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-3, 7, 9-12, 15 and 17-19 are rejected under 35 U.S.C. 103 as being unpatentable over Buniatyan (US 12,182,125 B1, hereinafter “Buniatyan”) in view of Lopatenko (US 9,009,146 B1, hereinafter “Lopatenko”) further in view of Sengupta et al. (US 12,591,606 B1, hereinafter “Sengupta”) and Deb et al. (US 12,292,811 B1, hereinafter “Deb”).
Regarding claim 1, Buniatyan teaches
A system, comprising: one or more computer processors; and (see Buniatyan, [col 12 lines 24-26] “The computing system 180 can include one or more servers, computing systems, data centers (e.g., cloud computing systems), processors, and memory”).
one or more computer-readable mediums storing instructions that, when executed by the one or more computer processors, cause the system to: (see Buniatyan, [col 30 lines 50-54] “Such instructions can be read into main memory 405 from another computer-readable medium, such as the storage device 415. Execution of the arrangement of instructions contained in main memory 405 causes the data processing system 105”).
receive a query from a user; (see Buniatyan, [col 19 lines 26-29] “the client device 130 or the data processing system 105 can provide the query to the query identifier via a graphical user interface”; [col 10 line 67-col 11 line 2] “a query may be any type of text input that may be used to retrieve relevant information from a corresponding text, image, video, or audio corpus”).
generate a vector embedding of the query; (see Buniatyan, [col 25 lines 24-26] “used to generate a corresponding vector search input query (e.g., text data used to generate training embeddings 138)”; [col 7 lines 15-17] “may include embeddings vectors that numerically represent data of any suitable modality (e.g., text, images, video, other data, etc.)”; [col 9 lines 40-44] “These embedding vectors can capture semantic or contextual characteristics of the input data samples, allowing for efficient computation, analysis, and downstream tasks”).
perform a semantic search between (see Buniatyan, [col 4 lines 7-19] “Embeddings can be used to transform arbitrary data, such as text or image data, to numerical data structures (e.g., vectors) in a continuous n-dimensional space (e.g., an embeddings space)… in which input embeddings are used to identify similar embeddings in a pre-generated corpus, to identify relevant portions of data without necessarily re-training or updating machine-learning models”) the vector embedding of the query (see Buniatyan, [col 25 lines 24-26] “used to generate a corresponding vector search input query (e.g., text data used to generate training embeddings 138)”; [col 7 lines 15-17] “may include embeddings vectors that numerically represent data of any suitable modality (e.g., text, images, video, other data, etc.)”; [col 9 lines 40-44] “These embedding vectors can capture semantic or contextual characteristics of the input data samples, allowing for efficient computation, analysis, and downstream tasks”) and vector embeddings of each of a plurality of historical queries… (see Buniatyan, [col 22 lines 14-23] “The set of query results may be stored in a historic query database… may receive requests to execute historic queries/vector search operations and may provide the historic queries (and indications of any cached query results corresponding thereto) to the query identifier 115”) wherein a model discovery database stores,… (see Buniatyan, [col 29 lines 49-] “The data processing system may store the set of query results in one or more data structures in the memory of the data processing system 105… The set of query results may be stored in a historic query database”) a historical response to the historical query received from the LLM,… (see Buniatyan, [col 13 lines 44-55] “The output data produced by the machine-learning model 182 can be accessed, retrieved, or otherwise stored by the data processing system… The machine-learning model 182 can be a general-purpose large language model, or a large language model that was specifically fine-tuned to generate text-based queries for input data”; [col 22 lines 14-23] “The set of query results may be stored in a historic query database… The set of query results may be stored in a historic query database, for example, in association with the multi-dimensional sample dataset to which the query corresponds… may receive requests to execute historic queries/vector search operations and may provide the historic queries (and indications of any cached query results corresponding thereto) to the query identifier 115”) the historical query; (see Buniatyan, [col 22 lines 14-23] “The set of query results may be stored in a historic query database… The set of query results may be stored in a historic query database, for example, in association with the multi-dimensional sample dataset to which the query corresponds).
for each LLM of the plurality of LLMs and… (see Buniatyan, [col 12 lines 43-47] “the machine-learning model(s) 182 include one or more large-language models (LLMs). Large language models include generative models that are pre-trained on large amounts of text and can generate corresponding text in response to an input context”).
for each LLM of the plurality of LLMs and… (see Buniatyan, [col 12 lines 43-47] “the machine-learning model(s) 182 include one or more large-language models (LLMs). Large language models include generative models that are pre-trained on large amounts of text and can generate corresponding text in response to an input context”).
for each LLM of the plurality of LLMs,… (see Buniatyan, [col 12 lines 43-47] “the machine-learning model(s) 182 include one or more large-language models (LLMs). Large language models include generative models that are pre-trained on large amounts of text and can generate corresponding text in response to an input context”).
transmit, to a user interface of a client device, data relating to queries (see Buniatyan, [col 6 lines 57-63] “the data processing system 105 can construct, generate, build, implement, or create one or more graphical user interfaces, which may be provided for display on the client device 130, and may display various data relating to queries, transformation data structures 134 (including training processes thereof), or any embeddings (or from which embeddings are generated)”).
Buniatyan does not explicitly teach to identify a predetermined number of the plurality of historical queries that are semantically related to the received query, a model discovery database stores for each pair of a historical query of the plurality of historical queries and a large language model (LLM) of a plurality of LLMs, associated historical metadata generated based on execution of the historical query by the LLM, and a quality rank that ranks the LLM from among the plurality of LLMs for the historical query; for each historical query of the identified predetermined number of historical queries, retrieve, from the model discovery database, the associated historical metadata and the quality rank corresponding to the pair of the historical query and the LLM; for each of a plurality of predetermined metrics,… aggregate the retrieved associated historical metadata and the quality ranks for the LLM across the identified predetermined number of historical queries to determine a score for the predetermined metric for the LLM; determine an overall score of the LLM based on the determined scores for the plurality of predetermined metrics for the LLM; and a ranked list of the plurality of LLMs based on the overall scores.
However, Lopatenko discloses similarity between queries and teaches
to identify a predetermined number of the plurality of historical queries that are semantically related to the received query, (see Lopatenko, [col 15 line 33 – col 16 line 46] “With a historical query selected for comparison, the selected historical query is compared to the input query (step 5012)… After determining the match score for the selected historical query, a determination is made as to whether there are more historical queries to compare to the input query (step 5014)… each historical query within the defined set of historical queries is compared… comparison is performed on queries within the set until a desired number of historical queries with a match score above a threshold value are found… a fixed number of historical queries can be selected (e.g., the historical queries having the top 2, 3, 4, 5, 10, 15, 20, 50, 100 match scores)… the historical queries with match scores in a certain percentile range can be selected (e.g., select the historical queries in the top 0.1%, 0.5%, 1%, 3%, 5%, 10% of match scores)”).
for each historical query of the identified predetermined number of historical queries, (see Lopatenko, [col 15 line 33 – col 16 line 46] “With a historical query selected for comparison, the selected historical query is compared to the input query (step 5012)… After determining the match score for the selected historical query, a determination is made as to whether there are more historical queries to compare to the input query (step 5014)… each historical query within the defined set of historical queries is compared… comparison is performed on queries within the set until a desired number of historical queries with a match score above a threshold value are found… a fixed number of historical queries can be selected (e.g., the historical queries having the top 2, 3, 4, 5, 10, 15, 20, 50, 100 match scores)… the historical queries with match scores in a certain percentile range can be selected (e.g., select the historical queries in the top 0.1%, 0.5%, 1%, 3%, 5%, 10% of match scores)”).
the identified predetermined number of historical queries (see Lopatenko, [col 15 line 33 – col 16 line 46] “With a historical query selected for comparison, the selected historical query is compared to the input query (step 5012)… After determining the match score for the selected historical query, a determination is made as to whether there are more historical queries to compare to the input query (step 5014)… each historical query within the defined set of historical queries is compared… comparison is performed on queries within the set until a desired number of historical queries with a match score above a threshold value are found… a fixed number of historical queries can be selected (e.g., the historical queries having the top 2, 3, 4, 5, 10, 15, 20, 50, 100 match scores)… the historical queries with match scores in a certain percentile range can be selected (e.g., select the historical queries in the top 0.1%, 0.5%, 1%, 3%, 5%, 10% of match scores)”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to include the functionality of identifying predetermined number of semantically related queries as being disclosed and taught by Lopatenko, in the system taught by Buniatyan to yield the predictable results of providing improved query results (see Lopatenko, [col 4 lines 64-67] “illustrating improvement of query results from a search engine through identification of historical queries that are similar to an input query for the purpose of borrowing user behavior data”).
The proposed combination of Buniatyan and Lopatenko does not explicitly teach a model discovery database stores for each pair of a historical query of the plurality of historical queries and a large language model (LLM) of a plurality of LLMs, associated historical metadata generated based on execution of the historical query by the LLM, and a quality rank that ranks the LLM from among the plurality of LLMs for the historical query; retrieve, from the model discovery database, the associated historical metadata and the quality rank corresponding to the pair of the historical query and the LLM; for each of a plurality of predetermined metrics, aggregate the retrieved associated historical metadata and the quality ranks for the LLM across to determine a score for the predetermined metric for the LLM; determine an overall score of the LLM based on the determined scores for the plurality of predetermined metrics for the LLM; and a ranked list of the plurality of LLMs based on the overall scores.
However, Sengupta discloses language processing machine learning and teaches
a conversation history data store for each pair of a historical query of the plurality of historical queries and a large language model (LLM) of a plurality of LLMs, (see Sengupta, [col 6 line 59- col 7 line 25] “A history data retrieval component 220 may retrieve conversation history data 222 from a conversation history data store 230 based on user query 102… conversation history data store 230 may be a database or other data storage entity that stores conversation history data, such as prior user queries (e.g., including prior enriched user queries)… associated responses that were generated to the prior user queries… History data retrieval component 220 may retrieve such related conversation history data for user query 102 based on one or more attributes related to user query 102, such as… a software tool that is being used to handle user query 102 and/or one or more previous queries (e.g., a software tool associated with language processing machine learning model 240), other types of metadata, and/or the like, such as based on one or more such attributes… Language processing machine learning model 240 may, for example be a large language model (LLM), such as having a transformer architecture, or may be any suitable type of machine learning model that has been trained to process natural language inputs”) associated historical metadata generated based on execution of the historical query by the LLM, and (see Sengupta, [col 6 line 59- col 7 line 25] “A history data retrieval component 220 may retrieve conversation history data 222 from a conversation history data store 230 based on user query 102… conversation history data store 230 may be a database or other data storage entity that stores conversation history data, such as prior user queries (e.g., including prior enriched user queries)… associated responses that were generated to the prior user queries… History data retrieval component 220 may retrieve such related conversation history data for user query 102 based on one or more attributes related to user query 102, such as… a software tool that is being used to handle user query 102 and/or one or more previous queries (e.g., a software tool associated with language processing machine learning model 240), other types of metadata, and/or the like, such as based on one or more such attributes… Language processing machine learning model 240 may, for example be a large language model (LLM), such as having a transformer architecture, or may be any suitable type of machine learning model that has been trained to process natural language inputs”).
retrieve, from the model discovery database, the associated historical metadata (see Sengupta, [col 6 line 59- col 7 line 10] “A history data retrieval component 220 may retrieve conversation history data 222 from a conversation history data store 230… may retrieve such related conversation history data for user query 102 based on one or more attributes related to user query 102, such as a conversation identifier, a user identifier, a session identifier… other types of metadata, and/or the like, such as based on one or more such attributes also being associated with conversation history data 222 in conversation history data store 230”) pair of the historical query and the LLM; (see Sengupta, [col 6 line 59- col 7 line 25] “A history data retrieval component 220 may retrieve conversation history data 222 from a conversation history data store 230 based on user query 102… conversation history data store 230 may be a database or other data storage entity that stores conversation history data, such as prior user queries (e.g., including prior enriched user queries)… History data retrieval component 220 may retrieve such related conversation history data for user query 102 based on one or more attributes related to user query 102, such as… a software tool that is being used to handle user query 102 and/or one or more previous queries (e.g., a software tool associated with language processing machine learning model 240)… Language processing machine learning model 240 may, for example be a large language model (LLM), such as having a transformer architecture, or may be any suitable type of machine learning model that has been trained to process natural language inputs”).
aggregate the retrieved associated historical metadata and (see Sengupta, [col 6 line 59- col 7 line 10] “A history data retrieval component 220 may retrieve conversation history data 222 from a conversation history data store 230… may retrieve such related conversation history data for user query 102 based on one or more attributes related to user query 102, such as a conversation identifier, a user identifier, a session identifier… other types of metadata, and/or the like, such as based on one or more such attributes also being associated with conversation history data 222 in conversation history data store 230”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to include the functionality of pair of historical query along with large language model, associated historical metadata, retrieving associated historical metadata and aggregate the retrieved associated historical metadata as being disclosed and taught by Sengupta, in the system taught by the proposed combination of Buniatyan and Lopatenko to yield the predictable results of improving the ability of machine learning models (see Sengupta, [col 3 lines 35-45] “Techniques described herein improve the technical field of automated response generation in computing applications in a number of ways… improve the ability of a machine learning model to understand and process the query, resulting in automated generation of responses that have a higher level of accuracy and relevance than would be possible if the query was provided in its original form to the machine learning model”).
The proposed combination of Buniatyan, Lopatenko and Sengupta does not explicitly teach a quality rank that ranks the LLM from among the plurality of LLMs for the historical query; and the quality rank corresponding to the pair of the historical query and the LLM; for each of a plurality of predetermined metrics,… the quality ranks for the LLM across to determine a score for the predetermined metric for the LLM; determine an overall score of the LLM based on the determined scores for the plurality of predetermined metrics for the LLM; and a ranked list of the plurality of LLMs based on the overall scores.
However, Deb discloses selection of models and teaches
a quality rank that ranks the LLM from among the plurality of LLMs for historical performance (see Deb, [col 45 lines 47-48] “to rank the AI models, identifying those that best meet the overall requirements of the request”; [col 46 lines 2-4] “the system prioritizes each AI model in the plurality of AI models based on historical performance data of each AI model in the plurality of AI models”; [col 42 lines 24-26] “at least one AI model in the plurality of AI models is a Large Language Model (LLM)”).
and the quality rank corresponding to the historical performance of LLM (see Deb, [col 45 lines 47-48] “to rank the AI models, identifying those that best meet the overall requirements of the request”; [col 46 lines 2-4] “the system prioritizes each AI model in the plurality of AI models based on historical performance data of each AI model in the plurality of AI models”; [col 42 lines 24-26] “at least one AI model in the plurality of AI models is a Large Language Model (LLM)”).
for each of a plurality of predetermined metrics,… the quality ranks for the LLM across based on performance (see Deb, [col 45 lines 47-48] “to rank the AI models, identifying those that best meet the overall requirements of the request”; [col 46 lines 2-4] “the system prioritizes each AI model in the plurality of AI models based on historical performance data of each AI model in the plurality of AI models”; [col 42 lines 24-26] “at least one AI model in the plurality of AI models is a Large Language Model (LLM)”) to determine a score for the predetermined metric for the LLM; (see Deb, [col 45 lines 35-58] “The system can compare the determined capabilities a first model of the plurality of models with the determined capabilities of a second model of the plurality of models. The system can use a scoring mechanism that assigns a compatibility score to each AI model based on how well its capabilities match the expected values. The scoring mechanism can use weighted criteria to prioritize certain attributes over others, depending on the specific requirements of the request… to rank the AI models, identifying those that best meet the overall requirements of the request. The system can normalize the performance metrics and expected values to a common scale to allow different metrics can be compared and aggregated. The system applies weights to each metric based on the importance of the corresponding attribute. The weights can be predefined based on the type of request or dynamically adjusted based on user preferences or contextual factors. For instance, a weight of 0.7 can be assigned to accuracy and 0.3 can be assigned to latency for a medical diagnosis task, reflecting the higher priority of accuracy”; [col 46 lines 2-4] “the system prioritizes each AI model in the plurality of AI models based on historical performance data of each AI model in the plurality of AI models”; [col 42 lines 24-26] “at least one AI model in the plurality of AI models is a Large Language Model (LLM)”).
determine an overall score of the LLM based on the determined scores for the plurality of predetermined metrics for the LLM; and (see Deb, [col 45 lines 59=63] “Once the weights are applied, the system calculates a weighted sum for each AI model, representing its overall compatibility score. The score is a composite measure that reflects how well the model's capabilities align with the expected values across all relevant attributes”; [col 42 lines 24-26] “at least one AI model in the plurality of AI models is a Large Language Model (LLM)”).
a ranked list of the plurality of LLMs based on the overall scores (see Deb, [col 45 lines 59-67] “Once the weights are applied, the system calculates a weighted sum for each AI model, representing its overall compatibility score. The score is a composite measure that reflects how well the model's capabilities align with the expected values across all relevant attributes. The system aggregates the scores to rank the AI models, identifying those that best meet the overall requirements of the request. The models with the highest compatibility scores are selected”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to include the functionality of ranking, scoring, metrics, prompt tokens, output tokens, comparison, weighted scores, cost, latency and rank as being disclosed and taught by Deb, in the system taught by the proposed combination of Buniatyan, Lopatenko and Sengupta to yield the predictable results of improving the accuracy of estimated performance metric value determination (see Deb, [col 26 lines 47-67] “generates the estimated performance metric value based on providing the prompt to an evaluation model… an indication of the first model ( e.g., LLM) to a performance metric evaluation model to generate the first estimated performance metric value… By doing so, the data generation platform 102 improves the accuracy of estimated performance metric value determination, thereby mitigating overuse of system resources”).
Claims 10 and 17 incorporates substantively all the limitations of claim 1 in a method form (see Buniatyan, [col 3 lines 66-67] “implementations of, techniques, approaches, methods, apparatuses, and systems”) and computer-readable medium form (see Buniatyan, [col 3 lines 22-27] “Aspects can be implemented in any convenient form, for example, by appropriate computer programs, which may be carried on appropriate carrier media ( computer readable media), which may be tangible carrier media (e.g., disks) or intangible carrier media (e.g., communications signals)”) and are rejected under the same rationale.
Regarding claim 2, the proposed combination of Buniatyan, Lopatenko, Sengupta and Deb teaches
wherein the instructions further cause the system to: (see Buniatyan, [col 30 lines 50-54] “Such instructions can be read into main memory 405 from another computer-readable medium, such as the storage device 415. Execution of the arrangement of instructions contained in main memory 405 causes the data processing system 105”).
input the received query to each of the plurality of LLMs to generate a corresponding sample response to the query; (see Buniatyan, [col 13 lines 42-55] “executes the machine-learning model 182 to produce output data… the input context for the machine-learning model 182 can cause the machine-learning model 182 to generate corresponding text-based queries that would result in the provided portions of a corpus being retrieved were the text-based query used in a search operation over said corpus of information. The machine-learning model 182 can be a general-purpose large language model, or a large language model that was specifically fine-tuned to generate text-based queries for input data”).
wherein the ranked list of the plurality of LLMs includes, (see Deb, [col 45 lines 59-67] “Once the weights are applied, the system calculates a weighted sum for each AI model, representing its overall compatibility score. The score is a composite measure that reflects how well the model's capabilities align with the expected values across all relevant attributes. The system aggregates the scores to rank the AI models, identifying those that best meet the overall requirements of the request. The models with the highest compatibility scores are selected”) for each LLM, the determined score for one or more of the plurality of predetermined metrics (see Deb, [col 45 lines 35-58] “The system can compare the determined capabilities a first model of the plurality of models with the determined capabilities of a second model of the plurality of models. The system can use a scoring mechanism that assigns a compatibility score to each AI model based on how well its capabilities match the expected values. The scoring mechanism can use weighted criteria to prioritize certain attributes over others, depending on the specific requirements of the request… to rank the AI models, identifying those that best meet the overall requirements of the request. The system can normalize the performance metrics and expected values to a common scale to allow different metrics can be compared and aggregated. The system applies weights to each metric based on the importance of the corresponding attribute. The weights can be predefined based on the type of request or dynamically adjusted based on user preferences or contextual factors. For instance, a weight of 0.7 can be assigned to accuracy and 0.3 can be assigned to latency for a medical diagnosis task, reflecting the higher priority of accuracy”; [col 46 lines 2-4] “the system prioritizes each AI model in the plurality of AI models based on historical performance data of each AI model in the plurality of AI models”; [col 42 lines 24-26] “at least one AI model in the plurality of AI models is a Large Language Model (LLM)”) and the corresponding sample response generated by the LLM (see Buniatyan, [col 13 lines 42-55] “executes the machine-learning model 182 to produce output data… the input context for the machine-learning model 182 can cause the machine-learning model 182 to generate corresponding text-based queries that would result in the provided portions of a corpus being retrieved were the text-based query used in a search operation over said corpus of information. The machine-learning model 182 can be a general-purpose large language model, or a large language model that was specifically fine-tuned to generate text-based queries for input data”). The motivation for the proposed combination is maintained.
Claims 11 and 18 incorporate substantively all the limitations of claim 2 in a method and computer-readable medium form and are rejected under the same rationale.
Regarding claim 3, the proposed combination of Buniatyan, Lopatenko, Sengupta and Deb teaches
wherein the instructions further cause the system to: (see Buniatyan, [col 30 lines 50-54] “Such instructions can be read into main memory 405 from another computer-readable medium, such as the storage device 415. Execution of the arrangement of instructions contained in main memory 405 causes the data processing system 105”).
receive, from the user interface of the client device, a selection of one or more of embedding model (see Buniatyan, [col 9 line 55 – col 10 line 4] “the client device 130 and/or the computing system 180 can transmit a request (e.g., in response to user input, etc.) to generate a set of corpus embeddings 136 from a specified corpus of information (which may be specified via a location in the data lake 160)… select the embeddings model for generation of the corpus of embeddings 136, for example, based on the type of data in the specified corpus. For example, the data processing system 105 can select a text-specific embeddings model if the corpus of input data includes text data, or an image-specific embeddings model if the corpus of input data includes images. Similar approaches may be utilized for other types of embeddings”) the LLMs from the ranked list; and (see Deb, [col 45 lines 59-67] “Once the weights are applied, the system calculates a weighted sum for each AI model, representing its overall compatibility score. The score is a composite measure that reflects how well the model's capabilities align with the expected values across all relevant attributes. The system aggregates the scores to rank the AI models, identifying those that best meet the overall requirements of the request. The models with the highest compatibility scores are selected”).
automatically configure a model serving endpoint based on the received selection (see Buniatyan, [col 9 line 55 – col 10 line 4] “the client device 130 and/or the computing system 180 can transmit a request (e.g., in response to user input, etc.) to generate a set of corpus embeddings 136 from a specified corpus of information (which may be specified via a location in the data lake 160)… select the embeddings model for generation of the corpus of embeddings 136, for example, based on the type of data in the specified corpus. For example, the data processing system 105 can select a text-specific embeddings model if the corpus of input data includes text data, or an image-specific embeddings model if the corpus of input data includes images. Similar approaches may be utilized for other types of embeddings”). The motivation for the proposed combination is maintained.
Claims 12 and 19 incorporate substantively all the limitations of claim 3 in a method and computer-readable medium form and are rejected under the same rationale.
Regarding claim 7, the proposed combination of Buniatyan, Lopatenko, Sengupta and Deb teaches
wherein the quality rank of the LLM based on historical performance (see Deb, [col 45 lines 47-48] “to rank the AI models, identifying those that best meet the overall requirements of the request”; [col 46 lines 2-4] “the system prioritizes each AI model in the plurality of AI models based on historical performance data of each AI model in the plurality of AI models”; [col 42 lines 24-26] “at least one AI model in the plurality of AI models is a Large Language Model (LLM)”) stored in the model discovery database for (see Buniatyan, [col 29 lines 49-] “The data processing system may store the set of query results in one or more data structures in the memory of the data processing system 105… The set of query results may be stored in a historic query database”) each pair of a historical query and an LLM is based on (see Sengupta, [col 6 line 59- col 7 line 25] “A history data retrieval component 220 may retrieve conversation history data 222 from a conversation history data store 230 based on user query 102… conversation history data store 230 may be a database or other data storage entity that stores conversation history data, such as prior user queries (e.g., including prior enriched user queries)… associated responses that were generated to the prior user queries… History data retrieval component 220 may retrieve such related conversation history data for user query 102 based on one or more attributes related to user query 102, such as… a software tool that is being used to handle user query 102 and/or one or more previous queries (e.g., a software tool associated with language processing machine learning model 240), other types of metadata, and/or the like, such as based on one or more such attributes… Language processing machine learning model 240 may, for example be a large language model (LLM), such as having a transformer architecture, or may be any suitable type of machine learning model that has been trained to process natural language inputs”) a comparison between expected output (see Deb, [col 37 line 56] “compare the expected output with the actual test output”) the historical response to the historical query (see Buniatyan, [col 22 lines 14-15] “The set of query results may be stored in a historic query database”) and test output (see Deb, [col 37 line 56] “compare the expected output with the actual test output”) a ground- truth response to an output of the historical query from a ground-truth model (see Buniatyan, [col 13 lines 64-67] “can utilize the same embeddings model to generate the synthetic training embeddings 138 as was used to generate the ground truth training embeddings 138”; [col 16 lines 52-55] “performing a similarity calculation between the output embeddings and the ground truth data (e.g., the corpus embeddings 136 identified in the mappings 140 corresponding to the input training data)”). The motivation for the proposed combination is maintained.
Claim 15 incorporates substantively all the limitations of claim 7 in a method form and is rejected under the same rationale.
Regarding claim 9, the proposed combination of Buniatyan, Lopatenko, Sengupta and Deb teaches
wherein the predetermined metrics include cost, (see Deb, [col 51 lines 48-50] “Cost representation 1622 can include metrics such as the cost per request, total cost over a specified period, and cost breakdown by model or resource type”) latency, and (see Deb, [col 44 lines 26-30] “The values of the set of estimated performance metrics for each particular AI model in the plurality of AI models can include, for example, response time, accuracy, and/or latency”) rank (see Deb, [col 23 lines 21-23] “enables the prioritization of relevant performance metrics (e.g., cost) over other metrics (e.g., memory usage) according to system requirements”). The motivation for the proposed combination is maintained.
Claims 4-5, 13 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Buniatyan, Lopatenko, Sengupta and Deb in view of Rasband et al. (US 2023/0140084 A1, hereinafter “Rasband”).
Regarding claim 4, the proposed combination of Buniatyan, Lopatenko, Sengupta and Deb teaches
wherein the instructions further cause the system to: (see Buniatyan, [col 30 lines 50-54] “Such instructions can be read into main memory 405 from another computer-readable medium, such as the storage device 415. Execution of the arrangement of instructions contained in main memory 405 causes the data processing system 105”).
in response to determining that the received selection includes a selection of two or more of the LLMs,… (see Buniatyan, [col 9 line 55 – col 10 line 4] “the client device 130 and/or the computing system 180 can transmit a request (e.g., in response to user input, etc.) to generate a set of corpus embeddings 136 from a specified corpus of information (which may be specified via a location in the data lake 160)… select the embeddings model for generation of the corpus of embeddings 136, for example, based on the type of data in the specified corpus. For example, the data processing system 105 can select a text-specific embeddings model if the corpus of input data includes text data, or an image-specific embeddings model if the corpus of input data includes images. Similar approaches may be utilized for other types of embeddings” – there are plurality of models being selected; [col 4 lines 40-43] “large language models (LLMs) or other machine-learning models, thereby improving overall computational performance of machine-learning systems that utilize embeddings”) the two or more LLMs (see Buniatyan, [col 9 line 55 – col 10 line 4] “the client device 130 and/or the computing system 180 can transmit a request (e.g., in response to user input, etc.) to generate a set of corpus embeddings 136 from a specified corpus of information (which may be specified via a location in the data lake 160)… select the embeddings model for generation of the corpus of embeddings 136, for example, based on the type of data in the specified corpus. For example, the data processing system 105 can select a text-specific embeddings model if the corpus of input data includes text data, or an image-specific embeddings model if the corpus of input data includes images. Similar approaches may be utilized for other types of embeddings” – there are plurality of models being selected; [col 4 lines 40-43] “large language models (LLMs) or other machine-learning models, thereby improving overall computational performance of machine-learning systems that utilize embeddings”) based on their respective overall scores (see Deb, [col 45 lines 59=63] “Once the weights are applied, the system calculates a weighted sum for each AI model, representing its overall compatibility score. The score is a composite measure that reflects how well the model's capabilities align with the expected values across all relevant attributes”; [col 42 lines 24-26] “at least one AI model in the plurality of AI models is a Large Language Model (LLM)”).
The proposed combination of Buniatyan, Lopatenko, Sengupta and Deb does not explicitly teach automatically determine a traffic routing weight for each of the two or more LLMs.
However, Rasband discloses monitoring traffic data and teaches
automatically determine a traffic routing weight for each of the system (see Rasband, [0056] “a weighting algorithm can be used to dynamically route traffic based on the output of the HMS… some markers, flags, and indicators from the HMS 120 can be given higher weights, while others can be given less weight in dynamic traffic routing… the output of the DTM 214 can be given higher weight in routing traffic than the output of the RSM 212 because a downtime duration can be more detrimental to the traffic flow than an up and running remote system 104, which may otherwise have a lower responsiveness score”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to include the functionality of traffic routing weight and scores being higher than the other as being disclosed and taught by Rasband in the system taught by the proposed combination of Buniatyan, Lopatenko, Sengupta and Deb to yield the predictable results of providing more accurate predictions specially in cases when services are time critical (see Rasband, [0056] “the output of the HMS 120 can be used for prediction of future unexpected outages of a remote system 104. Routing traffic based on such predictions can be beneficial for the service providers 102, particularly in cases where their underlying services are time critical.”).
Claims 13 and 20 incorporate substantively all the limitations of claim 4 in a method and computer-readable medium form and are rejected under the same rationale.
Regarding claim 5, the proposed combination of Buniatyan, Lopatenko, Sengupta, Deb and Rasband teaches
wherein a traffic routing weight of the system (see Rasband, [0056] “a weighting algorithm can be used to dynamically route traffic based on the output of the HMS… some markers, flags, and indicators from the HMS 120 can be given higher weights, while others can be given less weight in dynamic traffic routing”) a first LLM (see Buniatyan, [col 4 lines 40-43] “large language models (LLMs) or other machine-learning models, thereby improving overall computational performance of machine-learning systems that utilize embeddings” – there are plurality of LLMs) having a first overall score is higher than (see Rasband, [0031] “The RSM 212 can monitor the status of traffic to and from a remote system 104, over a period of time, and assign a responsiveness score to the remote system 104. The responsiveness score can be used by the local services 114 to dynamically route traffic to remote systems 104 that have obtained a higher responsiveness score in the recent past”) a traffic routing weight of remote system (see Rasband, [0056] “a weighting algorithm can be used to dynamically route traffic based on the output of the HMS… some markers, flags, and indicators from the HMS 120 can be given higher weights, while others can be given less weight in dynamic traffic routing… the output of the DTM 214 can be given higher weight in routing traffic than the output of the RSM 212 because a downtime duration can be more detrimental to the traffic flow than an up and running remote system 104, which may otherwise have a lower responsiveness score”; [0022] “for some traffic, multiple remote systems 104 can provide responses and servicing. The service provider 102 can utilize the described systems and methods to monitor the health and responsiveness of the remote systems 104 and direct its outgoing traffic accordingly”) a second LLM (see Buniatyan, [col 4 lines 40-43] “large language models (LLMs) or other machine-learning models, thereby improving overall computational performance of machine-learning systems that utilize embeddings” – there are plurality of LLMs) having a second overall score, (see Rasband, [0056] “because a downtime duration can be more detrimental to the traffic flow than an up and running remote system 104, which may otherwise have a lower responsiveness score”) the first overall score being higher than the second overall score (see Rasband, [0031] “The RSM 212 can monitor the status of traffic to and from a remote system 104, over a period of time, and assign a responsiveness score to the remote system 104. The responsiveness score can be used by the local services 114 to dynamically route traffic to remote systems 104 that have obtained a higher responsiveness score in the recent past”; [0056] “because a downtime duration can be more detrimental to the traffic flow than an up and running remote system 104, which may otherwise have a lower responsiveness score”). The motivation for the proposed combination is maintained.
Claims 6 and 14 are rejected under 35 U.S.C. 103 as being unpatentable over Buniatyan, Lopatenko, Sengupta and Deb further in view of Ramalingam et al. (US 10,133,775 B1, hereinafter “Ramalingam”).
Regarding claim 6, the proposed combination of Buniatyan, Lopatenko, Sengupta and Deb teaches
wherein the associated historical metadata stored in the model discovery database for each pair of a historical query and an LLM includes… (see Sengupta, [col 6 line 59- col 7 line 25] “A history data retrieval component 220 may retrieve conversation history data 222 from a conversation history data store 230 based on user query 102… conversation history data store 230 may be a database or other data storage entity that stores conversation history data, such as prior user queries (e.g., including prior enriched user queries)… associated responses that were generated to the prior user queries… History data retrieval component 220 may retrieve such related conversation history data for user query 102 based on one or more attributes related to user query 102, such as… a software tool that is being used to handle user query 102 and/or one or more previous queries (e.g., a software tool associated with language processing machine learning model 240), other types of metadata, and/or the like, such as based on one or more such attributes… Language processing machine learning model 240 may, for example be a large language model (LLM), such as having a transformer architecture, or may be any suitable type of machine learning model that has been trained to process natural language inputs”) the historical query, (see Buniatyan, [col 22 lines 20-21] “historic queries”) prompt tokens (see Deb, [col 25 lines 54-55] “can determine a number of input tokens within the input or prompt”) of the historical query, (see Buniatyan, [col 22 lines 20-21] “historic queries”) output tokens (see Deb, [col 23 line 62] “number of input or output tokens associated with LLMs”) of the historical response, (see Buniatyan, [col 22 lines 14-15] “The set of query results may be stored in a historic query database”) the historical response, and… (see Buniatyan, [col 22 lines 14-15] “The set of query results may be stored in a historic query database”) the historical query,… (see Buniatyan, [col 22 lines 20-21] “historic queries”) the prompt tokens and (see Deb, [col 25 lines 54-55] “can determine a number of input tokens within the input or prompt”) the output tokens (see Deb, [col 23 line 62] “number of input or output tokens associated with LLMs”).
The proposed combination of Buniatyan, Lopatenko, Sengupta and Deb does not explicitly teach an execution duration of the historical query, a cost of the historical query, the cost being determined based on the prompt tokens and the output tokens.
However, Ramalingam discloses query analysis and teaches
an execution duration of historical query (see Ramalingam, [col 2 line 11] “historical query execution time data”) a cost of the query (see Ramalingam, [col 2 line 10] “query cost data”) the cost being determined based on number of tables and CPU processing capacity prediction (see Ramalingam, [col 2 lines 26-33] “the cost of a data query may describe a proportion (e.g., percentage) of central processing unit (CPU) processing capacity predicted to be employed by the data query while it is executing… the cost may be determined based on a number of tables being joined or otherwise accessed during execution of the data query, the complexity of operations performed by the data query, or the type of operations performed by the data query”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to include the functionality of execution duration, cost and the cost being determined as being disclosed and taught by Ramalingam in the system taught by the proposed combination of Buniatyan, Lopatenko, Sengupta and Deb to yield the predictable results of improving the predictive value of the model (see Ramalingam, [col 9 lines 52-56] “The outlying data point(s) may represent execution instances of the data queries 104 that are sufficiently anomalous or rare, such that their omission from the model generation analysis may improve the predictive value of the model 116.”).
Claim 14 incorporates substantively all the limitations of claim 6 in a method form and is rejected under the same rationale.
Claims 8 and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Buniatyan, Lopatenko, Sengupta and Deb further in view of Edmonds et al. (US 2008/0201348 A1, hereinafter “Edmonds”).
Regarding claim 8, the proposed combination of Buniatyan, Lopatenko, Sengupta and Deb teaches
wherein the instructions that cause the system to, (see Buniatyan, [col 30 lines 50-54] “Such instructions can be read into main memory 405 from another computer-readable medium, such as the storage device 415. Execution of the arrangement of instructions contained in main memory 405 causes the data processing system 105”) for each LLM of the plurality of LLMs, (see Buniatyan, [col 12 lines 43-47] “the machine-learning model(s) 182 include one or more large-language models (LLMs). Large language models include generative models that are pre-trained on large amounts of text and can generate corresponding text in response to an input context”) determine the overall score of the LLM (see Deb, [col 45 lines 59=63] “Once the weights are applied, the system calculates a weighted sum for each AI model, representing its overall compatibility score. The score is a composite measure that reflects how well the model's capabilities align with the expected values across all relevant attributes”; [col 42 lines 24-26] “at least one AI model in the plurality of AI models is a Large Language Model (LLM)”) comprise instructions that cause the system to, for each LLM of the plurality of LLMs: (see Buniatyan, [col 30 lines 50-54] “Such instructions can be read into main memory 405 from another computer-readable medium, such as the storage device 415. Execution of the arrangement of instructions contained in main memory 405 causes the data processing system 105”; [col 4 lines 40-43] “large language models (LLMs) or other machine-learning models, thereby improving overall computational performance of machine-learning systems that utilize embeddings” – there are plurality of LLMs).
… scores of each of the predetermined metrics based on the quality ranks (see Deb, [col 45 lines 35-58] “The system can compare the determined capabilities a first model of the plurality of models with the determined capabilities of a second model of the plurality of models. The system can use a scoring mechanism that assigns a compatibility score to each AI model based on how well its capabilities match the expected values. The scoring mechanism can use weighted criteria to prioritize certain attributes over others, depending on the specific requirements of the request… to rank the AI models, identifying those that best meet the overall requirements of the request. The system can normalize the performance metrics and expected values to a common scale to allow different metrics can be compared and aggregated. The system applies weights to each metric based on the importance of the corresponding attribute. The weights can be predefined based on the type of request or dynamically adjusted based on user preferences or contextual factors. For instance, a weight of 0.7 can be assigned to accuracy and 0.3 can be assigned to latency for a medical diagnosis task, reflecting the higher priority of accuracy”; [col 46 lines 2-4] “the system prioritizes each AI model in the plurality of AI models based on historical performance data of each AI model in the plurality of AI models”; [col 42 lines 24-26] “at least one AI model in the plurality of AI models is a Large Language Model (LLM)”) and the retrieved associated historical metadata for the LLM (see Sengupta, [col 6 line 59- col 7 line 10] “A history data retrieval component 220 may retrieve conversation history data 222 from a conversation history data store 230… may retrieve such related conversation history data for user query 102 based on one or more attributes related to user query 102, such as a conversation identifier, a user identifier, a session identifier… other types of metadata, and/or the like, such as based on one or more such attributes also being associated with conversation history data 222 in conversation history data store 230”) across the identified predetermined number of historical queries; (see Lopatenko, [col 15 line 33 – col 16 line 46] “With a historical query selected for comparison, the selected historical query is compared to the input query (step 5012)… After determining the match score for the selected historical query, a determination is made as to whether there are more historical queries to compare to the input query (step 5014)… each historical query within the defined set of historical queries is compared… comparison is performed on queries within the set until a desired number of historical queries with a match score above a threshold value are found… a fixed number of historical queries can be selected (e.g., the historical queries having the top 2, 3, 4, 5, 10, 15, 20, 50, 100 match scores)… the historical queries with match scores in a certain percentile range can be selected (e.g., select the historical queries in the top 0.1%, 0.5%, 1%, 3%, 5%, 10% of match scores)”).
… of each of the predetermined metrics... (see Deb, [col 45 line 49] “can normalize the performance metrics”) for one or more of the predetermined metrics; and (see Deb, [col 45 line 49] “can normalize the performance metrics”).
determine the overall score of the LLM based on the weighted scores of each of the predetermined metrics (see Deb, [col 45 lines 59=63] “Once the weights are applied, the system calculates a weighted sum for each AI model, representing its overall compatibility score. The score is a composite measure that reflects how well the model's capabilities align with the expected values across all relevant attributes”; [col 42 lines 24-26] “at least one AI model in the plurality of AI models is a Large Language Model (LLM)”).
The proposed combination of Buniatyan, Lopatenko, Sengupta and Deb does not explicitly teach normalize scores, weight the normalized scores, based on user specified sensitivity values.
However, Edmonds discloses implicit metrics and teaches
normalize scores (see Edmonds, [0055] “the system normalizes the tag-based content producer scores”; [0061] “may introduce distortions that review scores and engagement metrics across content type… so each score is computed several times with different normalizations, once for each variant”).
weight the normalized scores… (see Edmonds, [0061] “these scores weighted to produce a final output”) based on user specified sensitivity values (see Edmonds, [0049] “The reviewer rank for the particular document reflects a weighted average of the reviews who submitted concerning the document, where the weighting factors are the reviewer ranks for those reviewers”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to include the functionality of normalize scores, weighing the normalized scores and user specified values being determined as being disclosed and taught by Edmonds in the system taught by the proposed combination of Buniatyan, Lopatenko, Sengupta and Deb to yield the predictable results of providing effective search engine retrieval algorithms based on content quality scores (see Edmonds, [0006 “The tag-mediated content quality scores, which are content quality scores for the subject matter area identified by the associated tags, can be exposed to end-users. The tag-mediated content quality scores and/or a composite content quality score can also be used to complement search engine retrieval algorithms based on traditional keyword-based information retrieval”).
Claim 16 incorporates substantively all the limitations of claim 8 in a method form and is rejected under the same rationale.
Citation of Relevant Prior Art
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
US Publication No. US 2023/0042738 A1 (Smith) teaches a query that is received from a client device and a mode is selected to process the query from a set of possible modes. The possible modes include a fast mode and a low-cost mode. If the fast mode is selected, the query is forwarded to a cloud database to retrieve responsive records.
US Patent No. US 12,099,502 B2 (Nagy et al.) teaches a machine learning model that may determine one or more previously input queries that are similar to the query input by the user device. This determination may be based on comparing the query input by the user device with queries stored in the database. This determination may be based on comparing a query that is semantically similar the query input by the user device with queries stored in the database.
US Patent No. US 12,524,452 B2 (Kranjc et al.) teaches each topic summary stored in the topic summary datastore that includes a corresponding summary of past query-response interactions related to a particular topic that were derived during one or more previous chat sessions between the user and the assistant LLM.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to VAISHALI SHAH whose telephone number is (571)272-8532. The examiner can normally be reached Monday - Friday (7:30 AM to 4:00 PM).
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, AJAY BHATIA can be reached at (571)272-3906. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/VAISHALI SHAH/Primary Examiner, Art Unit 2156