Prosecution Insights
Last updated: October 02, 2026
Application No. 19/202,612

GENERATIVE MODEL BASED DECOMPOSITION OF INPUT QUERY INTO SUB-QUERIES AND GENERATION OF COMPREHENSIVE RESPONSE BASED ON RESPONSES TO SUB-QUERIES

Final Rejection §102§103
Filed
May 08, 2025
Priority
May 12, 2024 — provisional 63/645,915
Examiner
ALMANI, MOHSEN
Art Unit
2159
Tech Center
2100 — Computer Architecture & Software
Assignee
Google LLC
OA Round
2 (Final)
50%
Grant Probability
Moderate
3-4
OA Rounds
2y 8m
Est. Remaining
72%
With Interview

Examiner Intelligence

Grants 50% of resolved cases
50%
Career Allowance Rate
191 granted / 381 resolved
-4.9% vs TC avg
Strong +22% interview lift
Without
With
+21.9%
Interview Lift
resolved cases with interview
Typical timeline
4y 1m
Avg Prosecution
21 currently pending
Career history
411
Total Applications
across all art units

Statute-Specific Performance

§101
13.0%
-27.0% vs TC avg
§103
51.0%
+11.0% vs TC avg
§102
21.5%
-18.5% vs TC avg
§112
10.4%
-29.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 381 resolved cases

Office Action

§102 §103
Detailed Action Applicant amended claims 1, 3, and 11, canceled claim 20, added claim 21 and presented claims 1-19 and 21 for reconsideration on 07/10/2026. Claim Rejections - 35 USC § 102 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 102 that forms the basis for all the rejections under this section made in this Office Action: A person shall be entitled to a patent unless— (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale or otherwise available to the public before the effective filing date of the claimed invention. Claims 1-12, 16-17 and 20 are rejected under 35 U.S.C. 102(a)(2) as being anticipated by Sen et al., Pub. No.: US 2025/0252319 A1 (Sen). Sen discloses: Claim 1. A method implemented by one or more processors, the method comprising: receiving an input query that is generated based on user interface input at a client device; (¶¶ 6-7, “Questions may be a simple one query question, or a more complex question that include multiple parts or requires an explanation of the answer… When the question or request is more complex…the question may be decomposed…The decomposing may include separating the question into multiple parts for a better understanding and to allow focus on each part of the question separately for producing the best results. Prompt design…which is a technique to select the right words and guide the LLM in generating high quality results may also be used”, ¶¶ 41, 101-105, “a question is received...the question may be a simple question with a single ask or query that may be inputted by a user in an input area on a user interface…. the question may be a multipart or complex question that may require multiple answers or answers that are based on multiple factors…the results obtained from analyzing the content and/or context of the question may be used to determine which ELLM to use”) decomposing the input query, decomposing the input query comprising processing the input query, using a first generative model and/or a second generative model, to determine: (a received question is analyzed and each part of a multipart question is processed by selected ELLM and LLMs: ¶¶ 6-7, ¶ 33, “The selected combination of ELLMs and LLMs are used to process a question inputted by a user. Answers obtained by processing the question via the combination of ELLMs and LLMs are blended and used as an input into an ensemble model to obtain a final answer”; ¶¶ 75, 102-110, “the question may be a multipart or complex question that may require multiple answers or answers that are based on multiple factors…the control circuitry 428 and/or 420 may analyze the content and/or context of the question….the control circuitry 428 and/or 420 may use the results from the analysis to determine which ELLM has been trained with data that is relevant to the content and/or context of the question…Analyzing the question may…involve using natural language processing (NLP) techniques…NLP techniques may also be used in conjunction with other tools, such as sentiment analysis tools, to capture the sentiment and mood of the user when asking the question…The question may also be analyzed using artificial intelligence (AI) and machine learning (ML) engines that executive AI and ML algorithms. Such engines and algorithms may provide recommendations based on the question, such as what the question is really seeking, what class or subclass the question fall under, or what enterprise departments are more relevant to the content and context of the question… Since context and content of the question received may apply to more than one topic and as such to more than on ELLM, any ELLMs that is available for use and connected to the topic, either the entire question or a portion of the question, may be included in a set of the narrowed n ELLMS”) a plurality of sub-queries, and for each of the sub-queries, one or more corresponding tools to utilize in processing the sub-query; (see above, “NLP techniques may also be used in conjunction with other tools…Since context and content of the question received may apply to more than one topic and as such to more than on ELLM, any ELLMs that is available for use and connected to the topic, either the entire question or a portion of the question, may be included in a set of the narrowed n ELLMS”) for each of the sub-queries: processing the sub-query, using the one or more corresponding tools for the subquery, to generate one or more corresponding sub-query responses; (see above, ¶¶ 75, 82, 87, 127-128, 106, “the control circuitry 428 and/or 420 may extract data from all the servers… Extracting data may be performed by the control circuitry by using existing techniques such as extract, transform, and load (ETL) techniques and extract, load, transform (ELT) techniques. Data may also be extracted by the control circuitry by using crawlers, scraping software tools, API integration, data mining, database querying, text pattern matching and other types of large data extraction techniques. Data may also be obtained from other existing LLMs and ELLMs…the control circuitry…may obtain answers from all the ELLMs and/or LLMs as part of the sequence and strategy deployed… the control circuitry 428 and/or 420 may blend the answers obtained at block 970. The blending process, in one embodiment, may include selecting portions of answers from different ELLMs and/or LLMs and combining them such that make logical sense”) generating an initial comprehensive response to the input query, generating the initial comprehensive response comprising processing, using the first generative model, the second generative model, or a third generative model, the one or more corresponding sub-query responses for each of the sub-queries; (¶¶ 75, 127-128, “the control circuitry…may obtain answers from all the ELLMs and/or LLMs as part of the sequence and strategy deployed… the control circuitry 428 and/or 420 may blend the answers obtained at block 970. The blending process, in one embodiment, may include selecting portions of answers from different ELLMs and/or LLMs and combining them such that make logical sense) determining whether the initial comprehensive response is responsive to the input query, determining whether the initial comprehensive response is responsive to the input query comprising: processing, using the first generative model, the second generative model, the third generative model, or a fourth generative model, the input query and the initial comprehensive response to generate a critique response that indicates whether the initial comprehensive response is responsive to the input query; and (¶¶ 130-131, generated responses are analyzed by an ensemble model for selecting a golden response: “the control circuitry…may input the blended answer into an ensemble model…the ensemble model may be a ruled-based model that has its own rules on how to determine the better answer…the control circuitry…may obtain a golden answer from the ensemble model and display it to the user from whom the question was received”) determining, based on the critique response, that the initial comprehensive response is not responsive to the input query; (initial results are not satisfactory; the results are refined by feeding them, as a further sub-query, into different models each using required tools for obtaining a refined result: ¶¶ 75, 123, “the control circuitry 428 and/or 420 may simultaneously feed the question received to a plurality of ELLMs and LLMs, use the results obtained from the ELLMs and LLMs, and refine the results by feeding them into a second set of ELLMs and LLMs to obtain a more refined answer”, ¶¶ 130-131, “The ensemble model may continuously learn from the parameters as they are added and the continuous learning may allow the ensemble model to further refine its ability to make predictions of the next work and accordingly determine a better answer for the question… golden answer may refer to an answer that has been refined one or more times through process 900. An iterative method may be used by the control circuitry 428 and/or 420 to refine the answer obtained from a sequence of ELLMs and/or LLMs”) in response to determining that the initial comprehensive response is not responsive to the input query: generating a refined comprehensive response that is based on a further sub-query response, the further sub-query response being generated based on the critique response; and causing the refined comprehensive response to be rendered at the client device as responsive to the input query. (obtained results are not satisfactory; the results are refined by feeding them, as a further sub-query, into different models using t for obtaining a refined result: ¶¶ 75, 82, 87, 123, “the control circuitry 428 and/or 420 may simultaneously feed the question received to a plurality of ELLMs and LLMs, use the results obtained from the ELLMs and LLMs, and refine the results by feeding them into a second set of ELLMs and LLMs to obtain a more refined answer”, ¶¶ 130-131, “The ensemble model may continuously learn from the parameters as they are added and the continuous learning may allow the ensemble model to further refine its ability to make predictions of the next work and accordingly determine a better answer for the question… golden answer may refer to an answer that has been refined one or more times through process 900. An iterative method may be used by the control circuitry 428 and/or 420 to refine the answer obtained from a sequence of ELLMs and/or LLMs”; ¶ 140, “a sample or preferred answer may also be inputted into the ensemble model along with the blended answer and the question 1010. The sample answer may be used as training data for the ensemble model to further refine the golden answer”; ¶¶ 106, 167, “The process of using one or more LLMs and ELLMs and using the results from one to process it again using a second LLMs and/or ELLMs or set second set of LLMs and ELLMs may result in the answer continuously getting refined based on additional information gathered along the process and use of different types of training data used in each separate LLMs and ELLMs to answer the question. Such iterative process may be predetermined or may be dynamically determined as results from each LLMs and ELLMs are analyzed”) Claim 21. A system, comprising: one or more processors; and memory, the memory storing computer readable instructions that, when executed by the one or more processors, cause the one or more processors to: receive an input query that is generated based on user interface input at a client device; (¶¶ 6-7, “Questions may be a simple one query question, or a more complex question that include multiple parts or requires an explanation of the answer… When the question or request is more complex…the question may be decomposed…The decomposing may include separating the question into multiple parts for a better understanding and to allow focus on each part of the question separately for producing the best results. Prompt design…which is a technique to select the right words and guide the LLM in generating high quality results may also be used”, ¶¶ 41, 101-105, “a question is received...the question may be a simple question with a single ask or query that may be inputted by a user in an input area on a user interface…. the question may be a multipart or complex question that may require multiple answers or answers that are based on multiple factors…the results obtained from analyzing the content and/or context of the question may be used to determine which ELLM to use”) decompose the input query, decomposing the input query comprising processing the input query, using a first generative model and/or a second generative model, to determine: (a received question is analyzed and each part of a multipart question is processed by selected ELLM and LLMs: ¶¶ 6-7, ¶ 33, “The selected combination of ELLMs and LLMs are used to process a question inputted by a user. Answers obtained by processing the question via the combination of ELLMs and LLMs are blended and used as an input into an ensemble model to obtain a final answer”; ¶¶ 75, 102-110, “the question may be a multipart or complex question that may require multiple answers or answers that are based on multiple factors…the control circuitry 428 and/or 420 may analyze the content and/or context of the question….the control circuitry 428 and/or 420 may use the results from the analysis to determine which ELLM has been trained with data that is relevant to the content and/or context of the question…Analyzing the question may…involve using natural language processing (NLP) techniques…NLP techniques may also be used in conjunction with other tools, such as sentiment analysis tools, to capture the sentiment and mood of the user when asking the question…The question may also be analyzed using artificial intelligence (AI) and machine learning (ML) engines that executive AI and ML algorithms. Such engines and algorithms may provide recommendations based on the question, such as what the question is really seeking, what class or subclass the question fall under, or what enterprise departments are more relevant to the content and context of the question… Since context and content of the question received may apply to more than one topic and as such to more than on ELLM, any ELLMs that is available for use and connected to the topic, either the entire question or a portion of the question, may be included in a set of the narrowed n ELLMS”) a plurality of sub-queries, and for each of the sub-queries, one or more corresponding tools to utilize in processing the sub-query; (see above, “NLP techniques may also be used in conjunction with other tools…Since context and content of the question received may apply to more than one topic and as such to more than on ELLM, any ELLMs that is available for use and connected to the topic, either the entire question or a portion of the question, may be included in a set of the narrowed n ELLMS”) for each of the sub-queries: process the sub-query, using the one or more corresponding tools for the subquery, to generate one or more corresponding sub-query responses; (see above, ¶¶ 75, 82, 87, 127-128, 106, “the control circuitry 428 and/or 420 may extract data from all the servers… Extracting data may be performed by the control circuitry by using existing techniques such as extract, transform, and load (ETL) techniques and extract, load, transform (ELT) techniques. Data may also be extracted by the control circuitry by using crawlers, scraping software tools, API integration, data mining, database querying, text pattern matching and other types of large data extraction techniques. Data may also be obtained from other existing LLMs and ELLMs…the control circuitry…may obtain answers from all the ELLMs and/or LLMs as part of the sequence and strategy deployed… the control circuitry 428 and/or 420 may blend the answers obtained at block 970. The blending process, in one embodiment, may include selecting portions of answers from different ELLMs and/or LLMs and combining them such that make logical sense”) generate an initial comprehensive response to the input query, generating the initial comprehensive response comprising processing, using the first generative model, the second generative model, or a third generative model, the one or more corresponding sub-query responses for each of the sub-queries; (¶¶ 75, 127-128, “the control circuitry…may obtain answers from all the ELLMs and/or LLMs as part of the sequence and strategy deployed… the control circuitry 428 and/or 420 may blend the answers obtained at block 970. The blending process, in one embodiment, may include selecting portions of answers from different ELLMs and/or LLMs and combining them such that make logical sense) determine whether the initial comprehensive response is responsive to the input query, determining whether the initial comprehensive response is responsive to the input query comprising: process, using the first generative model, the second generative model, the third generative model, or a fourth generative model, the input query and the initial comprehensive response to generate a critique response that indicates whether the initial comprehensive response is responsive to the input query; and (¶¶ 130-131, generated responses are analyzed by an ensemble model for selecting a golden response: “the control circuitry…may input the blended answer into an ensemble model…the ensemble model may be a ruled-based model that has its own rules on how to determine the better answer…the control circuitry…may obtain a golden answer from the ensemble model and display it to the user from whom the question was received”) determine, based on the critique response, whether the initial comprehensive response is responsive to the input query; (¶¶ 130-131, generated responses are analyzed by an ensemble model for selecting a golden response: “the control circuitry…may input the blended answer into an ensemble model…the ensemble model may be a ruled-based model that has its own rules on how to determine the better answer…the control circuitry…may obtain a golden answer from the ensemble model and display it to the user from whom the question was received”) in response to determining that the initial comprehensive response is responsive to the input query: cause the initial comprehensive response to be rendered at the client device as responsive to the input query; and (¶¶ 130-131, generated responses are analyzed by an ensemble model for selecting a golden response: “the control circuitry…may input the blended answer into an ensemble model…the ensemble model may be a ruled-based model that has its own rules on how to determine the better answer…the control circuitry…may obtain a golden answer from the ensemble model and display it to the user from whom the question was received”) in response to determining that the initial comprehensive response is not responsive to the input query: generate a refined comprehensive response that is based on a further sub-query response, the further sub-query response being generated based on the critique response; and cause the refined comprehensive response to be rendered at the client device as responsive to the input query. (obtained results are not satisfactory; the results are refined by feeding them, as a further sub-query, into different models using t for obtaining a refined result: ¶¶ 75, 82, 87, 123, “the control circuitry 428 and/or 420 may simultaneously feed the question received to a plurality of ELLMs and LLMs, use the results obtained from the ELLMs and LLMs, and refine the results by feeding them into a second set of ELLMs and LLMs to obtain a more refined answer”, ¶¶ 130-131, “The ensemble model may continuously learn from the parameters as they are added and the continuous learning may allow the ensemble model to further refine its ability to make predictions of the next work and accordingly determine a better answer for the question… golden answer may refer to an answer that has been refined one or more times through process 900. An iterative method may be used by the control circuitry 428 and/or 420 to refine the answer obtained from a sequence of ELLMs and/or LLMs”; ¶ 140, “a sample or preferred answer may also be inputted into the ensemble model along with the blended answer and the question 1010. The sample answer may be used as training data for the ensemble model to further refine the golden answer”; ¶¶ 106, 167, “The process of using one or more LLMs and ELLMs and using the results from one to process it again using a second LLMs and/or ELLMs or set second set of LLMs and ELLMs may result in the answer continuously getting refined based on additional information gathered along the process and use of different types of training data used in each separate LLMs and ELLMs to answer the question. Such iterative process may be predetermined or may be dynamically determined as results from each LLMs and ELLMs are analyzed”) Claim 2. The method of claim 1, wherein generating the refined comprehensive response, that is based on the further sub-query response, comprises: determining, based on the critique response, a further sub-query and one or more further tools to utilize in processing the further sub-query; processing the further sub-query, using the one or more further tools for the further subquery, to generate one or more further sub-query responses; generating the refined comprehensive response based on processing, using the first generative model, the second generative model, or the third generative model, the one or more further subquery responses and the initial comprehensive response or the one or more corresponding sub-query responses for each of the sub-queries. (obtained results are not satisfactory; the results are refined by feeding them, as a further sub-query, into different ELLMs and LLMs using required tools for obtaining a refined result: ¶¶ 75, 123, “the control circuitry 428 and/or 420 may simultaneously feed the question received to a plurality of ELLMs and LLMs, use the results obtained from the ELLMs and LLMs, and refine the results by feeding them into a second set of ELLMs and LLMs to obtain a more refined answer”, ¶¶ 130-131, “The ensemble model may continuously learn from the parameters as they are added and the continuous learning may allow the ensemble model to further refine its ability to make predictions of the next work and accordingly determine a better answer for the question… golden answer may refer to an answer that has been refined one or more times through process 900. An iterative method may be used by the control circuitry 428 and/or 420 to refine the answer obtained from a sequence of ELLMs and/or LLMs”; ¶ 140, “a sample or preferred answer may also be inputted into the ensemble model along with the blended answer and the question 1010. The sample answer may be used as training data for the ensemble model to further refine the golden answer”; ¶ 167, “The process of using one or more LLMs and ELLMs and using the results from one to process it again using a second LLMs and/or ELLMs or set second set of LLMs and ELLMs may result in the answer continuously getting refined based on additional information gathered along the process and use of different types of training data used in each separate LLMs and ELLMs to answer the question. Such iterative process may be predetermined or may be dynamically determined as results from each LLMs and ELLMs are analyzed”) Claim 3. The method of claim 2, further comprising: determining whether the refined comprehensive response is responsive to the input query, determining whether the refined comprehensive response is responsive to the input query comprising: processing, using the first generative model, the second generative model, the third generative model, or the fourth generative model, the input query and the refined comprehensive response to generate an additional critique response that indicates whether the refined comprehensive response is responsive to the input query; and determining, based on the critique response, whether the refined comprehensive response is responsive to the input query; wherein causing the refined comprehensive response to be rendered at the client device as responsive to the input query is in response to determining that the refined comprehensive response is responsive to the input query. (refining a response is an iterative process of evaluating a response with respect to the input query: ¶¶ 75, 123, “the control circuitry 428 and/or 420 may simultaneously feed the question received to a plurality of ELLMs and LLMs, use the results obtained from the ELLMs and LLMs, and refine the results by feeding them into a second set of ELLMs and LLMs to obtain a more refined answer”, ¶¶ 130-131, “The ensemble model may continuously learn from the parameters as they are added and the continuous learning may allow the ensemble model to further refine its ability to make predictions of the next work and accordingly determine a better answer for the question… golden answer may refer to an answer that has been refined one or more times through process 900. An iterative method may be used by the control circuitry 428 and/or 420 to refine the answer obtained from a sequence of ELLMs and/or LLMs”; ¶ 140, “a sample or preferred answer may also be inputted into the ensemble model along with the blended answer and the question 1010. The sample answer may be used as training data for the ensemble model to further refine the golden answer”; ¶ 167, “The process of using one or more LLMs and ELLMs and using the results from one to process it again using a second LLMs and/or ELLMs or set second set of LLMs and ELLMs may result in the answer continuously getting refined based on additional information gathered along the process and use of different types of training data used in each separate LLMs and ELLMs to answer the question. Such iterative process may be predetermined or may be dynamically determined as results from each LLMs and ELLMs are analyzed”) Claim 4. The method of claim 2, wherein the critique response directly indicates one or both of the further sub-query and the one or more further tools to utilize in processing the further sub-query. (An unsatisfactory result is provided, as a generated further sub-query, into different models each using required tools for obtaining a refined result: ¶¶ 75, 82, 87, 123, “the control circuitry 428 and/or 420 may simultaneously feed the question received to a plurality of ELLMs and LLMs, use the results obtained from the ELLMs and LLMs, and refine the results by feeding them into a second set of ELLMs and LLMs to obtain a more refined answer”, ¶¶ 130-131, “The ensemble model may continuously learn from the parameters as they are added and the continuous learning may allow the ensemble model to further refine its ability to make predictions of the next work and accordingly determine a better answer for the question… golden answer may refer to an answer that has been refined one or more times through process 900. An iterative method may be used by the control circuitry 428 and/or 420 to refine the answer obtained from a sequence of ELLMs and/or LLMs”) Claim 5. The method of claim 2, further comprising: determining, based on processing the critique response, the further sub-query and the one or more further tools to utilize in processing the further sub-query. (an unsatisfactory result is provided, as a generated further sub-query, into different models each using required tools for obtaining a refined result: ¶¶ 75, 82, 87, 123, “the control circuitry 428 and/or 420 may simultaneously feed the question received to a plurality of ELLMs and LLMs, use the results obtained from the ELLMs and LLMs, and refine the results by feeding them into a second set of ELLMs and LLMs to obtain a more refined answer”, ¶¶ 130-131, “The ensemble model may continuously learn from the parameters as they are added and the continuous learning may allow the ensemble model to further refine its ability to make predictions of the next work and accordingly determine a better answer for the question… golden answer may refer to an answer that has been refined one or more times through process 900. An iterative method may be used by the control circuitry 428 and/or 420 to refine the answer obtained from a sequence of ELLMs and/or LLMs”) Claim 6. The method of claim 1, wherein processing, using the first generative model, the second generative model, or the third generative model, the input query and the initial comprehensive response to generate the critique response that indicates whether the initial comprehensive response is responsive to the input query further comprises: processing, using the first generative model, the second generative model, or the third generative model, and along with the input query and the initial comprehensive response: each of the sub-queries, each of the corresponding tools utilized in processing the sub-queries, and/or each of the corresponding sub-query responses. (see claim 1, ¶¶ 130-131, “the control circuitry…may input the blended answer into an ensemble model…the ensemble model may be a ruled-based model that has its own rules on how to determine the better answer…the control circuitry…may obtain a golden answer from the ensemble model and display it to the user from whom the question was received”) Claim 7. The method of claim 1, further comprising: determining whether to provide an LLM-only response to the input query in lieu of a comprehensive response; wherein generating the initial comprehensive response is performed responsive to determining to not provide the LLM-only response to the input query. (¶ 116, a simple question is answered using an LLM: “A user may be willing to live with an average or low level of accuracy for simple questions, such as a middle school math problem or for writing a thank you letter and desire a high level of accuracy for critical problems. For example, if the question posed to the LLM is desiring to seek a solution that would impact a company's sales, a job prospect, debugging of a bug in a critical software, then the user may desire a higher level of accuracy and be willing to pay for the higher level of accuracy. Since accuracy may relate to computing power, e.g., a higher level of accuracy for a complex problem requiring higher usage of computational resources and thereby incurring more costs, the user may reserve a higher level of accuracy for more important and critical tasks. Accordingly, in an example where accuracy parameters are described and the question is presented to LLMs such as ChatGPT™, Bard™, Llama™, Bing chat™, Claude™, and Jasper™, the system may narrow the selection of the LLMs based on the accuracy parameters”) Claim 8. The method of claim 7, further comprising: prior to generating the initial comprehensive response: generating the LLM-only response based on processing, in a single LLM pass, an LLM prompt that is based on the input query; wherein determining whether to provide the LLM-only response to the input query in lieu of the comprehensive response comprises processing the LLM-only response. (a simple question is answered using an LLM while answering a complex question requires using and blending multiple LLMs generated answers: ¶¶ 52, 116, “A user may be willing to live with an average or low level of accuracy for simple questions, such as a middle school math problem or for writing a thank you letter and desire a high level of accuracy for critical problems. For example, if the question posed to the LLM is desiring to seek a solution that would impact a company's sales, a job prospect, debugging of a bug in a critical software, then the user may desire a higher level of accuracy and be willing to pay for the higher level of accuracy. Since accuracy may relate to computing power, e.g., a higher level of accuracy for a complex problem requiring higher usage of computational resources and thereby incurring more costs, the user may reserve a higher level of accuracy for more important and critical tasks. Accordingly, in an example where accuracy parameters are described and the question is presented to LLMs such as ChatGPT™, Bard™, Llama™, Bing chat™, Claude™, and Jasper™, the system may narrow the selection of the LLMs based on the accuracy parameters”; ¶¶ 127- 128, “the control circuitry 428 and/or 420 may obtain answers from all the ELLMs and/or LLMs…The blending process, in one embodiment, may include selecting portions of answers from different ELLMs and/or LLMs and combining them such that make logical sense”) Claim 9. The method of claim 7, wherein processing the LLM-only response in determining whether to provide the LLM-only response to the input query in lieu of the comprehensive response comprises: processing, using the first generative model, the second generative model, the third generative model, or the fourth generative model, the input query and the LLM-only response to generate an initial critique response that indicates whether the LLM-only response is responsive to the input query; and determining, based on the initial critique response, whether to provide the LLM-only response to the input query in lieu of the comprehensive response. (A generated result is evaluated for refining the result by selecting LLMs: ¶¶ 52, 116, “A user may be willing to live with an average or low level of accuracy for simple questions, such as a middle school math problem or for writing a thank you letter and desire a high level of accuracy for critical problems. For example, if the question posed to the LLM is desiring to seek a solution that would impact a company's sales, a job prospect, debugging of a bug in a critical software, then the user may desire a higher level of accuracy and be willing to pay for the higher level of accuracy. Since accuracy may relate to computing power, e.g., a higher level of accuracy for a complex problem requiring higher usage of computational resources and thereby incurring more costs, the user may reserve a higher level of accuracy for more important and critical tasks. Accordingly, in an example where accuracy parameters are described and the question is presented to LLMs such as ChatGPT™, Bard™, Llama™, Bing chat™, Claude™, and Jasper™, the system may narrow the selection of the LLMs based on the accuracy parameters”; ¶ 123, “the control circuitry 428 and/or 420 may simultaneously feed the question received to a plurality of ELLMs and LLMs, use the results obtained from the ELLMs and LLMs, and refine the results by feeding them into a second set of ELLMs and LLMs to obtain a more refined answer”, ¶¶ 130-131, “The ensemble model may continuously learn from the parameters as they are added and the continuous learning may allow the ensemble model to further refine its ability to make predictions of the next work and accordingly determine a better answer for the question”) Claim 10. The method of claim 9, wherein processing, using the first generative model, the second generative model, the third generative model, or the fourth generative model, the input query and the LLM-only response to generate the initial critique response further comprises: processing, using the first generative model, the second generative model, the third generative model, or the fourth generative model, and along with the input query and the LLM-only response: each of the sub-queries, and/or each of the corresponding tools utilized in processing the sub-queries. (A generated result is evaluated for refining the result by selecting LLMs: ¶¶ 52, 116, “A user may be willing to live with an average or low level of accuracy for simple questions, such as a middle school math problem or for writing a thank you letter and desire a high level of accuracy for critical problems. For example, if the question posed to the LLM is desiring to seek a solution that would impact a company's sales, a job prospect, debugging of a bug in a critical software, then the user may desire a higher level of accuracy and be willing to pay for the higher level of accuracy. Since accuracy may relate to computing power, e.g., a higher level of accuracy for a complex problem requiring higher usage of computational resources and thereby incurring more costs, the user may reserve a higher level of accuracy for more important and critical tasks. Accordingly, in an example where accuracy parameters are described and the question is presented to LLMs such as ChatGPT™, Bard™, Llama™, Bing chat™, Claude™, and Jasper™, the system may narrow the selection of the LLMs based on the accuracy parameters”; ¶ 123, “the control circuitry 428 and/or 420 may simultaneously feed the question received to a plurality of ELLMs and LLMs, use the results obtained from the ELLMs and LLMs, and refine the results by feeding them into a second set of ELLMs and LLMs to obtain a more refined answer”, ¶¶ 130-131, “The ensemble model may continuously learn from the parameters as they are added and the continuous learning may allow the ensemble model to further refine its ability to make predictions of the next work and accordingly determine a better answer for the question”) Claim 16. The method of claim 1, wherein the plurality of sub-queries include a first sub-query and a second sub-query that is distinct from the first sub-query and wherein the corresponding tools include a first tool to utilize in processing the first sub-query and a second tool, that is distinct from the first tool, to utilize in processing the second sub-query. (¶¶ 75, 82, 87, 122-125, 127-128, 106, a complex query is processed by multiple models each using a corresponding tool and wherein a response generated by a model is given as input to another model: “the control circuitry 428 and/or 420 may determine which ELLMs and LLMs, from the set of n ELLMs and LLMs, to use first and then use the results from such ELLMs and LLMs to feed into a second set of ELLMs and LLMs, from the set of n ELLMs and LLMs”) Claim 17. The method of claim 1, wherein the plurality of sub-queries include a first sub-query and a second sub-query that is distinct from the first sub-query and that is conditioned on the corresponding sub-query response generated based on the first sub-query. (¶¶ 75, 82, 87, 122-125, 127-128, 106, a complex query is processed by multiple models each using a corresponding tool and wherein a response generated by a model is given as input to another model: “the control circuitry 428 and/or 420 may determine which ELLMs and LLMs, from the set of n ELLMs and LLMs, to use first and then use the results from such ELLMs and LLMs to feed into a second set of ELLMs and LLMs, from the set of n ELLMs and LLMs”) Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 11-15 and 18-19 are rejected under 35 U.S.C. 103(a) as being unpatentable over Sen as applied to above claims 1 above in view of Kim et al., Pub. No.: US 2025/0007870 A1 (Kim). Claim 11. Sen thought the method of claim 1; Sen did not specifically teach but Kim teaches: prior to processing each of the sub-queries to generate the corresponding sub-query responses: causing a prompt to be rendered, at the client device, that characterizes the plurality of sub-queries; and determining that affirmative user interface input is received responsive to the prompt; wherein processing each of the sub-queries to generate the corresponding sub-prompt responses is contingent on receiving the affirmative user interface input. (Kim, wherein a user query is divided into suggested questions and wherein a prompt is generated for the user to select a suggested question: ¶ 332, “in some implementations, the natural language user input is divided into multiple sub-parts or portions, each portion used to generate a separate prompt. …In some cases, natural language processing is performed on the user input to identify potentially divisible requests that may be serviced using separate prompts…the multiple requests or prompts are dependent such that the result of one prompt is used to generate another prompt” and ¶¶ 369-371, “FIG. 31A depicts an example graphical user interface 3100 a that includes an initial user-generated request message 3102 including a natural language user input provided to a chat interface of a messaging platform…the interface 3100 b includes a response 3120 that includes a computed completeness score 3233 based on the user input of the initial request message 3120…If the user…selecting one or more of the proposed questions, the sequence may be repeated until a satisfactory question completeness score 3122 is achieved. In some embodiments, the system generates the series of suggested questions or proposed queries 3120 regardless of the completeness score such that the user can review additional questions even when the initial request may be predicted to be sufficiently complete to obtain a successful resolution”) Sen ¶ 6 discloses that “When the question or request is more complex…the question may be decomposed at 202. The decomposing may include separating the question into multiple parts for a better understanding and to allow focus on each part of the question separately for producing the best results”. It would have been obvious before the effective filling date of the claimed invention to a person having ordinary skill in the art to combine the applied references for disclosing prior to processing each of the sub-queries to generate the corresponding sub-query responses: causing a prompt to be rendered, at the client device, that characterizes the plurality of sub-queries; and determining that affirmative user interface input is received responsive to the prompt; wherein processing each of the sub-queries to generate the corresponding sub-prompt responses is contingent on receiving the affirmative user interface input because doing so would further increase usability of Sen by utilizing a user feedback for better understating the user question for generating a successful result. Claim 12. The method of claim 11, wherein the prompt further characterizes the corresponding tools for the plurality of sub-queries. (Kim, wherein, ¶ 332, “in some implementations, the natural language user input is divided into multiple sub-parts or portions, each portion used to generate a separate prompt. …In some cases, natural language processing is performed on the user input to identify potentially divisible requests that may be serviced using separate prompts…the multiple requests or prompts are dependent such that the result of one prompt is used to generate another prompt” and wherein ¶ 372, “the automated chat service may conduct a search of a knowledge base or other content store using the original user input or a modified user input resulting from an exchange similar to as described above with respect to FIG. 31B” and Sen, ¶¶ 50, 55, wherein “in response to determining that the user who has inputted the query in the prompt is an employee, the question relates to enterprise finance, and that the employee has authorization to receive confidential data that is below a level 6 (on a 1-10 scale where 10 may be the most confidential date), then automatically determining a strategy to use a combination of ELLMs that are finance related and include confidential data below the level 6…computing device 418 may receive a user input like a question, query, or task to answer a math question, to perform algorithm testing to detect any bugs, determine financial projections for an enterprise, etc.” suggests characterizing a corresponding tool for suggested multiple sub-parts or portions) Claim 13. The method of claim 1, further comprising: prior to or while processing each of the sub-queries to generate the corresponding subquery responses: causing a notification to be rendered, at the client device, that characterizes that there will be a time delay before a comprehensive response is provided. (Kim, Fig. 31A wherein “Hang tight as I search the knowledge base for relevant content and help find the answer. This can sometimes take a few seconds” suggests providing a time delay notification to the user) Claim 14. The method of claim 13, wherein the notification further characterizes an anticipated duration of the time delay. (Kim, fig. 31A, wherein “Hang tight as I search the knowledge base for relevant content and help find the answer. This can sometimes take a few seconds” suggests providing an anticipated duration of the time delay) Claim 15. The method of claim 14, further comprising: determining the anticipated duration of the time delay as a function of at least one of the corresponding tools. (Kim, fig. 31A, wherein “Hang tight as I search the knowledge base for relevant content and help find the answer. This can sometimes take a few seconds” suggests providing a time delay with respect to using search tool for searching a knowledge base) Claim 18. The method of claim 1, wherein receiving the input query comprises: receiving an initial input query that is generated based on initial user interface input at the client device; causing a specification prompt to be provided, at the client device, requesting further specification of the initial input query; receiving a refinement of the initial input query that is based on further user interface input provided responsive to the specification prompt; and generating the input query based on the refinement and the initial input query. (Kim, ¶¶ 369-371, a specification prompt is provided to the user: “FIG. 31A depicts an example graphical user interface 3100 a that includes an initial user-generated request message 3102 including a natural language user input provided to a chat interface of a messaging platform…the interface 3100 b includes a response 3120 that includes a computed completeness score 3233 based on the user input of the initial request message 3120…If the user modifies the request manually or selecting one or more of the proposed questions, the sequence may be repeated until a satisfactory question completeness score 3122 is achieved. In some embodiments, the system generates the series of suggested questions or proposed queries 3120 regardless of the completeness score such that the user can review additional questions even when the initial request may be predicted to be sufficiently complete to obtain a successful resolution”) Claim 19. The method of claim 18, further comprising: determining, based on processing the initial input query, to provide the specification prompt; wherein causing the specification prompt to be provided is in response to determining, based on processing the initial input query, to provide the specification prompt. (Kim, ¶¶ 369-371, based on processing an initial user query for completeness, a specification prompt is provided to the user: “FIG. 31A depicts an example graphical user interface 3100 a that includes an initial user-generated request message 3102 including a natural language user input provided to a chat interface of a messaging platform…the interface 3100 b includes a response 3120 that includes a computed completeness score 3233 based on the user input of the initial request message 3120…If the user modifies the request manually or selecting one or more of the proposed questions, the sequence may be repeated until a satisfactory question completeness score 3122 is achieved. In some embodiments, the system generates the series of suggested questions or proposed queries 3120 regardless of the completeness score such that the user can review additional questions even when the initial request may be predicted to be sufficiently complete to obtain a successful resolution”) Response to Amendment and Arguments In light of amendments, claim objections are withdrawn. Applicant’s arguments with respect to amended claims have been considered but are not persuasive because the applied references disclosed the amended feature as shown above. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. It is suggested that the Applicant review these documents before submitting any amendments. THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any extension fee pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to Mohsen Almani whose telephone number is (571)270-7722. The examiner can normally be reached on M-F, 9 AM-5 PM, ET. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Ann J. Lo can be reached on 571-272-9767. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see https://ppair-my.uspto.gov/pair/PrivatePair. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /MOHSEN ALMANI/Primary Examiner, Art Unit 2159
Read full office action

Prosecution Timeline

May 08, 2025
Application Filed
Apr 09, 2026
Non-Final Rejection mailed — §102, §103
Jul 10, 2026
Response Filed
Sep 23, 2026
Final Rejection mailed — §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12718914
DATABASE RECORD LINKAGE USING ADAPTIVE DYNAMIC BLOCKING
3y 5m to grant Granted Aug 25, 2026
Patent 12705239
SYSTEM AND METHOD FOR GENERATING WEIGHTED QUERY REPRESENTATIONS FOR ENHANCED RETRIEVAL AUGMENTED GENERATION
2y 2m to grant Granted Aug 11, 2026
Patent 12699725
HIERARCHICAL DICTIONARY WITH STATISTICAL FILTERING BASED ON WORD FREQUENCY
1y 9m to grant Granted Aug 04, 2026
Patent 12675438
CROSS-SILO DATA STORAGE AND DEDUPLICATION
1y 10m to grant Granted Jul 07, 2026
Patent 12657216
SCALABLE INDEXING ARCHITECTURE
6y 9m to grant Granted Jun 16, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
50%
Grant Probability
72%
With Interview (+21.9%)
4y 1m (~2y 8m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 381 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month