DETAILED ACTION
Application No. 19/294,444 filed on 08/08/2025 has been examined. In this Office Action, claims 1-20 are pending.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 10/08/25, 02/26/26, 04/03/26, 06/05/26, 06/12/26 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or
composition of matter, or any new and useful improvement thereof, may obtain a patent
therefor, subject to the conditions and requirements of this title.
Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to a judicial exception (i.e., a law of nature, a natural phenomenon, or an abstract idea) without significantly more.
Based upon consideration of all of the relevant factors with respect to the claims as a whole, claims 1-20 are determined to be directed to an abstract idea and not significantly more than the abstract idea itself. The rationale for this determination is explained below:
At Step 1:
Regarding with independent claims 1, 8 recite receive an indication of negative feedback for a first query language query generated by a generative artificial intelligence (AI) model based on a first natural language query, generate a corrected pair based on the indication and the first natural language query, the corrected pair comprising the first natural language query and a second query language query, the second query language query a syntactically valid conversion of the first natural language query, and cause a language conversion engine to utilize the corrected pair to convert a second natural language query to a third query language query, and independent claim 15 recites receive a request to convert a first natural language query to a first query language query, generate a first prompt based on the first natural language query and synthetic data, the synthetic data comprising a first pair comprising a second natural language query and a second query language query, utilize a generative artificial intelligence (Al) model to convert the first natural language query to the first query language query based at least on the first prompt, receive an indication of negative feedback for the first query language query, cause a synthetic data generator to update the synthetic data based at least on the indication of negative feedback, resulting in updated synthetic data comprising a second pair comprising the second natural language query and a third query language query, and utilize the updated synthetic data to generate a second prompt based on a third natural language query, the second prompt structured to cause the generative Al model to convert the third natural language query to a fourth natural language query based at least on the updated synthetic data.
At Step 2A, Prong One:
Claims 1, 8 recite the following limitations directed to an abstract idea:
“receive an indication of negative feedback for a first query language query generated by a generative artificial intelligence (AI) model based on a first natural language query,”
“generate a corrected pair based on the indication and the first natural language query, the corrected pair comprising the first natural language query and a second query language query, the second query language query a syntactically valid conversion of the first natural language query,”
and “cause a language conversion engine to utilize the corrected pair to convert a second natural language query to a third query language query.”
The limitations of receiving negative feedback, evaluating a first natural language query and a first query language query, generating a corrected pair in response to the negative feedback, and determining a syntactically valid conversion recite concepts of evaluation, judgment, and correction of information. Such concepts can be practically performed in the human mind, for example, by reviewing a natural-language question and a corresponding query-language expression, determining that the query-language expression is incorrect, and preparing a corrected query-language expression corresponding to the natural-language question. Accordingly, these limitations recite a mental process, which is an abstract idea under MPEP § 2106.04(a)(2). The USPTO identifies observations, evaluations, judgments, and opinions that can practically be performed in the human mind as mental processes.
Claim 15 recites the following limitations directed to an abstract idea:
The limitation ““receive a request to convert a first natural language query to a first query language query,” “generate a first prompt based on the first natural language query and synthetic data, the synthetic data comprising a first pair comprising a second natural language query and a second query language query,” “utilize a generative artificial intelligence (AI) model to convert the first natural language query to the first query language query based at least on the first prompt,”
“receive an indication of negative feedback for the first query language query,”
“cause a synthetic data generator to update the synthetic data based at least on the indication of negative feedback, resulting in updated synthetic data comprising a second pair comprising the second natural language query and a third query language query,”
and “utilize the updated synthetic data to generate a second prompt based on a third natural language query, the second prompt structured to cause the generative AI model to convert the third natural language query to a fourth natural language query based at least on the updated synthetic data.”
These limitations recite the abstract idea of a mental process, including receiving information, converting linguistic information, formulating instructions based on example information, receiving evaluative feedback, correcting associated information based on the feedback, and using the corrected information in a subsequent linguistic conversion.
A person can receive a request to convert a natural-language query, consider an example natural-language/query-language pair, formulate a corresponding query, receive feedback that the query is incorrect, correct the associated pair, and use the corrected information when converting another query.
In particular, “convert the first natural language query to the first query language query” can be performed by a person familiar with the applicable query language by interpreting the natural-language request and writing the corresponding query.
Likewise, “convert the third natural language query to a fourth natural language query” can practically be performed in the human mind by interpreting, rewriting, paraphrasing, or otherwise converting one natural-language query into another natural-language query.
Accordingly, the recited receiving, conversion, prompting, feedback evaluation, correction, updating, and subsequent conversion operations fall within the mental-process grouping of abstract ideas.
Under Step 2A, Prong Two, claims 1, 8, 15 additionally recites:
“a processor circuit;” “a memory device that stores program code structured to cause the processor circuit to:” a “generative artificial intelligence (AI) model,” and “a language conversion engine.”, and “a synthetic data generator”. These additional elements use computer components to perform the recited evaluation, correction, storage, and conversion of information. The claim does not recite a particular improvement to the processor circuit, memory device, architecture of the generative AI model, or architecture of the language conversion engine or the synthetic data generator. Rather, the computer components are used as tools to carry out the recited abstract information-processing operations. Therefore, the claim does not integrate the judicial exception into a practical application.
Under Step 2B, the additional elements, considered individually and as an ordered combination, do not add significantly more than the abstract idea because they perform their ordinary functions of processing, storing, generating, and converting information.
Accordingly, claims 1, 8 and 15 are directed to patent-ineligible subject matter.
Regarding claims 2, 9 further recite “determine the corrected pair satisfies criterion of a synthetic data store; and “store the corrected pair as synthetic data in the synthetic data store.”. The limitation “determine the corrected pair satisfies criterion of a synthetic data store” recites evaluating information according to a criterion, which is an evaluation or judgment that can practically be performed in the human mind.
The limitation “store the corrected pair as synthetic data in the synthetic data store” merely stores the result of the abstract evaluation. The claim does not recite an improvement to the operation or structure of the synthetic data store.
Accordingly, claims 2, 9 do not integrate the mental process into a practical application and does not add significantly more to the abstract idea. Therefore, claims are patent ineligible for the reasons discussed with respect to claims 1, 8.
Regarding claims 3, 10 further recite “utilize the generative AI model to generate a candidate natural language query based on the second query language query; and
“determine a level of similarity between the first natural language query and the candidate natural language query satisfies a consistency criterion.”
The limitation “utilize the generative AI model to generate a candidate natural language query based on the second query language query” is not characterized as a mental process merely because an AI model performs the operation.
However, the limitation “determine a level of similarity between the first natural language query and the candidate natural language query satisfies a consistency criterion”
recites comparing two items of linguistic information and determining whether the comparison satisfies a criterion. At the level of generality claimed, a person can compare the meaning of two natural-language statements and determine whether they are sufficiently similar or consistent. Thus, the limitation recites an evaluation or judgment that falls within the mental-process grouping.
Claims 3, 10 do not recite a particular similarity algorithm, computational technique, or improvement to the operation of the generative AI model. Accordingly, the additional limitations do not integrate the abstract idea into a practical application or provide significantly more.
Regarding claims 4, 11 further recite “generate a first prompt to cause the generative AI model to generate a candidate pair based on the indication and the first natural language query;” “determine the candidate pair is syntactically invalid;” “generate a second prompt to cause the generative AI model to generate the corrected pair based on the indication, the first natural language query, and the determination of the candidate pair being syntactically invalid; and “determine the corrected pair is syntactically valid.”
The limitations “determine the candidate pair is syntactically invalid” and “determine the corrected pair is syntactically valid” recite evaluation and judgment of information according to syntactic criteria. A person can review a query-language expression and determine whether its syntax complies with established syntax rules.
The limitations directed to “generate a first prompt” and “generate a second prompt” amount to preparing instructions based on the information and results of the preceding evaluation. Although the generative AI model itself is a computer implementation, claims 4, 11 do not recite a particular technical improvement to the AI model or its operation.
Accordingly, claims 4, 11 do not integrate the recited mental process into a practical application or add significantly more.
Regarding claims 5, 12 further recite “wherein the indication comprises a session identifier identifying a session in which the first natural language query was provided to the generative AI model, and the program code is further structured to cause the processor circuit to:” “identify the first query language query based at least on the session identifier; and “determine the first query language query corresponds to the first natural language query.”
The limitation “determine the first query language query corresponds to the first natural language query” recites evaluating the relationship between two items of information and determining that they correspond, which is an evaluation or judgment that can practically be performed mentally.
The limitation “identify the first query language query based at least on the session identifier” uses an identifier to identify associated information. The claim does not recite a particular improvement to session-management, indexing, retrieval, or computer-storage technology.
Therefore, claims 5, 12 do not integrate the abstract mental process into a practical application or provide significantly more.
Regarding claims 6, 13 further recites “determine a level of similarity between the corrected pair and a dataset pair stored in the synthetic data store fails to satisfy similarity criteria, the dataset pair comprising a synthetic natural language query and a synthetic query language query.” This limitation recites comparing the corrected pair with another pair, evaluating their level of similarity, and determining whether that level “fails to satisfy similarity criteria.” At the breadth of the claim, no particular similarity algorithm or mathematical technique is required. A person can compare two natural-language/query-language examples and judge whether the examples are sufficiently similar according to a given criterion. Accordingly, the limitation recites an evaluation or judgment within the mental-process grouping.
The claimed synthetic data store merely supplies information on which the abstract comparison is performed. Claims 6, 13 do not recite an improvement to the operation of the data store or to computer similarity-processing technology.
Accordingly, claims 6, 13 do not integrate the abstract idea into a practical application or provide significantly more.
Regarding claims 7, 14 further recite “determine the first natural language query is eligible to be converted based on at least one of a permission of a user account or an available table in a database.”. This limitation recites determining eligibility according to specified conditions, namely: “a permission of a user account” or “an available table in a database.” determining whether permission exists or whether required information is available and, based on that determination, deciding whether an operation is permitted recites evaluation and judgment that can practically be performed mentally.
The user account and database merely provide information used in making the determination. Claims 7, 14 do not recite an improvement to computer-security technology, database architecture, or database-access mechanisms.
Accordingly, claims 7, 14 do not integrate the mental process into a practical application or add significantly more.
Regarding claim 16 further recites “wherein the third query language query is a syntactically valid conversion of the second natural language query.”
Under Step 2A, Prong One, this limitation recites the mental process of converting linguistic information and evaluating whether the resulting query-language expression is syntactically valid.
A person familiar with the applicable query language can interpret the second natural language query, formulate the corresponding third query language query, and determine whether that query complies with applicable syntax rules.
Thus, the limitation recites mental activities involving conversion, evaluation, and judgment.
Under Step 2A, Prong Two, the additional elements inherited from claim 15, the processors, memories, language conversion engine, synthetic data generator, and generative AI model—merely implement the abstract operations and do not recite a particular technological improvement.
Under Step 2B, these additional elements do not provide significantly more than the judicial exception.
Therefore, claim 16 is patent ineligible for the reasons stated with respect to claim 15.
Regarding claim 17 further recites “a second processor circuit; and” “a second memory device comprising second programming instructions structured to cause the second processor circuit to:” “receive, from the language conversion engine, the indication of negative feedback,” “generate the second pair based on the indication of negative feedback, and” “update the synthetic data with the second pair, resulting in the updated synthetic data.”
Under Step 2A, Prong One, the limitations “receive, from the language conversion engine, the indication of negative feedback,” “generate the second pair based on the indication of negative feedback,” and “update the synthetic data with the second pair” recite receiving evaluative information, correcting or generating associated information based on that evaluation, and recording the corrected information.
A person can receive negative feedback regarding a prior natural-language/query-language association, generate a corrected pair based on the feedback, and update the existing information with the corrected pair. These activities involve evaluation, judgment, correction, and organization of information and therefore recite a mental process.
The following are additional elements:
“a second processor circuit” and “a second memory device comprising second programming instructions.”
Under Step 2A, Prong Two, the second processor circuit and second memory device merely automate the recited mental operations and do not provide a particular improvement to computer technology.
Under Step 2B, the additional computer components do not add significantly more than the abstract idea.
Therefore, claim 17 is patent ineligible.
Regarding claim 18 further recites, “generate a candidate natural language query based on the third query language query, and” “determine a level of similarity between the second natural language query and the candidate natural language query satisfies a consistency criterion; and
“update, based on the determination of the level of similarity between the second natural language query and the candidate natural language query satisfying the consistency criterion, the synthetic data with the second pair.”
Under Step 2A, Prong One, “generate a candidate natural language query based on the third query language query” recites converting query-language information into natural-language information. A person familiar with the query language can read a query and express its meaning in natural language.
The limitation “determine a level of similarity between the second natural language query and the candidate natural language query satisfies a consistency criterion” recites comparing two items of linguistic information and judging whether they are sufficiently similar or consistent.
The limitation “update ... the synthetic data with the second pair” recites recording information based on the result of that evaluation.
Accordingly, these limitations recite mental processes involving conversion, comparison, evaluation, judgment, and organization of information.
Under Step 2A, Prong Two, the processors, memories, synthetic data generator, and generative AI model inherited from the preceding claims are additional elements that merely implement these abstract operations and do not recite a particular technological improvement.
Under Step 2B, the additional elements do not provide significantly more than the judicial exception.
Therefore, claim 18 is patent ineligible.
Regarding claim 19 further recites “generate a second prompt to cause the generative AI model to generate a candidate pair based on the indication of negative feedback and the second natural language query;” “determine the candidate pair is syntactically invalid;” “generate a third prompt to cause the generative AI model to generate the corrected pair based on the indication, the second natural language query, and the determination of the candidate pair being syntactically invalid; and” “determine the corrected pair is syntactically valid.”
Under Step 2A, Prong One, these limitations recite formulating instructions based on information, evaluating a candidate pair according to syntax rules, correcting the pair based on that evaluation, and evaluating whether the corrected pair satisfies the applicable syntax rules.
A person can formulate instructions based on negative feedback and a natural-language query, review a candidate query to determine that it is syntactically invalid, formulate corrective instructions, and determine that a corrected query is syntactically valid.
Thus, these limitations recite mental processes involving instruction formulation, evaluation, judgment, correction, and validation.
The “generative AI model” is treated as an additional element used to implement these mental operations.
Under Step 2A, Prong Two, the generative AI model and inherited computer components merely perform the recited abstract information-processing operations and do not recite a particular technological improvement.
Under Step 2B, the additional elements do not amount to significantly more than the judicial exception.
Therefore, claim 19 is patent ineligible.
Regarding claim 20 further recites “provide, to the synthetic data generator, a session identifier identifying a session in which the first natural language query was provided to the generative AI model, causing the synthetic data generator to identify the first query language query based at least on the session identifier and to determine the first query language query corresponds to the first natural language query.”
Under Step 2A, Prong One, this limitation recites identifying information using an associated identifier and determining whether two items of information correspond.
A person can use a session identifier or label to locate information associated with a particular session and can evaluate whether a particular query-language query corresponds to a particular natural-language query.
Thus, the limitation recites mental processes involving identification, association, evaluation, and judgment.
The session identifier, synthetic data generator, generative AI model, processors, and memories are additional elements used to implement the abstract process.
Under Step 2A, Prong Two, those elements do not recite a particular improvement to session management, indexing, retrieval, processor operation, memory operation, or AI architecture; rather, they are used as tools to implement the abstract identification and correspondence determination.
Under Step 2B, the additional elements do not provide significantly more than the judicial exception.
Therefore, claim 20 is patent ineligible.
Double Patenting
The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969).
A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on nonstatutory double patenting provided the reference application or patent either is shown to be commonly owned with the examined application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. See MPEP § 717.02 for applications subject to examination under the first inventor to file provisions of the AIA as explained in MPEP § 2159. See MPEP § 2146 et seq. for applications not subject to examination under the first inventor to file provisions of the AIA . A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b).
The filing of a terminal disclaimer by itself is not a complete reply to a nonstatutory double patenting (NSDP) rejection. A complete reply requires that the terminal disclaimer be accompanied by a reply requesting reconsideration of the prior Office action. Even where the NSDP rejection is provisional the reply must be complete. See MPEP § 804, subsection I.B.1. For a reply to a non-final Office action, see 37 CFR 1.111(a). For a reply to final Office action, see 37 CFR 1.113(c). A request for reconsideration while not provided for in 37 CFR 1.113(c) may be filed after final for consideration. See MPEP § 706.07(e) and 714.13.
The USPTO Internet website contains terminal disclaimer forms which may be used. Please visit www.uspto.gov/patent/patents-forms. The actual filing date of the application in which the form is filed determines what form (e.g., PTO/SB/25, PTO/SB/26, PTO/AIA /25, or PTO/AIA /26) should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto- processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to www.uspto.gov/patents/apply/applying-online/eterminal-disclaimer.
Claims 1, 4, 8, 11, 15-17, 19 are rejected on the ground of nonstatutory double patenting as being unpatentable over claims 1, 3, 5, 12-13, 15 of U.S. Patent No. 12,430,328 B1. Although the claims at issue are not identical, they are not patentably distinct from each other because they are cover the same recitations, limitations that are combined or worded slightly differently from the original patented claims but in essence they convey the same subject matter. The instant application is a broader scope of recitation of the claims of the already patented application.
Present Application 19/294444
US PAT. 12430328 B1
1. A synthetic generation system, comprising: a processor circuit; and a memory device that stores program code structured to cause the processor circuit to: receive an indication of negative feedback for a first query language query generated by a generative artificial intelligence (AI) model based on a first natural language query, generate a corrected pair based on the indication and the first natural language query, the corrected pair comprising the first natural language query and a second query language query, the second query language query a syntactically valid conversion of the first natural language query, and cause a language conversion engine to utilize the corrected pair to convert a second natural language query to a third query language query.
4. The synthetic generation system of claim 1, wherein to generate the corrected pair, the program code is further structured to cause the processor circuit to: generate a first prompt to cause the generative Al model to generate a candidate pair based on the indication and the first natural language query; determine the candidate pair is syntactically invalid; generate a second prompt to cause the generative Al model to generate the corrected pair based on the indication, the first natural language query, and the determination of the candidate pair being syntactically invalid; and determine the corrected pair is syntactically valid.
8. A computer implemented method, comprising: receiving an indication of negative feedback for a first query language query generated by a generative artificial intelligence (AI) model based on a first natural language query; generating a corrected pair based on the indication and the first natural language query, the corrected pair comprising the first natural language query and a second query language query, the second query language query a syntactically valid conversion of the first natural language query; and causing a language conversion engine to utilize the corrected pair to convert a second natural language query to a third query language query.
11. The computer implemented method of claim 8, wherein said generating the corrected pair comprises: generating a first prompt to cause the generative Al model to generate a candidate pair based on the indication and the first natural language query; determining the candidate pair is syntactically invalid; generating a second prompt to cause the generative Al model to generate the corrected pair based on the indication, the first natural language query, and said determining the candidate pair is syntactically invalid; and determining the corrected pair is syntactically valid.
15. A system, comprising: a language conversion engine comprising a first processor and a first memory device, the first memory device comprising first programming instructions structured to cause the first processor to: receive a request to convert a first natural language query to a first query language query, generate a first prompt based on the first natural language query and synthetic data, the synthetic data comprising a first pair comprising a second natural language query and a second query language query, utilize a generative artificial intelligence (Al) model to convert the first natural language query to the first query language query based at least on the first prompt, receive an indication of negative feedback for the first query language query, cause a synthetic data generator to update the synthetic data based at least on the indication of negative feedback, resulting in updated synthetic data comprising a second pair comprising the second natural language query and a third query language query, and utilize the updated synthetic data to generate a second prompt based on a third natural language query, the second prompt structured to cause the generative Al model to convert the third natural language query to a fourth natural language query based at least on the updated synthetic data.
16. The system of claim 15, wherein the third query language query is a syntactically valid conversion of the second natural language query.
17. The system of claim 15, further comprising the synthetic data generator, the synthetic data generator comprising: a second processor circuit; and a second memory device comprising second programming instructions structured to cause the second processor circuit to: receive, from the language conversion engine, the indication of negative feedback, generate the second pair based on the indication of negative feedback, and update the synthetic data with the second pair, resulting in the updated synthetic data.
19. The system of claim 17, wherein to generate the second pair, the second programming instructions are further structured to cause the processor circuit to: generate a second prompt to cause the generative Al model to generate a candidate pair based on the indication of negative feedback and the second natural language query; determine the candidate pair is syntactically invalid; generate a third prompt to cause the generative Al model to generate the corrected pair based on the indication, the second natural language query, and the determination of the candidate pair being syntactically invalid; and determine the corrected pair is syntactically valid.
1. A system for generating synthetic data for use in performance benchmarking, comprising: a processor circuit; and a memory device that stores program code to be executed by the processor circuit, the program code comprising: a synthetic data generator that: obtains a dataset pair comprising a first natural language query and a first query language query; utilizes the dataset pair and first predicted catalog information to generate a first prompt to cause a large language model (LLM) to generate a variation of the dataset pair; responsive to providing the first prompt to the LLM, receives a first augmented pair comprising a first augmented natural language query and a first augmented query language query, the first augmented natural language query a variation of the first natural language query and the first augmented query language query a variation of the first query language query; generates synthetic data comprising the first augmented pair; and causes a language conversion engine to utilize the synthetic data to: generate a second prompt based on a second natural language query and the synthetic data, and provide the second prompt to the LLM to cause the LLM to convert the second natural language query to a second query language query.
3. The system of claim 1, wherein the second prompt is generated based on the first augmented pair and the program code further comprises a pair correction component that: receives an indication of negative feedback for the second query language query; generates a corrected pair based on the indication and the second natural language query, the corrected pair comprising the second natural language query and a corrected query language query, the corrected query language query a syntactically valid conversion of the second natural language query; and updates the synthetic data to comprise the corrected pair.
5. The system of claim 4, wherein to prompt the LLM to generate the corrected pair, the pair correction component: generates a second prompt to cause the LLM to generate a candidate pair based on the indication and the second natural language query; determines the candidate pair is syntactically invalid; generates a second prompt to cause the LLM to generate the corrected pair based on the indication, the second natural language query, and the determination that the candidate pair is syntactically invalid; and determines the corrected pair is syntactically valid.
12. A method for generating synthetic data for use in performance benchmarking, comprising: obtaining a dataset pair comprising a first natural language query and a first query language query; utilizing the dataset pair and first predicted catalog information to generate a first prompt to cause a generative artificial intelligence (AI) model to generate a variation of the dataset pair; responsive to providing the first prompt to the generative AI model, receiving a first augmented pair comprising a first augmented natural language query and a first augmented query language query, the first augmented natural language query a variation of the first natural language query and the first augmented query language query a variation of the first query language query; generating synthetic data comprising the first augmented pair; and causing a language conversion engine to utilize the synthetic data to: generate a second prompt based on a second natural language query and the synthetic data, and provide the second prompt to the generative AI model to cause the generative AI model to convert the second natural language query to a second query language query.
13. The method of claim 12, wherein the second prompt is generated based on the first augmented pair and the method further comprises: receiving an indication of negative feedback for the second query language query; generating a corrected pair based on the indication and the second natural language query, the corrected pair comprising the second natural language query and a corrected query language query, the corrected query language query a syntactically valid conversion of the second natural language query; and updating the synthetic data to comprise the corrected pair.
15. The method of claim 14, wherein to prompt the LLM to generate the corrected pair, the pair correction component: generates a second prompt to cause the LLM to generate a candidate pair based on the indication and the second natural language query; determines the candidate pair is syntactically invalid; generates a second prompt to cause the LLM to generate the corrected pair based on the indication, the second natural language query, and the determination that the candidate pair is syntactically invalid; and determines the corrected pair is syntactically valid.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1 and 8 are rejected under 35 U.S.C. 103 as being unpatentable over Dharnidharka et al (US 2025/0021767 A1) in view of Truong et al (US 2024/0330279 A1, which claims priority to provisional application no 63/493,693 filed on 03/31/2023).
As per claim 1, Dharnidharka teaches a synthetic generation system, comprising: a processor circuit; and a memory device that stores program code structured to cause the processor circuit to (see Fig. 1A, teaches a data processing environment including model training system 170 and computing resources 160 for generating training data from natural-language descriptions, LLM-generated search-query translations, and user feedback.):
receive an indication of negative feedback for a first query language query generated by a generative artificial intelligence (AI) model based on a first natural language query (see Figs. 2–4 and corresponding description, teaches receiving a natural-language description of a search query, transmitting the natural-language description to LLMs, receiving generative-AI translations into executable search queries, and receiving user feedback indicating whether each translation is correct, incorrect, or partially incorrect. ’767 identifies examples of the models as generative pretrained transformers including GPT-2, GPT-3, GPT-4 and Codex. Thus, an “incorrect” or “partially incorrect” indication corresponds to negative feedback for the query generated by the generative AI model from the natural-language description), generate a corrected pair based on the indication and the first natural language query, the corrected pair comprising the first natural language query and a second query language query, the second query language query a syntactically valid conversion of the first natural language query (see Figs. 1B, 3B, and 4; claims 1, 4, and 6. Fig. 3B teaches that search query 312 is a translation of user input 304 into SPL generated by a first LLM and provides input box 314 for receiving user feedback indicating whether search query 312 is correct, partially correct, or incorrect. Fig. 4 teaches receiving a natural-language description of a search query in text box 402, displaying LLM-generated translations in result display boxes 408a-408d, receiving feedback concerning correctness through UI elements 410a-410d, and receiving in text box 412 an “expected response,” i.e., a translation of the natural-language description provided in text box 402. Fig. 1B teaches that model training system 170 generates training data from the user input, translations provided by the LLMs, and user feedback; the training data may be stored as a table in which each row corresponds to a natural-language description and includes the natural-language description, LLM responses, and user feedback concerning syntactic and/or semantic correctness. Claim 1 expressly teaches obtaining a natural-language description of a search query, requesting from the LLMs a “syntactically correct version of the search query” corresponding to the natural-language description, and receiving user feedback indicating whether the LLM results are syntactically correct. Claim 4 further teaches receiving additional text-based user input corresponding to the “syntactically correct version of the search query,” and claim 6 identifies that syntactically correct version as a pipelined search-query statement. Thus, the natural-language description corresponds to the first natural language query, and the syntactically correct/expected translation corresponds to the second query language query that is a syntactically valid conversion of the first natural language query),
Dharnidharka does not explicitly teach cause a language conversion engine to utilize the corrected pair to convert a second natural language query to a third query language query.
However, Truong teaches cause a language conversion engine to utilize the corrected pair to convert a second natural language query to a third query language query (see claim 5, teaches displaying a generated query to a user, receiving from the user “a corrected query that corrects the query,” and storing the user input as a new predefined input and the corrected query as a new predefined query. See claim 1, teaches subsequently receiving a user input, selecting a predefined input based on similarity to the user input, and prompting a trained machine-learning model to generate a query based on the user input and a predefined query associated with the selected predefined input. See claim 7, teaches that the prompt includes the user input and the predefined query as an example; claim 10 identifies the trained machine-learning model as an LLM. Thus, the stored user input/corrected query corresponds to the corrected pair, the subsequently received user input corresponds to the second natural language query, and the query generated by the LLM corresponds to the third query language query).
Thus, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to apply the teachings of Truong with the teachings of Dharnidharka in order to generate more accurate database queries that capture user intent by learning from example queries associated with example requests that are most similar to subsequent user requests (Truong).
Regarding claim 8, claim 8 is rejected for substantially the same reason as claim 1 above.
Claims 2, 9 are rejected under 35 U.S.C. 103 as being unpatentable over Dharnidharka et al (US 2025/0021767 A1) in view of Truong et al (US 2024/0330279 A1, which claims priority to provisional application no 63/493,693 filed on 03/31/2023) further in view of Misra et al (WO 2023/107744 A2).
As per claim 2, Dharnidharka and Truong do not explicitly teach wherein to generate the corrected pair, the program code is further structured to cause the processor circuit to: determine the corrected pair satisfies criterion of a synthetic data store; and store the corrected pair as synthetic data in the synthetic data store.
However, Misra teaches wherein to generate the corrected pair, the program code is further structured to cause the processor circuit to: determine the corrected pair satisfies criterion of a synthetic data store; and store the corrected pair as synthetic data in the synthetic data store (see Fig. 7, steps 710–714 and claim 1, teaches receiving from a machine-learning instance a generated natural-language query corresponding to a structured query, validating the generated natural-language query, and only after validation adding the structured query and validated natural-language query “as a training pair to the training set.” Misra further teaches that validation may include determining that the generated natural-language query is a valid representation of its corresponding structured query, or repairing the generated natural-language query so that it better corresponds to the structured query. Thus, validation is the criterion for admission into the training set, and the validated/repaired natural-language query and corresponding structured query are stored as a training pair in the training set).
Thus, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to apply the teachings of Misra with the teachings of Dharnidharka and Truong in order to ensure that generated training samples are validated before being added as training pairs to the training set (Misra).
Regarding claim 9, claim 9 is rejected for substantially the same reason as claim 2 above.
Claims 3, 10 are rejected under 35 U.S.C. 103 as being unpatentable over Dharnidharka et al (US 2025/0021767 A1) in view of Truong et al (US 2024/0330279 A1, which claims priority to provisional application no 63/493,693 filed on 03/31/2023) in view of Misra et al (WO 2023/107744 A2) further in view of Sheinin et al (US 2020/0257679 A1).
As per claim 3, wherein to determine the corrected pair satisfies criterion of the synthetic data store, the program code is further structured to cause the processor circuit to: utilize the generative Al model to generate a candidate natural language query based on the second query language query (see Misra, Fig. 7, steps 708–710 and claims 1–2, teaches training a machine-learning instance using structured-query/natural-language-query training pairs, providing a structured query to the trained machine-learning instance, and receiving from the machine-learning instance a natural-language query corresponding to that structured query. Claim 2 expressly states that the machine-learning instance is GPT-3. Thus, the structured query corresponds to the second query language query and the GPT-3-generated natural-language query corresponds to the candidate natural language query); but does not explicitly teach determine a level of similarity between the first natural language query and the candidate natural language query satisfies a consistency criterion. However, Sheinin teaches determine a level of similarity between the first natural language query and the candidate natural language query satisfies a consistency criterion (see Fig. 2 and claim 13, teaches receiving an input question in natural language, generating natural-language utterances corresponding to possible SQL queries, and executing a paraphrase model that “measures a similarity” between the utterances generated for a SQL query and utterances generated from the original input question. Sheinin then determines which possible SQL query has the highest similarity. Thus, the original input question corresponds to the first natural language query, the natural-language utterance corresponding to the SQL query corresponds to the candidate natural language query, and the measured/ranked similarity corresponds to the consistency criterion).
Thus, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to apply the teachings of Sheinin with the teachings of Dharnidharka, Truong, and Misra in order to determine whether the natural-language meaning represented by a generated database query is consistent with the original natural-language question (Sheinin).
Regarding claim 10, claim 10 is rejected for substantially the same reason as claim 3 above.
Claims 4, 11 are rejected under 35 U.S.C. 103 as being unpatentable over Dharnidharka et al (US 2025/0021767 A1) in view of Truong et al (US 2024/0330279 A1, which claims priority to provisional application no 63/493,693 filed on 03/31/2023) further in view of Liu et al (US 2025/0005018 A1).
As per claim 4, Truong teaches wherein to generate the corrected pair, the program code is further structured to cause the processor circuit to: generate a first prompt to cause the generative Al model to generate a candidate pair based on the indication and the first natural language query (see, Fig. 4, step 416 and paragraphs [0045]-[0047], teaches that after a generated query produces an error, request processing module 202 generates a prompt that includes the user request, the previously generated query, and the error message and prompts language model 204 to generate a new query correcting the error. The user request corresponds to the first natural language query, the error corresponds to the indication, and the user request together with the newly generated query corresponds to the candidate input/query pair); but it does not explicitly teach determine the candidate pair is syntactically invalid; generate a second prompt to cause the generative Al model to generate the corrected pair based on the indication, the first natural language query, and the determination of the candidate pair being syntactically invalid; and determine the corrected pair is syntactically valid.
However, Liu teaches determine the candidate pair is syntactically invalid; generate a second prompt to cause the generative Al model to generate the corrected pair based on the indication, the first natural language query, and the determination of the candidate pair being syntactically invalid; and determine the corrected pair is syntactically valid (see claim 8 and corresponding description, teaches that in response to determining that a syntax error exists, an error log is input into the LLM and syntax correction of the initial SQL statement is performed based on the error log. The error log may include the query request corresponding to the SQL statement, the error reason, and the error type of the syntax error. Thus, the corrective input to the LLM is based on the natural-language query request and the determination that the SQL statement contains a syntax error and performing syntax correction based on the error log “until a final SQL statement is obtained,” and expressly states that the LLM performs syntax correction and adjustment “until a correct SQL statement is finally obtained”).
Thus, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to apply the teachings of Liu with the teachings of Dharnidharka and Truong in order to avoid syntax errors in an LLM-generated SQL statement and ensure the accuracy of the SQL statement before execution (Liu).
Regarding claim 11, claim 11 is rejected for substantially the same reason as claim 4 above.
Claims 5, 7, 12, 14 are rejected under 35 U.S.C. 103 as being unpatentable over Dharnidharka et al (US 2025/0021767 A1) in view of Truong et al (US 2024/0330279 A1, which claims priority to provisional application no 63/493,693 filed on 03/31/2023) further in view of Jain et al (US 12,430,333 B2).
As per claim 5, Dharnidharka and Truong do not explicitly teach wherein the indication comprises a session identifier identifying a session in which the first natural language query was provided to the generative Al model, and the program code is further structured to cause the processor circuit to: identify the first query language query based at least on the session identifier; and determine the first query language query corresponds to the first natural language query.
However, Jain teaches wherein the indication comprises a session identifier identifying a session in which the first natural language query was provided to the generative Al model (see col. 15, lines 17-31, teaches initiating a session with an LLM, receiving a query having a natural-language portion, prompting the LLM “in the session” with the processed natural-language portion, and receiving native database query content from the LLM; see also col. 17, lines. 48-58, teaches “a user-specific inner session identifier” and retrieving context for a user using the user-specific inner session identifier before processing a natural-language query for that user), and the program code is further structured to cause the processor circuit to: identify the first query language query based at least on the session identifier (see col. 17, lines 48-58, teaches retrieving user context based on a user-specific inner session identifier; see also col. 19, lines 56-67 through col. 20, lines 1-8, teaches treating a database session as an LLM conversation and maintaining a history of past prompts and responses in the current database session); and determine the first query language query corresponds to the first natural language query (see col. 15, lines 17-31, teaches receiving a query having a natural-language portion, prompting the LLM in the session with that processed natural-language portion, and receiving native database query content from the LLM).
Thus, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to apply the teachings of Jain with the teachings of the Dharnidharka and Truong in order to prevent generation or execution of queries against unauthorized or unavailable database structures by limiting natural-language query processing according to user privileges and database structures available to the user (Jain).
As per claim 7, wherein the program code is further structured to cause the processor circuit to: determine the first natural language query is eligible to be converted based on at least one of a permission of a user account or an available table in a database (see col. 14, lines 35-49, teaches retrieving available schema information for a user based on one or more roles, profiles, and/or privileges, determining which database objects are available to the user for query execution based on the user's roles/profiles/privileges, and adding the available tables to the schema information used for prompting the LLM and col. 18, lines 15-35, teaches storing security configuration information for a session that determines “whether the session is allowed to execute natural language queries,” and teaches that privileges may be managed on a user-by-user or session-by-session basis and permissions may be checked for natural-language query processing, Jain).
Regarding claims 12, 14, claims 12, 14 are rejected for substantially the same reason as claims 5, 7 above.
Claims 6, 13 are rejected under 35 U.S.C. 103 as being unpatentable over Dharnidharka et al (US 2025/0021767 A1) in view of Truong et al (US 2024/0330279 A1, which claims priority to provisional application no 63/493,693 filed on 03/31/2023) further in view of BHAISAHEB et al (US 2024/0119046 A1).
As per claim 6, Dharnidharka and Truong do not explicitly teach wherein the program code is further structured to cause the processor circuit to: determine a level of similarity between the corrected pair and a dataset pair stored in the synthetic data store fails to satisfy similarity criteria, the dataset pair comprising a synthetic natural language query and a synthetic query language query.
However, BHAISAHEB teaches wherein the program code is further structured to cause the processor circuit to: determine a level of similarity between the corrected pair and a dataset pair stored in the synthetic data store fails to satisfy similarity criteria, the dataset pair comprising a synthetic natural language query and a synthetic query language query (see paragraphs [0017], [0032], [0047]-[0054], Fig. 4C, Algorithm 1, and claims 3 and 6, teaches maintaining training data comprising Natural Language (NL)-SQL query pairs; synthesizing SQL programs and feeding the synthesized SQL programs to a backward model to generate corresponding synthetic Natural Language queries; thereby forming synthetic NL-query/SQL-program pairs; extracting a sentence-BERT representation for a generated synthetic NL query and comparing the representation with representations of NL queries in the existing dataset using cosine similarity; determining a maximum similarity score; and filtering the generated synthetic NL query when the maximum similarity score is below a predefined similarity threshold, while maintaining the corresponding SQL program as the associated query-language component of the pair).
Thus, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to apply the similarity-based filtering teachings of BHAISAHEB with the teachings of Dharnidharka and Truong in order to compare newly generated NL/query-language data with NL/query-language data already represented in a training dataset and reject or filter generated data that fails a predefined similarity criterion, thereby controlling the composition of synthetic training data and preventing unsuitable synthetic examples from being added to the dataset (BHAISAHEB).
Regarding claim 13, claim 13 is rejected for substantially the same reason as claim 6 above.
Claims 15-17 are rejected under 35 U.S.C. 103 as being unpatentable over Truong et al (US 2024/0330279 A1, which claims priority to provisional application no 63/493,693 filed on 03/31/2023) in view of Dharnidharka et al (US 2025/0021767 A1) in view of Relan et al (US 2022/0374420 A1) further in view of Bista et al (US 2025/0094455 A1).
As per claim 15, Truong teaches a system, comprising: a language conversion engine comprising a first processor and a first memory device, the first memory device comprising first programming instructions structured to cause the first processor to (see claim 20, teaches a system comprising one or more memories storing instructions and one or more processors coupled to the memories for performing the disclosed natural-language-to-query generation operations): receive a request to convert a first natural language query to a first query language query (see claim 1 and corresponding description, teaches receiving a user input and prompting a trained machine-learning model to generate a database query for the user input; Truong expressly identifies query generation as translating natural-language user intent into a database query), generate a first prompt based on the first natural language query and synthetic data, the synthetic data comprising a first pair comprising a second natural language query and a second query language query (see claims 1 and 7, teaches selecting a predefined/example input based on similarity to the current user input and generating a prompt that includes the current user input and a predefined query associated with the selected predefined input as an example. The stored predefined input and corresponding predefined query constitute an NL-query/QL-query example pair), utilize a generative artificial intelligence (Al) model to convert the first natural language query to the first query language query based at least on the first prompt (see claims 1, 7, and 10, teaches inputting the prompt containing the user input and example query into a trained machine-learning model, wherein the trained model comprises an LLM, to generate the query corresponding to the user input; ’ Truong identifies GPT/Generative Pre-trained Transformer as an exemplary language model),
and utilize the updated synthetic data to generate a second prompt based on a third natural language query (see claims 1, 5, and 7, teaches receiving a user input, selecting a predefined input based on similarity to the user input, and generating a prompt that includes the user input and a predefined query associated with the selected predefined input as an example; claim 5 further teaches receiving a corrected query and storing the user input as a new predefined input and the corrected query as a new predefined query. Thus, the stored corrected input/query pair constitutes updated example data that is subsequently available for selection and inclusion in a prompt generated for a later user input),
Truong does not explicitly teach receive an indication of negative feedback for the first query language query,
However, Dharnidharka teaches receive an indication of negative feedback for the first query language query (see claim 1, teaches receiving user feedback indicating whether LLM-generated search-query results are syntactically correct, and claim 4 teaches additional user input corresponding to a syntactically correct version of the query),
Thus, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to apply the teachings of Dharnidharka with the teachings of Truong in order to receive negative feedback concerning an LLM-generated query and use a corresponding syntactically correct query as corrective information, thereby enabling the system to improve subsequent natural-language-to-query-language conversions and reduce the generation of incorrect or syntactically invalid queries (Dharnidharka).
Truong and Dharnidharka do not explicitly teach cause a synthetic data generator to update the synthetic data based at least on the indication of negative feedback, resulting in updated synthetic data comprising a second pair comprising the second natural language query and a third query language query,
However, Relan teaches cause a synthetic data generator to update the synthetic data based at least on the indication of negative feedback, resulting in updated synthetic data comprising a second pair comprising the second natural language query and a third query language query (see paragraphs [0018], [0040]-[0046], teaches training data comprising natural-language-question/database-query tuples; generating a recommended query for a natural-language question; receiving analyst review that may correct or modify the recommended query; and, based on the analyst’s acceptance, improvement, or correction, adding the question and corresponding accepted, improved, or corrected query to the training dataset as new training data for subsequent model training),
Thus, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to apply the feedback-based training teachings of Relan with the teachings of Truong and Dharnidharka in order to update the training/synthetic data with a corrected natural-language-question/database-query pair after receiving negative feedback, thereby allowing the model to learn from the corrected query and improve the accuracy of query generation when subsequently presented with the same or a similar natural-language question (Relan).
Truong, Dharnidharka and Relan do not explicitly teach the second prompt structured to cause the generative Al model to convert the third natural language query to a fourth natural language query based at least on the updated synthetic data.
However, Bista teaches the second prompt structured to cause the generative Al model to convert the third natural language query to a fourth natural language query based at least on the updated synthetic data (see, paragraph [0065], teaches that CQR model 214C generates a modified query enriched with contextual information from previous input in a dialog session, that CQR model 214C may be an LLM, including a specialized Mistral-7B-v0.1 LLM, and that based on an original or rewritten query the execution engine generates a prompt that can include dialog history and information from a context and memory store and paragraphs [0078] and [0082]–[0084], teaches that the CQR model is applied to rewrite a query based on conversational history; a user provides an nth natural-language query qn, the query qn is provided to CQR model 408, CQR model 408 rewrites qn using conversation history 406 to produce rewritten natural-language query qn∗, and rewritten query qn∗ may then be provided as part of a prompt to a backend language model. Thus, qn corresponds to the claimed third natural language query and rewritten qn∗ corresponds to the claimed fourth natural language query).
Thus, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to apply the teachings of Bista with the teachings of Truong, Dharnidharka and Relan in order to use contextual information from prior interactions when processing a subsequent natural-language query, cause an LLM-based contextual query rewriting model to convert the subsequent natural-language query into a modified natural-language query, and use the modified query for further language-model processing, thereby resolving ambiguity and omitted context and improving the accuracy of subsequent query processing (Bista).
As per claim 16, wherein the third query language query is a syntactically valid conversion of the second natural language query (see claims 1, 4, and 6, teaches receiving a natural-language description of a search query, obtaining or receiving a “syntactically correct version of the search query” corresponding to the natural-language description, and, in claim 6, specifies that the syntactically correct version is a query-language statement, Dharnidharka).
As per claim 17 , Relan teaches further comprising the synthetic data generator, the synthetic data generator comprising: a second processor circuit; and a second memory device comprising second programming instructions structured to cause the second processor circuit to: receive, from the language conversion engine, the indication of negative feedback, generate the second pair based on the indication of negative feedback, and update the synthetic data with the second pair, resulting in the updated synthetic data (see [0018], [0041], and [0044], teaches training data comprising natural-language-phrase/database-query tuples and, following analyst review of a generated query, accepting, improving, correcting, or modifying the query such that the corrected natural-language-question/query information constitutes new training data and see [0045]-[0046], teaches adding the natural-language question and corresponding accepted, improved, or corrected database query to the training dataset, thereby generating new training data from the corrected NL2Query output for subsequent model training).
Claim 18 is rejected under 35 U.S.C. 103 as being unpatentable over Truong et al (US 2024/0330279 A1, which claims priority to provisional application no 63/493,693 filed on 03/31/2023) in view of Dharnidharka et al (US 2025/0021767 A1) in view of Relan et al (US 2022/0374420 A1) in view of Bista et al (US 2025/0094455 A1) further in view of BHAISAHEB et al (US 2024/0119046 A1).
As per claim 18, Truong, Dharnidharka, Relan and Bista do not explicitly teach wherein to generate the second pair, the second programming instructions are further structured to cause the second processor circuit to: generate a candidate natural language query based on the third query language query, and determine a level of similarity between the second natural language query and the candidate natural language query satisfies a consistency criterion; and to update the synthetic data with the second pair, the second programming instructions are further structured to cause the second processor circuit to: update, based on the determination of the level of similarity between the second natural language query and the candidate natural language query satisfying the consistency criterion, the synthetic data with the second pair.
However, BHAISAHEB teaches generate a candidate natural language query based on the third query language query, and determine a level of similarity between the second natural language query and the candidate natural language query satisfies a consistency criterion; and to update the synthetic data with the second pair, the second programming instructions are further structured to cause the second processor circuit to: update, based on the determination of the level of similarity between the second natural language query and the candidate natural language query satisfying the consistency criterion, the synthetic data with the second pair (see paragraphs [0017], [0032], [0047]-[0054], Fig. 4C, Algorithm 1, and claims 3 and 6, teaches maintaining training data comprising Natural Language (NL)-SQL query pairs; synthesizing SQL programs and feeding the synthesized SQL programs to a backward model to generate corresponding synthetic Natural Language queries; thereby forming synthetic NL-query/SQL-program pairs; extracting a sentence-BERT representation for a generated synthetic NL query and comparing the representation with representations of NL queries in the existing dataset using cosine similarity; determining a maximum similarity score; and filtering the generated synthetic NL query when the maximum similarity score is below a predefined similarity threshold, while maintaining the corresponding SQL program as the associated query-language component of the pair).
Thus, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to apply the similarity-based filtering teachings of BHAISAHEB with the teachings of Truong, Dharnidharka, Relan and Bista in order to compare newly generated NL/query-language data with NL/query-language data already represented in a training dataset and reject or filter generated data that fails a predefined similarity criterion, thereby controlling the composition of synthetic training data and preventing unsuitable synthetic examples from being added to the dataset (BHAISAHEB).
Claim 19 is rejected under 35 U.S.C. 103 as being unpatentable over Truong et al (US 2024/0330279 A1, which claims priority to provisional application no 63/493,693 filed on 03/31/2023) in view of Dharnidharka et al (US 2025/0021767 A1) in view of Relan et al (US 2022/0374420 A1) in view of Bista et al (US 2025/0094455 A1) further in view of Liu et al (US 2025/0005018 A1).
As per claim 19, Truong teaches wherein to generate the second pair, the second programming instructions are further structured to cause the processor circuit to: generate a second prompt to cause the generative Al model to generate a candidate pair based on the indication of negative feedback and the second natural language query (see claims 8, 17, and 18, teaches executing a generated query and, responsive to an error, prompting a trained machine-learning model to generate a corrected query based on the user input, the generated query, and the error; if another error occurs, the model is again prompted to generate another corrected query);
Dharnidharka, Truong, Relan and Bista do not explicitly teach determine the candidate pair is syntactically invalid; generate a third prompt to cause the generative Al model to generate the corrected pair based on the indication, the second natural language query, and the determination of the candidate pair being syntactically invalid; and determine the corrected pair is syntactically valid.
However, Liu teaches determine the candidate pair is syntactically invalid; generate a third prompt to cause the generative Al model to generate the corrected pair based on the indication, the second natural language query, and the determination of the candidate pair being syntactically invalid; and determine the corrected pair is syntactically valid (see claim 8 and corresponding description, teaches that in response to determining that a syntax error exists, an error log is input into the LLM and syntax correction of the initial SQL statement is performed based on the error log. The error log may include the query request corresponding to the SQL statement, the error reason, and the error type of the syntax error. Thus, the corrective input to the LLM is based on the natural-language query request and the determination that the SQL statement contains a syntax error and performing syntax correction based on the error log “until a final SQL statement is obtained,” and expressly states that the LLM performs syntax correction and adjustment “until a correct SQL statement is finally obtained”).
Thus, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to apply the teachings of Liu with the teachings of Dharnidharka, Truong, Relan and Bista in order to avoid syntax errors in an LLM-generated SQL statement and ensure the accuracy of the SQL statement before execution (Liu).
Claim 20 is rejected under 35 U.S.C. 103 as being unpatentable over Truong et al (US 2024/0330279 A1, which claims priority to provisional application no 63/493,693 filed on 03/31/2023) in view of Dharnidharka et al (US 2025/0021767 A1) in view of Relan et al (US 2022/0374420 A1) in view of Bista et al (US 2025/0094455 A1) further in view of Jain et al (US 12,430,333 B2).
As per claim 20, Dharnidharka, Truong, Relan and Bista do not explicitly teach wherein to cause the synthetic data generator to update the synthetic data, the program code is further structured to cause the processor circuit to: provide, to the synthetic data generator, a session identifier identifying a session in which the first natural language query was provided to the generative Al model, causing the synthetic data generator to identify the first query language query based at least on the session identifier and to determine the first query language query corresponds to the first natural language query.
However, Jain teaches provide, to the synthetic data generator, a session identifier identifying a session in which the first natural language query was provided to the generative Al model, causing the synthetic data generator to identify the first query language query based at least on the session identifier and to determine the first query language query corresponds to the first natural language query (see col. 15, lines 17-31, teaches initiating a session with an LLM, receiving a query having a natural-language portion, prompting the LLM “in the session” with the processed natural-language portion, and receiving native database query content from the LLM; see also col. 17, lines. 48-58, teaches “a user-specific inner session identifier” and retrieving context for a user using the user-specific inner session identifier before processing a natural-language query for that user), and the program code is further structured to cause the processor circuit to: identify the first query language query based at least on the session identifier (see col. 17, lines 48-58, teaches retrieving user context based on a user-specific inner session identifier; see also col. 19, lines 56-67 through col. 20, lines 1-8, teaches treating a database session as an LLM conversation and maintaining a history of past prompts and responses in the current database session); and determine the first query language query corresponds to the first natural language query (see col. 15, lines 17-31, teaches receiving a query having a natural-language portion, prompting the LLM in the session with that processed natural-language portion, and receiving native database query content from the LLM).
Thus, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to apply the teachings of Jain with the teachings of the Dharnidharka, Truong, Relan and Bista in order to prevent generation or execution of queries against unauthorized or unavailable database structures by limiting natural-language query processing according to user privileges and database structures available to the user (Jain).
It is noted that any citation [[s]] to specific, pages, columns, lines, or figures in the prior art references and any interpretation of the references should not be considered to be limiting in any wav. A reference is relevant for all it contains and may be relied upon for all that it would have reasonably suggested to one having ordinary skill in the art. [[See, MPEP 2123]].
Pertinent Prior Art
The prior art made of record and not relied upon is considered pertinent to
applicant's disclosure.
Mostafazadeh et al US 20220343903 A1 discloses a system and method that enable data-informed decision making for anyone without the need for any coding ability are disclosed. This artificial intelligence platform has domain-generality, interoperability across heterogeneous sources of data, and controllability by tracking provenance. The artificial intelligence platform works by receiving a natural language query, converts the natural language query into executable code grounded in the deep semantic understanding of the underlying data, using a natural language artificial intelligence engine, runs the executable code on a distributed runtime engine to generate data output, and augments the data with a generated natural language report which becomes the ultimate output to the user.
Lee et al US 20230315856 A1 discloses a processor receives natural language data for performing an identified cybersecurity task. The processor can provide the natural language data to a first machine learning (ML) model. The first ML model can automatically infer a template query based on the natural language data. The processor can receive user input indicating a finalized query and to provide the finalized query as input to a system configured to perform the identified computational task. The processor can provide the finalized query as a reference phrase to a second ML model, the second ML model configured to generate a set of natural language phrases similar to the reference phrase. The processor can generate supplemental training data using the set of natural language phrases similar to the reference phrase to augment training data used to improve performance of the first ML model and/or the second ML model.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Mohammad A Sana whose telephone number is (571)270-1753. The examiner can normally be reached Monday-Friday 9-5.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Sanjiv Shah can be reached at 5712724098. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/Mohammad A Sana/Primary Examiner, Art Unit 2166