DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim(s) 1-20 is/are pending and has/have been examined.
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 01/10/2025 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
The Examiner notes that most of the submitted Gupta reference is covered by a subscription pop-up, making it difficult to determine what of the reference, if anything, is relevant. As such, that particular reference was not considered.
Drawings
The drawings are objected to because of the following informalities:
Fig. 6 - element 621 is not in the spec
Fig. 8 - element 216 is not in the spec
Corrected drawing sheets in compliance with 37 CFR 1.121(d) are required in reply to the Office action to avoid abandonment of the application. Any amended replacement drawing sheet should include all of the figures appearing on the immediate prior version of the sheet, even if only one figure is being amended. The figure or figure number of an amended drawing should not be labeled as “amended.” If a drawing figure is to be canceled, the appropriate figure must be removed from the replacement sheet, and where necessary, the remaining figures must be renumbered and appropriate changes made to the brief description of the several views of the drawings for consistency. Additional replacement sheets may be necessary to show the renumbering of the remaining figures. Each drawing sheet submitted after the filing date of an application must be labeled in the top margin as either “Replacement Sheet” or “New Sheet” pursuant to 37 CFR 1.121(d). If the changes are not accepted by the examiner, the applicant will be notified and informed of any required corrective action in the next Office action. The objection to the drawings will not be held in abeyance.
Claim Objections
Claims 2, 3, 9, 13, and 17 are objected to because of the following informalities:
Claims 2 and 3 recite “a generative LLM”, and claim 2 recites “the LLM”. The Examiner suggests amending the claim(s) to recite –the generative LLM-- in order to maintain clear antecedent basis.
Claim 9 recites “a text-based query” in line 3, and “the query” in line 5. The Examiner suggests amending the claim(s) to recite –the text-based query-- in order to maintain clear antecedent basis.
Claim 13 recites “the query” in line 4. The Examiner suggests amending the claim(s) to recite –the text-based query-- in order to maintain clear antecedent basis.
Claim 17 recites “the pre-trained LLM” and “the LLM”. The Examiner suggests amending the claim to recite language consistent with whatever amendments are made/final claim language determinations that address the 112b rejection below for claim 16. Please see the corresponding rejection below for further detail.
Appropriate correction is required.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 16-19 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claim 16 recites “a pre-trained Language Learning Model (LLM) within the GAN pipeline”, which leads to unclear antecedent basis. In claim 1, an LLM is already recited (“a generative Language Learning Model (LLM)”) that performs the same functionality of generating training samples. It is unclear whether the pretrained LLM is supposed to be the same or different from the generative LLM of the independent claim. In the interest of compact prosecution, the Examiner is treating the two LLMs as being the same LLM, where claim 16 specifies that the LLM is part of a GAN architecture.
Claims 17-19 are rejected as being dependent upon a rejected claim.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Regarding claim(s) 1 and 20, the limitation(s) of receiving, performing, generating, labeling, and aggregating, as drafted, are processes that, under broadest reasonable interpretation, covers performance of the limitation in the mind and/or with pen and paper but for the recitation of generic computer components. More specifically, the mental process of a human looking at a set of sentences, determining how to change the wording of the sentences so that they maintain their meaning, using a learned set of rules to write out the re-worded sentences, recognizing entities in the each sentence and writing the type of entity next to the word, and writing everything out in an organized list on a piece of paper for future use. The generative LLM reads on a set of rules a human learns to understand and process human language to obtain specific results. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind and/or with pen and paper but for the recitation of generic computer components, then it falls within the --Mental Processes-- grouping of abstract ideas. Accordingly, the claim(s) recite(s) an abstract idea.
This judicial exception is not integrated into a practical application because the recitation of a system, processor, and memory of claim 20 reads to generalized computer components, based upon the claim interpretation wherein the structure is interpreted using [0030-49] in the specification. Accordingly, these additional elements do not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claim(s) is/are directed to an abstract idea.
The claim(s) do(es) not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to the integration of the abstract idea into a practical application, the additional element of using generalized computer components to receive, perform, generate, label, and aggregate, amounts to no more than mere instructions to apply the exception using a generic computer component. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. The claim(s) is/are not patent eligible.
With respect to claim(s) 2, the claim(s) recite(s) training the generative LLM with a prompt, which reads on a human learning how to re-word sentences using specific instructions. No additional limitations are present.
With respect to claim(s) 3, the claim(s) recite(s) training an LLM to produce multiple versions, which reads on a human learning how to re-word a sentence in several different ways. No additional limitations are present.
With respect to claim(s) 4 and 5, the claim(s) recite(s) applying a random noise/rephrasing function, which reads on a human performing a specific action to re-word the sentence. No additional limitations are present.
With respect to claim(s) 6, the claim(s) recite(s) applying a function, which reads on a human re-writing the re-worded sentences with placeholder terms and listing potential words that would fit where the placeholder term is beneath it. No additional limitations are present.
With respect to claim(s) 7, the claim(s) recite(s) replacing, which reads on a human writing out the re-worded sentences, where each sentence has a different potential word from the list of words in it. No additional limitations are present.
With respect to claim(s) 8, the claim(s) recite(s) the expanded labeled dataset is used to train models, which reads on a human using the list of labeled sentences to learn another language-related task. No additional limitations are present.
With respect to claim(s) 9, the claim(s) recite(s) receiving, interpreting, converting, mapping, executing, and communicating, which reads on a human looking at a written request for information, using a learned set of rules to identify entities, writing out the request in a specific format to include the entity information, writing out the request with the information in a format that makes it easier to find matching information in a reference, looking through the reference, and writing out the found information to show the requesting person. The NER model reads to a learned set of rules for performing a specific language-related task. The recitation of a user device reads on a generalized computer component as per [0030-49] in the specification.
With respect to claim(s) 10, the claim(s) recite(s) tokenizing the query, which reads on a human segmenting the request in a specific manner and using learned rules to write down an entity type for each section. No additional limitations are present.
With respect to claim(s) 11, the claim(s) recite(s) classifies, which reads on a human using learned rules to write a previously-seen category for each entity. No additional limitations are present.
With respect to claim(s) 12, the claim(s) recite(s) receiving, using…to select, and feeding, which reads on a human looking at the set of re-worded sentences, using a learned set of rules to identify specific sentences to use in the next step, and using the selected sentences to learn and improve on a set of rules for identifying entities in text. The generator, GAN pipeline, and NER model, each read to a different learned set of rules for performing a specific language-related task. No additional limitations are present.
With respect to claim(s) 13, the claim(s) recite(s) optimizing, converting, mapping, and executing, which reads on a human improving the different learned sets of rules using specific information and methods, writing out a request for information in a specific format to include identified entity information, writing out the request with the information in a format that makes it easier to find matching information in a reference, and looking through the reference to find the pertinent information related to the request. No additional limitations are present.
With respect to claim(s) 14, the claim(s) recite(s) evaluating the authenticity, which reads on a human specific learned rules (i.e. the discriminator) to determine an original sentence versus a re-worded sentence. No additional limitations are present.
With respect to claim(s) 15, the claim(s) recite(s) the structured format has specific characteristics, which reads on a human writing out the sentence in a specific way. No additional limitations are present.
With respect to claim(s) 16, the claim(s) recite(s) generating, fine-tuning, evaluating, and updating, which reads on a human writing out sentences using specific learned rules (i.e. generator, GAN pipeline), using specific information to improve a set of learned rules (i.e. LLM) for writing out additional sentences, using another set of learned rules to determine how good the sentences are (i.e. discriminator, NER model), and improving one set of learned rules using the resulting determination. No additional limitations are present.
With respect to claim(s) 17, the claim(s) recite(s) generating, which reads on a human using specific information to adjust the set of learned rules and write out sentences with specific characteristics. No additional limitations are present.
With respect to claim(s) 18 and 19, the claim(s) recite(s) specific processes for updating and improving the generator, which reads on a human using specific information and steps to improve one of the learned sets of rules. No additional limitations are present.
These claims further do not remedy the judicial exception being integrated into a practical application and further fail to include additional elements that are sufficient to amount to significantly more than the judicial exception.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-5, 8, and 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Lev (U.S. PG Pub No. 2025/0021766), hereinafter Lev, in view of Rossiello et al. (U.S. PG Pub No. 2024/0202447), hereinafter Rossiello.
Regarding claims 1 and 20, Lev teaches
(claim 1) A method for generating a training dataset for training a model (a method of enriching a dataset for training [0051]), comprising:
(claim 20) A system for generating a training dataset for training a model (a system for augmenting a dataset [0005],[0056]), comprising:
(claim 20) a processor (the computing environment includes a processor [0056-9]);
(claim 20) a memory operatively connected to the processor and storing instructions which, when executed by the processor, cause the system to perform (the computer includes storage media storing instructions that are accessed by the processor to perform the methods [0056-9]):
receiving a set of input samples (textual contents may be received, and a plurality of original text items, i.e. set of input samples, may be extracted from the textual contents, i.e. receiving [0078-9],[0103]);
performing a rephrasing operation to produce new versions of the set of input samples, wherein the new versions preserve semantic equivalence as the set of input samples but have different phrasing (the text items are combined and processed with a plurality of prompts to generate a plurality of queries, where the query is an instruction to rephrase the text item, i.e. performing a rephrasing operation…of the set of input samples, to generate new text similar to the prior one but without specific words/phrases, in a specific dialect, change the length, language style change emotional tone, use slang, write similarly to some famous writer's style and/or the like, thus generating a different text having similar meaning, but with different vocabularies and styles, i.e. produce new versions of the set of input samples wherein the new versions preserve semantic equivalence as the set of input samples but have different phrasing [0043],[0083],[0088],[0096],[0101-4],[0109-10]);
generating a dataset of generated versions of the input samples using a generative Language Learning Model (LLM) (the conversational language model may be a generative AI/LLM, i.e. using a generative Language Learning Model (LLM), may receive prompts comprising instructions to rephrase one or more textual contents, where the model generates a plurality of synthetic text items, i.e. generating a dataset of generated versions of the input samples Fig. 2,[0043],[0073],[0085],[0088],[0092-6],[0104]).
While Lev provides generating augmented datasets of rephrased text with an LLM, Lev does not specifically teach labeling entity references in text, and thus does not teach
labeling all entity references present in the generated versions of the input samples; and
aggregating the generated versions of the input samples and their corresponding labeled versions to form an expanded labeled dataset.
Rossiello, however, teaches labeling all entity references present in the generated versions of the input samples (text from a document of a document database, i.e. generated versions of the input samples, is aligned with entities from a knowledge graph, and the dataset is enhanced to include a named entity label associated with the named entity, i.e. labeling all entity references present [0025-6]); and
aggregating the generated versions of the input samples and their corresponding labeled versions to form an expanded labeled dataset (an enhanced training dataset of labeled text is generated using the updated training dataset that includes the named entity labels for the document text, i.e. aggregating…their corresponding labeled versions to form an expanded labeled dataset, and where the training dataset also comprises alignment data that maps unstructured text to corresponding triples, i.e. aggregating the generated versions of the input samples [0025-27]).
Lev and Rossiello are analogous art because they are from a similar field of endeavor in developing training datasets for training language models. Thus, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify the generating augmented datasets of rephrased text with an LLM teachings of Lev with labeling training datasets with named entity labels to generate an enhanced dataset as taught by Rossiello. It would have been obvious to combine the references to reduce error propagation and improve performance for NER and RE tasks (Rossiello [0024]).
Regarding claim 2, Lev in view of Rossiello teaches claim 1, and Lev further teaches
training a generative LLM in-context by providing the LLM with a prompt that instructs the LLM to create rephrased versions of a target sample (the text items are combined and processed with a plurality of prompts to generate a plurality of queries, where the query is an instruction for the LLM to rephrase the text item to generate new text similar to the prior one but without specific words/phrases, in a specific dialect, change the length, language style change emotional tone, use slang, write similarly to some famous writer's style and/or the like, thus generating a different text having similar meaning, but with different vocabularies and styles, i.e. providing the LLM with a prompt that instructs the LLM to create rephrased versions of a target sample, where the LLM is trained to augment the text phrases, i.e. training a generative LLM in-context [0043],[0083],[0085-6],[0088],[0092-6],[0101-4],[0109-10]).
Regarding claim 3, Lev in view of Rossiello teaches claim 1, and Lev further teaches
training a generative LLM in-context to produce multiple rephrased versions of a single input sentence (the text items are combined and processed with a plurality of prompts to generate a plurality of queries, where the query is an instruction for the LLM to rephrase the text item to generate a plurality of new text similar to the prior one but without specific words/phrases, in a specific dialect, change the length, language style change emotional tone, use slang, write similarly to some famous writer's style and/or the like, thus generating a different text having similar meaning, but with different vocabularies and styles, i.e. produce multiple rephrased versions of a single input sentence, where the LLM is trained to augment the text phrases, i.e. training a generative LLM in-context [0043],[0083],[0085-6],[0088],[0092-6],[0101-4],[0109-10]).
Regarding claim 4, Lev in view of Rossiello teaches claim 1, and Lev further teaches
generating a modified version of an input sample by applying a random noise based on a random parameter (iterations on model parameters such as temperature may be applied to change the randomness, i.e. applying a random noise based on a random parameter, of the variants of the rephrased text item generated by the model, i.e. generating a modified version of an input sample [0042-4],[0088],[0090-5]).
Regarding claim 5, Lev in view of Rossiello teaches claim 1, and Lev further teaches
generating a modified version of an input sample by applying a rephrasing function based on a random parameter (iterations on model parameters such as temperature may be applied to change the randomness, i.e. applying a rephrasing function based on a random parameter, of the variants of the rephrased text item generated by the model as instructed by the prompt, i.e. generating a modified version of an input sample by applying a rephrasing function [0042-4],[0088],[0090-5]).
Regarding claim 8, Lev in view of Rossiello teaches claim 1, and Rossiello further teaches
the expanded labeled dataset is used to train models in text-to-structured tasks (the enhanced training dataset, i.e. expanded labeled dataset, is used to train an NLP model, i.e. is used to train models, where the trained model converts unstructured text into structured data, i.e. train models in text-to-structured tasks [0028],[0076-7]).
Where the motivation to combine is the same as previously presented.
Claim(s) 6 and 7 is/are rejected under 35 U.S.C. 103 as being unpatentable over Lev, in view of Rossiello, and further in view of Hwang et al. (U.S. PG Pub No. 2015/0127319), hereinafter Hwang.
Regarding claim 6, Lev in view of Rossiello teaches claim 1.
While Lev in view of Rossiello provides prompting the LLM to avoid specific keywords in a rephrasing, Lev in view of Rossiello does not specifically teach replacing entity placeholders with potential values, and thus does not teach
applying a function to the generated versions of the input samples and corresponding placeholders for entity values, where the function replaces the corresponding placeholders with a list of potential values.
Hwang, however, teaches applying a function to the generated versions of the input samples and corresponding placeholders for entity values, where the function replaces the corresponding placeholders with a list of potential values (slot-tag abstraction is performed on the training data, where the slot labels and slot values of the training data, i.e. applying a function to the generated versions of the input samples and corresponding placeholders for entity values, are replaced with an abstract token that may be further replaced with one or more entities determined from content sources as corresponding to the slot type, i.e. function replaces the corresponding placeholders with a list of potential values [0018-21],[0029],[0033]).
Where Lev specifically teaches the training data are generated rephrasing of text items Fig. 2,[0043],[0073],[0085],[0088],[0092-6],[0104].
Lev, Rossiello, and Hwang are analogous art because they are from a similar field of endeavor in developing training datasets for training language models. Thus, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify the prompting the LLM to avoid specific keywords in a rephrasing teachings of Lev, as modified by Rossiello, with slot-tag abstraction to replace a single slot value with more than one as taught by Hwang. It would have been obvious to combine the references to enable the generation of larger amounts of training data (Hwang [0030]).
Regarding claim 7, Lev in view of Rossiello and Hwang teaches claim 6, and Hwang further teaches
the function replaces the corresponding placeholders with actual values for a list of potential values for each entity (slot-tag abstraction is performed on the training data, where the slot labels and slot values of the training data, i.e. applying a function to the generated versions of the input samples and corresponding placeholders for entity values, are replaced with an abstract token that may be further replaced with one or more entities determined from content sources as corresponding to the slot type, i.e. function replaces the corresponding placeholders with actual values for a list of potential values for each entity [0018-21],[0029],[0033]).
Where the motivation to combine is the same as previously presented.
Claim(s) 9-11 is/are rejected under 35 U.S.C. 103 as being unpatentable over Lev, in view of Rossiello, and further in view of Nezami et al. (U.S. PG Pub No. 2024/0062112), hereinafter Nezami.
Regarding claim 9, Lev in view of Rossiello teaches claim 1, and Rossiello further teaches
converting a text-based query into an executable database query by:
receiving a text-based query (unstructured text data from a document is prepared [0075]);
interpreting the text-based query using a pre-trained Named Entity Recognition (NER) model to classify entities within the query thereby generating identified entities, wherein the NER model is trained on the expanded labeled dataset (the trained model, such as a model for NER, converts unstructured text, i.e. interpreting the text-based query using a pre-trained Named Entity Recognition (NER) model, into structured data by generating surface forms, entity labels, and type information in identified entities, i.e. classify entities within the query thereby generating identified entities [0024],[0075-7], where the NER is trained using the enhanced training dataset, i.e. the NER model is trained on the expanded labeled dataset [0024-30]);
converting the identified entities into a predetermined standardized format to create a structured representation of the text-based query (the target sequence is output according to a specific schema, i.e. create a structured representation of the text-based query, to represent the semantic annotations of the entities, i.e. converting the identified entities into a predetermined standardized format [0073],[0075-7]);
mapping the structured representation to a query format compatible with a target database to generate an executable query (the target sequence can be further encoded into a feature vector, i.e. mapping the structured representation to a query format, where semantic triples are similarity encoded in a structured database for comparison, i.e. compatible with a target database to generate an executable query [0075-9]).
While Lev in view of Rossiello provides converting unstructured text into structured data for further processing, Lev in view of Rossiello does not specifically teach performing a search and providing the response to a user, and thus does not teach
executing the executable query on the target database to perform a requested search or transaction; and
communicating a response from the target database back to a user device for presentation to a user.
Nezami, however, teaches executing the executable query on the target database to perform a requested search or transaction (when a user types in a query, i.e. requested search, the NER model predicts a class label for a token and outputs an utterance labeled with the predicted class labels that identify the named entities, i.e. executable query, where the system uses the labeled output utterance to perform an operation such as query a database, i.e. executing the executable query on the target database [0029],[0031],[0128],[0155]); and
communicating a response from the target database back to a user device for presentation to a user (the system uses the labeled output utterance to query a database, generate a response, i.e. response from the target database, and display the response on a display device of a client device, i.e. communicating a response…to a user device for presentation to a user [0029],[0144],[0155],[0159]).
Lev, Rossiello, and Nezami, are analogous art because they are from a similar field of endeavor in developing training datasets for training language models. Thus, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify the converting unstructured text into structured data for further processing teachings of Lev, as modified with Rossiello, with the use of the trained NER model to output a labeled utterance for querying a database as taught by Nezami. It would have been obvious to combine the references to improve the performance of NER models trained on adaptively augmented training data (Nezami [0034]).
Regarding claim 10, Lev in view of Rossiello and Nezami teaches claim 9, and Nezami further teaches
the text-based query is tokenized into tokens, and the NER model tags each token with corresponding entity labels (the utterance may be tokenized into tokens, i.e. the text-based query is tokenized into tokens, and the NER model analyzes the utterance and tokens to predict a class label/value for the named entities for each token, i.e. the NER model tags each token with corresponding entity labels [0029],[0144],[0155]).
Where the motivation to combine is the same as previously presented.
Regarding claim 11, Lev in view of Rossiello and Nezami teaches claim 9, and Nezami further teaches
the NER model classifies the identified entities into respective categories based on labels used during training of the NER model (the utterance may be tokenized into tokens, and the NER model analyzes the utterance and tokens to predict a class label/value for the named entities for each token, i.e. the NER model classifies the identified entities into respective categories, where the model has been trained with adaptively augmented training data that has a balance of distributions of named entity category labels represented in the data, i.e. respective categories based on labels used during training of the NER model [0029],[0134-5],[0139],[0155]).
Where the motivation to combine is the same as previously presented.
Claim(s) 12, 14-16, and 19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Lev, in view of Rossiello, and further in view of Zhang et al. (U.S. PG Pub No. 2025/0103813), hereinafter Zhang.
Regarding claim 12, Lev in view of Rossiello teaches claim 1, and Rossiello further teaches
receiving the expanded labeled dataset, wherein the expanded labeled dataset comprises rephrased versions of the set of input samples (the enhanced training dataset is generated, i.e. receiving the expanded labeled dataset [0026]).
Where Lev further teaches that the training dataset are rephrased versions of the textual contents Fig. 2,[0043],[0073],[0085],[0088],[0092-6],[0104].
While Lev in view of Rossiello provides the LLM may be a GAN network, Lev in view of Rossiello does not specifically teach using part of a GAN pipeline to select training samples, and thus does not teach
using a generator in a Generative Adversarial Network (GAN) pipeline to select particular samples from the rephrased versions of the set of input samples, thereby generating selected samples; and
feeding the selected samples along with corresponding entity values into a Named Entity Recognition (NER) model to train the NER model.
Zhang, however, teaches using a generator in a Generative Adversarial Network (GAN) pipeline to select particular samples from the rephrased versions of the set of input samples, thereby generating selected samples (a self-cleaning named entity recognition system includes a named entity recognition model manager and a discriminator model manager, i.e. GAN pipeline, where the entity recognition system reweights training sentences and may remove a training sentence from the training dataset or retrain the NER model with the reweighted training sentence, i.e. using a generator…to select particular samples…thereby generating selected samples, and the named entity recognition model manager trains the NER model and manages training data, i.e. using a generator Fig. 7,8,[0059],[0061-3],[0113]); and
feeding the selected samples along with corresponding entity values into a Named Entity Recognition (NER) model to train the NER model (the training sentences, which have ground truth entity labels, i.e. feeding the selected samples along with corresponding entity values, are used by an NER model to generate predicted labels for the training sentences during training, i.e. into a Named Entity Recognition (NER) model to train the NER model [0042],[0051],[0053],[0059],[0061-3]).
Where Lev specifically teaches that the training dataset are rephrased versions of the textual contents Fig. 2,[0043],[0073],[0085],[0088],[0092-6],[0104].
Lev, Rossiello, and Zhang, are analogous art because they are from a similar field of endeavor in developing training datasets for training language models. Thus, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify the LLM may be a GAN network teachings of Lev, as modified by Rossiello, with the use of a generation and discrimination network to remove low quality training sentences as taught by Zhang. It would have been obvious to combine the references to improve NER learning on noisy training data and guide an NER model’s training (Zhang [0064]).
Regarding claim 14, Lev in view of Rossiello and Zhang teaches claim 12, and Zhang further teaches
evaluating the authenticity of the selected samples, using a discriminator of the GAN pipeline, by distinguishing between real and generated data (the discriminator model classifies the output of another machine learning model as authentic or not authentic, i.e. evaluating the authenticity…using a discriminator of the GAN pipeline by distinguishing between real and generated data, where the discriminator model determines if the training sentence is authentic/correct, and training sentences with low scores have been removed from the training data set [0043],[0046-7],[0057-9]).
Where the motivation to combine is the same as previously presented.
Regarding claim 15, Lev in view of Rossiello and Zhang teaches claim 12, and Rossiello further teaches
the structured format of the query includes representing each identified entity as a key-value pair (the target sequence is output according to a specific schema, i.e. the structured format of the query, to represent the semantic annotations of the entities, including entities, types, and relation labels, and where the label-to-identifier map is stored as a key-value, i.e. representing each identified entity as a key-value pair [0073],[0075-7],[0086]).
Where the motivation to combine is the same as previously presented.
Regarding claim 16, Lev in view of Rossiello teaches claim 1, and Lev further teaches
generating synthetic text samples using a generator within a Generative Adversarial Network (GAN) pipeline (the conversational language model may be based on a GAN model, where the conversational language model generates, i.e. a generator within a Generative Adversarial Network (GAN) pipeline, a plurality of synthetic items that are rephrases of a text item, i.e. generating synthetic text samples [0092-3],[0104-7]);
fine-tuning, using a pre-trained Language Learning Model (LLM) within the GAN pipeline the LLM with…a subsequent model used to generate subsequent model output, the fine-tuning causing the LLM to generate more training samples for the subsequent model (model parameters of the conversational language model, such as an LLM, i.e. using a pre-trained Language Learning Model (LLM) within the GAN pipeline, can be used in a feedback loop in a trained ML model that classifies the rephrased text, i.e. subsequent model used to generate subsequent model output, where an indication that the lower accuracy of the synthetic training items may indicate the parameters of the LLM should be adjusted, and new text phrases augmented at the next iteration of the LLM, i.e. fine-tuning the LLM…causing the LLM to generate more training samples for the subsequent model [0043-4],[0092-5],[0104-7]);
evaluating a quality of –output-- performed by the subsequent model … (the accuracy of the classification of the synthetic text performed by a ML model trained on the augmented data, i.e. –output-- performed by the subsequent model, is used to indicate whether the text item was altered to a lesser representation of the associated label, i.e. evaluating the quality of output [0043-4],[0048-50],[0092-5],[0104-7]); and
updating the generator based on the evaluation of the generated samples …, the updating involving modifying internal parameters or changing a prompt to produce alternative samples (model parameters of the conversational language model, such as an LLM, i.e. internal parameters, can be used in a feedback loop in a trained ML model that classifies the rephrased text, where an indication that the lower accuracy of the synthetic training items, i.e. based on the evaluation of the generated samples, may indicate the parameters of the LLM should be adjusted, such as by lowering the temperature, i.e. updating the generator…involving modifying internal parameters [0043-4],[0048-50],[0092-5],[0104-7]).
While Lev in view of Rossiello provides the LLM may be a GAN network, Lev in view of Rossiello does not specifically teach using a discriminator including an NER model, and thus does not teach
Zhang, however, teaches fine-tuning, … with an inverse loss function of a subsequent model… (the predicted label loss for the NER model is used with the discriminator weight, which is determined using the predicted label from the NER model, i.e. with an inverse loss function of a subsequent model, to reweight or remove training sentences based on an identification of the training data as incorrect [0042-3],[0045],[0049],[0051]);
evaluating a quality of entity recognition performed by the subsequent model using a discriminator, wherein the discriminator includes the NER model (the predicted label loss for the NER model, i.e. quality of entity recognition performed by the subsequent model, is used with the discriminator weight, which is determined using the predicted label from the NER model, i.e. the discriminator includes the NER model, to reweight the predicted label loss of the NER model, i.e. evaluating a quality of entity recognition performed by the subsequent model using a discriminator [0042-3],[0045],[0049],[0051]); and
updating the generator based on the evaluation of the generated samples by the discriminator…(the predicted label loss for the NER model is used with the discriminator weight, which is determined using the predicted label from the NER model, i.e. based on the evaluation of the generated samples by the discriminator, to reweight or remove training sentences based on an identification of the training data as incorrect [0042-3],[0045],[0049],[0051]).
Where Lev teaches that the accuracy of the synthetic training items and a feedback loop with the trained ML model is used to adjust the parameters of the LLM [0043-4],[0092-5],[0104-7].
Lev, Rossiello, and Zhang, are analogous art because they are from a similar field of endeavor in developing training datasets for training language models. Thus, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify the LLM may be a GAN network teachings of Lev, as modified by Rossiello, with the use of a generation and discrimination network to remove or reweight low quality training sentences as taught by Zhang. It would have been obvious to combine the references to improve NER learning on noisy training data and guide an NER model’s training (Zhang [0064]).
Regarding claim 19, Lev in view of Rossiello and Zhang teaches claim 16, and Lev further teaches
the GAN pipeline includes an iterative cycle of generating new samples, evaluating them using the discriminator, and updating the generator based on the evaluation (model parameters of the conversational language model, such as an LLM, i.e. GAN pipeline, can be used in a series of iterations of generating augmentations, i.e. includes an iterative cycle of generating new samples, along with a feedback loop in a trained ML model that classifies the rephrased text, where an indication that the lower accuracy of the synthetic training items, i.e. evaluating them, may indicate the parameters of the LLM should be adjusted, such as by lowering the temperature, i.e. updating the generator based on the evaluation [0043-4],[0048-50],[0092-5],[0104-7]).
Where Zhang further teaches that the evaluation of the training sentences is specifically performed by a discriminator [0042-3],[0045],[0049],[0051].
And where the motivation to combine is the same as previously presented.
Claim(s) 13 and 17 is/are rejected under 35 U.S.C. 103 as being unpatentable over Lev, in view of Rossiello, in view of Zhang, and further in view of Nezami.
Regarding claim 13, Lev in view of Rossiello and Zhang teaches claim 12, and Zhang further teaches
optimizing the generator and the NER model through backpropagation using a loss function (training sentences are managed based on reweighting based on discriminative reweight loss, where the named recognition entity model manager trains the NER model utilizing predicted label loss and backpropagation, i.e. optimizing the generator and the NER model through backpropagation using a loss function [0061]).
Where Rossiello teaches converting, using the trained NER model, a text-based query into a structured format by identifying and classifying entities within the text-based query (the trained model, such as a model for NER, converts unstructured text, i.e. converting using the trained NER model a text-based query, into structured data by generating surface forms, entity labels, and type information in identified entities, i.e. a structured format by identifying and classifying entities within the text-based query [0024],[0075-7]);
mapping the structured representation of the query to a query format compatible with a target database (the target sequence can be further encoded into a feature vector, i.e. mapping the structured representation of the query to a query format, where semantic triples are similarity encoded in a structured database for comparison, i.e. compatible with a target database [0075-9]).
While Lev in view of Rossiello and Zhang provides converting unstructured text into structured data for further processing, Lev in view of Rossiello and Zhang does not specifically teach performing a search and providing the response to a user, and thus does not teach
executing the mapped query on the target database to perform a requested search or transaction.
Nezami, however, teaches executing the mapped query on the target database to perform a requested search or transaction (when a user types in a query, i.e. requested search, the NER model predicts a class label for a token and outputs an utterance labeled with the predicted class labels that identify the named entities, i.e. mapped query, where the system uses the labeled output utterance to perform an operation such as query a database, i.e. executing the executable query on the target database [0029],[0031],[0128],[0155]).
Lev, Rossiello, Zhang, and Nezami, are analogous art because they are from a similar field of endeavor in developing training datasets for training language models. Thus, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify the converting unstructured text into structured data for further processing teachings of Lev, as modified with Rossiello and Zhang, with the use of the trained NER model to output a labeled utterance for querying a database as taught by Nezami. It would have been obvious to combine the references to improve the performance of NER models trained on adaptively augmented training data (Nezami [0034]).
Regarding claim 17, Lev in view of Rossiello and Zhang teaches claim 16, and Lev further teaches
generating, using the pre-trained LLM, synthetic text samples that resemble the initial known dataset by fine-tuning the --parameters-- within the LLM causing a change in a prompt (the text items are combined and processed with a plurality of prompts to generate a plurality of queries, where the query is an instruction for the LLM to rephrase the text item to generate new text similar to the prior one, i.e. generating using the pre-trained LLM synthetic text samples that resemble the initial known dataset, where the LLM is trained to augment the text phrases, i.e. pre-trained LLM, and where the parameters of the language model may be adjusted in different iterations based on accuracy of the synthetic items, i.e. fine-tuning the --parameters-- within the LLM, and each iteration may include a prompt that optionally adjusts model parameters, i.e. causing a change in prompt [0043-4],[0083],[0085-6],[0088],[0092-6],[0101-4],[0109-10]).
While Lev in view of Rossiello and Zhang provides adjusting model parameters, Lev in view of Rossiello and Zhang does not specifically teach the parameters include weights and biases, and thus does not teach
fine-tuning the weights and biases within the LLM….
Nezami, however, teaches fine-tuning the weights and biases within the LLM… (model training component trains the prediction model by selecting model parameters such as weights and biases for the prediction model [0126]).
Where Lev teaches the updated model parameters are for the LLM [0043-4],[0085-6],[0092-6].
Lev, Rossiello, Zhang, and Nezami are analogous art because they are from a similar field of endeavor in developing training datasets for training language models. Thus, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify the adjusting model parameters teachings of Lev, as modified with Rossiello and Zhang, with the parameters specifically including weights and biases as taught by Nezami. It would have been obvious to combine the references to improve the performance of NER models trained on adaptively augmented training data (Nezami [0034]).
Claim(s) 18 is/are rejected under 35 U.S.C. 103 as being unpatentable over Lev, in view of Rossiello, in view of Zhang, and further in view of Nguyen et al. (U.S. Patent No. 12,412,037), hereinafter Nguyen.
Regarding claim 18, Lev in view of Rossiello and Zhang teaches claim 16, and Lev further teaches
the --accuracy determination-- is propagated back through the GAN pipeline to the generator, …indicating parameters of the generator to be adjusted (model parameters of the conversational language model, such as an LLM, i.e. parameters of the generator, can be used in a feedback loop in a trained ML model that classifies the rephrased text, where an indication that the lower accuracy of the synthetic training items may indicate the parameters of the LLM should be adjusted, such as by lowering the temperature, i.e. --accuracy determination-- is propagated back through the GAN pipeline to the generator…indicating parameters of the generator to be adjusted [0043-4],[0048-50],[0092-5],[0104-7]).
Where Zhang further teaches the inverse loss function is propagated back…(the predicted label loss for the NER model is used with the discriminator weight, which is determined using the predicted label from the NER model, i.e. inverse loss function, to reweight or remove training sentences based on an identification of the training data as incorrect, is propagated back [0042-3],[0045],[0049],[0051]).
While Lev in view of Rossiello and Zhang provides using losses and adjusting model parameters, Lev in view of Rossiello and Zhang does not specifically teach the inverse loss function is propagated back providing a gradient, and thus does not teach
the inverse loss function is propagated back…providing a gradient indicating parameters of the –model-- to be adjusted.
Nguyen, however, teaches the inverse loss function is propagated back…providing a gradient indicating parameters of the –model-- to be adjusted (the model generation unit corrects the values of the parameters of the model, i.e. indicating parameters of the –model-- to be adjusted, based on the error calculated, where the model generation unit calculates an error gradient, i.e. inverse loss function…providing a gradient, that is propagated from the end to the beginning of the neural network, i.e. propagated back (18:35-44)).
Where Lev teaches the updated model parameters are for the LLM [0043-4],[0085-6],[0092-6].
Lev, Rossiello, Zhang, and Nguyen are analogous art because they are from a similar field of endeavor in training language models. Thus, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify the using losses and adjusting model parameters teachings of Lev, as modified by Rossiello and Zhang, with the propagation of an error gradient to correct model parameter values as taught by Nguyen. It would have been obvious to combine the references to enable training a model until an error dropped to a value equal to or less than a threshold value (Nguyen (18:35-56)).
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to NICOLE A K SCHMIEDER whose telephone number is (571)270-1474. The examiner can normally be reached 8:00 - 5:00 M-F.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Pierre-Louis Desir can be reached at (571) 272-7799. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/NICOLE A K SCHMIEDER/Primary Examiner, Art Unit 2659