DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Title
The title of the invention is not descriptive. A new title is required that is clearly indicative of the invention to which the claims are directed. Examiner believes that the title of the invention is imprecise. A descriptive title indicative of the invention will help in proper indexing, classifying, searching, etc. See MPEP 606.01. However, the title of the invention should be limited to 500 characters. Examiner suggests including the aspect(s) of the claims which Applicant believes to be novel or nonobvious over the prior art.
Specification
The disclosure is objected to because of the following informalities:
Paragraph 3 appears to contain a typo in the following phrase, “Thus, rather than manually each document in the corpus,”
Appropriate correction is required.
Drawings
The drawings are objected to because of the following informalities.
Referring to Figure 5, box 525 contains the misspelling “EXMAPLE” and box 525 contains the misspelling “EMEBED”
Corrected drawings in compliance with 37 CFR 1.121(d) are required in reply to the Office action to avoid abandonment of the application. The replacement sheet(s) should be labeled “Replacement Sheet” in the page header (as per 37 CFR 1.84(c)) so as not to obstruct any portion of the drawing figures. If the changes are not accepted by the examiner, the applicant will be notified and informed of any required corrective action in the next Office action. The objection to the drawings will not be held in abeyance.
Claim Rejections - 35 USC § 112(b)
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 1-20 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claim 1 recites presenting, via one or more processors, a user interface coupled to a large language model (LLM) via which a user interfaces with the LLM. However, the second use of the term “via” is not a correct usage of the term, which renders the scope of the claim unclear. “Via” in this context essentially means “by means of.” While via one or more processors indicates how presenting the UI is done, which a user interfaces with the LLM does not indicate how presenting a user interface coupled to a large language model (LLM) is done.
Claim 1 recites causing, via the one or more processors, the LLM to generate additional synthetic data based upon the feedback. The term causing as used in this limitation is indefinite because it does not clearly indicate the scope of the limitation. The scope could include any Rube Goldberg machine type approach to causing the LLM to generate additional synthetic data. Additionally, a complete description of every way this could be done is not found in the specification.
Claim 9 recites causing, via the one or more processors, the LLM to generate further additional synthetic data. The term causing as used in this limitation is indefinite because it does not clearly indicate the scope of the limitation. The scope could include any Rube Goldberg machine type approach to causing the LLM to generate additional synthetic data. Additionally, a complete description of every way this could be done is not found in the specification.
For this reason, the above listed claims are rejected for containing this language or being dependent on a claim that contains this language.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1: Claim 1 is a method claim. Claim 11 is a system claim. Claim 20 is a CRM claim. Therefore, claims 1, 11, and 20 are directed to either a process, machine, manufacture or composition of matter.
With respect to Claim 1:
Step 2A Prong 1:
detecting, via the user interface, a request to generate synthetic data, the request indicating one or more characteristics of the synthetic data (mental process – user can manually detect a request to generate synthetic data, the request indicating one or more characteristics of the synthetic data)
detecting, via the user interface, feedback on the example synthetic data (mental process – user can manually detect feedback on the example synthetic data)
embedding, via the one or more processors, the additional synthetic data to generate an embedding space for training a classifier (mental process – user can manually embed the additional synthetic data to generate an embedding space for training a classifier)
Step 2A Prong 2: This judicial exception is not integrated into a practical application. Additional elements:
presenting, via one or more processors, a user interface coupled to a large language model (LLM) via which a user interfaces with the LLM (Adding insignificant extra-solution activity to the judicial exception - see MPEP 2106.05(g))
presenting, via one or more processors, a user interface coupled to a large language model (LLM) via which a user interfaces with the LLM (mere instructions to apply the exception using a generic computer component)
presenting, via one or more processors, a user interface coupled to a large language model (LLM) via which a user interfaces with the LLM (mere instructions to apply the exception using a generic computer component)
inputting, via the one or more processors, the request into the LLM to generate example synthetic data having the one or more characteristics (Adding insignificant extra-solution activity to the judicial exception - see MPEP 2106.05(g))
causing, via the one or more processors, the LLM to generate additional synthetic data based upon the feedback (Adding insignificant extra-solution activity to the judicial exception - see MPEP 2106.05(g))
Step 2B: The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. Additional elements:
presenting, via one or more processors, a user interface coupled to a large language model (LLM) via which a user interfaces with the LLM (MPEP 2106.05(d)(II) indicate that merely “storing and retrieving information in memory” or “receiving or transmitting data over a network” is a well‐understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed step is well-understood, routine, conventional activity is supported under Berkheimer)
presenting, via one or more processors, a user interface coupled to a large language model (LLM) via which a user interfaces with the LLM (mere instructions to apply the exception using a generic computer component)
presenting, via one or more processors, a user interface coupled to a large language model (LLM) via which a user interfaces with the LLM (mere instructions to apply the exception using a generic computer component)
inputting, via the one or more processors, the request into the LLM to generate example synthetic data having the one or more characteristics (MPEP 2106.05(d)(II) indicate that merely “storing and retrieving information in memory” or “receiving or transmitting data over a network” is a well‐understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed step is well-understood, routine, conventional activity is supported under Berkheimer)
causing, via the one or more processors, the LLM to generate additional synthetic data based upon the feedback (MPEP 2106.05(d)(II) indicate that merely “storing and retrieving information in memory” or “receiving or transmitting data over a network” is a well‐understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed step is well-understood, routine, conventional activity is supported under Berkheimer)
Conclusion: The claim is not patent eligible.
Claims 11 and 20 are rejected on the same grounds as claim 1. Additionally for claims 11 and 20: Claim 11 has the additional elements of one or more processors and one or more memories storing non-transitory, computer readable instructions. These elements are mere instructions to apply the exception using a generic computer component under Step 2A prong 2 and Step 2B. Claim 20 has the additional element of a non-transitory computer-readable storage medium. This element is mere instructions to apply the exception using a generic computer component under Step 2A prong 2 and Step 2B.
Regarding Claim 2: The limitation(s), as drafted, are a process that, under its broadest reasonable interpretation, covers performance of the limitation(s) in the mind. That is, nothing in the claim limitation(s) precludes the step from practically being performed in the mind.
The limitation(s) encompasses the user manually wherein the one or more characteristics include one or more of a sentiment conveyed by the synthetic data, a topic referenced by the synthetic data, a format for the synthetic data, or a domain associated with the synthetic data.
These judicial exceptions are not integrated into a practical application. In particular, the claims do not recite any additional elements. Accordingly, this does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea.
The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, no additional elements are cited. Accordingly, the claim is not patent eligible.
Regarding Claim 3: The limitation(s), as drafted, are a process that, under its broadest reasonable interpretation, covers performance of the limitation(s) in the mind. That is, nothing in the claim limitation(s) precludes the step from practically being performed in the mind.
The limitation(s) encompasses the user manually wherein the request indicates a number of examples to include in the example synthetic data.
These judicial exceptions are not integrated into a practical application. In particular, the claims do not recite any additional elements. Accordingly, this does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea.
The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, no additional elements are cited. Accordingly, the claim is not patent eligible.
Regarding Claim 4: The limitation(s), as drafted, is a process that, under its broadest reasonable interpretation, covers performance of the limitation(s) in the mind. That is, other than the additional elements, nothing in the claim limitation(s) precludes the step from practically being performed in the mind.
The limitation(s) includes the additional elements of presenting, via the user interface, one or more elements that enable the user to provide the feedback as to a responsiveness of the example synthetic data to the one or more characteristics indicated by the request.
These judicial exceptions are not integrated into a practical application. The additional element(s) of presenting, via the user interface, one or more elements that enable the user to provide the feedback as to a responsiveness of the example synthetic data to the one or more characteristics indicated by the request recite adding insignificant extra-solution activity to the judicial exception - see MPEP 2106.05(g). Accordingly, this does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea.
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional element(s) of presenting, via the user interface, one or more elements that enable the user to provide the feedback as to a responsiveness of the example synthetic data to the one or more characteristics indicated by the request recite merely “storing and retrieving information in memory” or “receiving or transmitting data over a network” is a well‐understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim) (MPEP 2106.05(d)(II)). Thereby, a conclusion that the claimed storing step is well-understood, routine, conventional activity is supported under Berkheimer. Accordingly, the claims are not patent eligible.
Regarding Claim 5: The limitation(s), as drafted, are a process that, under its broadest reasonable interpretation, covers performance of the limitation(s) in the mind. That is, nothing in the claim limitation(s) precludes the step from practically being performed in the mind.
The limitation(s) encompasses the user manually wherein the classifier is (i) configured to identify documents most closely related to an inquiry, or (ii) configured to segment the embedding space into two or more segments.
These judicial exceptions are not integrated into a practical application. In particular, the claims do not recite any additional elements. Accordingly, this does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea.
The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, no additional elements are cited. Accordingly, the claim is not patent eligible.
Regarding Claim 6: The limitation(s), as drafted, is a process that, under its broadest reasonable interpretation, covers performance of the limitation(s) in the mind. That is, other than the additional elements, nothing in the claim limitation(s) precludes the step from practically being performed in the mind.
The limitation(s) encompasses the user manually mapping, via the one or more processors, the embedding space into a second embedding space generated for a corpus of documents.
The limitation(s) includes the additional elements of mapping, via the one or more processors, the embedding space into a second embedding space generated for a corpus of documents; and
tuning, via the one or more processors, the classifier based upon the second embedding space.
These judicial exceptions are not integrated into a practical application. The additional element(s) of via the one or more processors are recited at a high-level of generality such that it amounts no more than mere instructions to apply the exception using a generic computer component. The additional element(s) of tuning, via the one or more processors, the classifier based upon the second embedding space recite merely adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f). Accordingly, this does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea.
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional element(s) of via the one or more processors amount to no more than mere instructions to apply the exception using a generic computer component or operation. Mere instructions to apply an exception using a generic computer component or operation cannot provide an inventive concept. The additional element(s) of tuning, via the one or more processors, the classifier based upon the second embedding space recite adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f). Accordingly, the claims are not patent eligible.
Regarding Claim 7: The limitation(s), as drafted, are a process that, under its broadest reasonable interpretation, covers performance of the limitation(s) in the mind. That is, nothing in the claim limitation(s) precludes the step from practically being performed in the mind.
The limitation(s) encompasses the user manually wherein the corpus of documents includes privileged and/or confidential information.
These judicial exceptions are not integrated into a practical application. In particular, the claims do not recite any additional elements. Accordingly, this does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea.
The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, no additional elements are cited. Accordingly, the claim is not patent eligible.
Regarding Claim 8: The limitation(s), as drafted, is a process that, under its broadest reasonable interpretation, covers performance of the limitation(s) in the mind. That is, other than the additional elements, nothing in the claim limitation(s) precludes the step from practically being performed in the mind.
The limitation(s) encompasses the user manually applying, via the one or more processors, a logistic regression.
The limitation(s) includes the additional elements of applying, via the one or more processors, a logistic regression.
These judicial exceptions are not integrated into a practical application. The additional element(s) of via the one or more processors are recited at a high-level of generality such that it amounts no more than mere instructions to apply the exception using a generic computer component. Accordingly, this does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea.
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional element(s) of via the one or more processors amount to no more than mere instructions to apply the exception using a generic computer component or operation. Mere instructions to apply an exception using a generic computer component or operation cannot provide an inventive concept. Accordingly, the claims are not patent eligible.
Regarding Claim 9: The limitation(s), as drafted, is a process that, under its broadest reasonable interpretation, covers performance of the limitation(s) in the mind. That is, other than the additional elements, nothing in the claim limitation(s) precludes the step from practically being performed in the mind.
The limitation(s) encompasses the user manually validating, via the one or more processors, the classifier against one or more validation criteria;
determining, via the one or more processors, that the classifier does not satisfy the one or more validation criteria.
The limitation(s) includes the additional elements of via the one or more processors
causing, via the one or more processors, the LLM to generate further additional synthetic data.
These judicial exceptions are not integrated into a practical application. The additional element(s) of via the one or more processors are recited at a high-level of generality such that it amounts no more than mere instructions to apply the exception using a generic computer component. The additional element(s) of causing, via the one or more processors, the LLM to generate further additional synthetic data recite merely adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f). Accordingly, this does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea.
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional element(s) of via the one or more processors amount to no more than mere instructions to apply the exception using a generic computer component or operation. Mere instructions to apply an exception using a generic computer component or operation cannot provide an inventive concept. The additional element(s) of causing, via the one or more processors, the LLM to generate further additional synthetic data recite adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f). Accordingly, the claims are not patent eligible.
Regarding Claim 10: The limitation(s), as drafted, is a process that, under its broadest reasonable interpretation, covers performance of the limitation(s) in the mind. That is, other than the additional elements, nothing in the claim limitation(s) precludes the step from practically being performed in the mind.
The limitation(s) encompasses the user manually detecting, via the second user interface, a second request to generate synthetic data, the second request indicating one or more second characteristics of the synthetic data;
embedding, via the one or more processors, the second synthetic data to incorporate the embedded second synthetic data into the embedding space for tuning the classifier.
The limitation(s) includes the additional elements of presenting, via the one or more processors, a second user interface coupled to the LLM to a second user;
inputting, via the one or more processors, the request into the LLM to generate second synthetic data having the one or more second characteristics.
These judicial exceptions are not integrated into a practical application. The additional element(s) of via the one or more processors and a second user interface are recited at a high-level of generality such that it amounts no more than mere instructions to apply the exception using a generic computer component. The additional element(s) of presenting, via the one or more processors, a second user interface coupled to the LLM to a second user; and inputting, via the one or more processors, the request into the LLM to generate second synthetic data having the one or more second characteristics recite adding insignificant extra-solution activity to the judicial exception - see MPEP 2106.05(g). Accordingly, this does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea.
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional element(s) of via the one or more processors and a second user interface amount to no more than mere instructions to apply the exception using a generic computer component or operation. Mere instructions to apply an exception using a generic computer component or operation cannot provide an inventive concept. The additional element(s) of presenting, via the one or more processors, a second user interface coupled to the LLM to a second user; and inputting, via the one or more processors, the request into the LLM to generate second synthetic data having the one or more second characteristics recite merely “storing and retrieving information in memory” or “receiving or transmitting data over a network” is a well‐understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim) (MPEP 2106.05(d)(II)). Thereby, a conclusion that the claimed storing step is well-understood, routine, conventional activity is supported under Berkheimer. Accordingly, the claims are not patent eligible.
Claims 11-19 are rejected on the same grounds as claims 1-6, 8-10 respectively.
Claim 20 is rejected on the same grounds as claim 1.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-5, 9-15, 18-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Mannino et al. (hereinafter Mannino), Synner: Generating Realistic Synthetic Data, in view of Nagaraju et al. (hereinafter Nagaraju), U.S. Patent Application Publication 2024/0185001, further in view of Tang et al. (hereinafter Tang), Does Synthetic Data Generation of LLMs Help Clinical Text Mining?
Regarding Claim 1, Mannino discloses a method for generating synthetic data to train a machine learning model, the method comprising:
presenting, via one or more processors, a user interface coupled to a large language model (LLM) via which a user interfaces [“through a user-friendly UI” §1 ¶2; “the main interface features” §2 ¶1; Fig. 1] with the LLM [“LLM may then generate a conversational (e.g., natural language) query in response to the prompt that can be included in the generated dataset” §19; Figs. 1, 4B];
detecting, via the user interface, a request to generate synthetic data, the request indicating one or more characteristics of the synthetic data [“To generate her data, Marcia simply begins by adding a new column. As soon as she labels it Country, Synner suggests using a Country domain generator in the details pane.” §2 ¶3];
inputting, via the one or more processors, the request into the LLM [“LLM may then generate a conversational (e.g., natural language) query in response to the prompt that can be included in the generated dataset” §19; Figs. 1, 4B] to generate example synthetic data having the one or more characteristics [“To generate her data, Marcia simply begins by adding a new column. As soon as she labels it Country, Synner suggests using a Country domain generator in the details pane.” §2 ¶3];
presenting, via the user interface, the example synthetic data [“the preview column is populated with a few randomly selected countries. A histogram at the top of the column shows the distribution of different countries.” §2 ¶3; Fig. 1];
detecting, via the user interface, feedback on the example synthetic data [“Every user interaction that changes the properties of the data set automatically updates a declarative data generation script.” §2.1 ¶1; “Users can directly enter a few values to refine Synner’s data generation suggestions.” Fig. 1];
causing, via the one or more processors, the LLM [“LLM may then generate a conversational (e.g., natural language) query in response to the prompt that can be included in the generated dataset” §19; Figs. 1, 4B] to generate additional synthetic data based upon the feedback [“The engine returns a sample of the data to the front-end user interface for preview” §2.1 ¶2]; and
embedding, via the one or more processors, the additional synthetic data to generate an embedding space [“to obtain embeddings for both the original and synthetic data” §7 ¶1; Fig. 4] for training a classifier [“large language models (LLMs) (e.g., large generative language models) to generate robust and varied datasets that can be used to effectively train MLMs” ¶17; “machine-learning model (MLM)” ¶2].
However, Mannino fails to explicitly disclose presenting, via one or more processors, a user interface coupled to a large language model (LLM) via which a user interfaces with the LLM;
inputting, via the one or more processors, the request into the LLM to generate example synthetic data having the one or more characteristics;
causing, via the one or more processors, the LLM to generate additional synthetic data based upon the feedback; and
embedding, via the one or more processors, the additional synthetic data to generate an embedding space for training a classifier.
Nagaraju discloses presenting, via one or more processors, a user interface coupled to a large language model (LLM) via which a user interfaces with the LLM [“LLM may then generate a conversational (e.g., natural language) query in response to the prompt that can be included in the generated dataset” §19; Figs. 1, 4B];
inputting, via the one or more processors, the request into the LLM [“LLM may then generate a conversational (e.g., natural language) query in response to the prompt that can be included in the generated dataset” §19; Figs. 1, 4B] to generate example synthetic data having the one or more characteristics;
causing, via the one or more processors, the LLM [“LLM may then generate a conversational (e.g., natural language) query in response to the prompt that can be included in the generated dataset” §19; Figs. 1, 4B] to generate additional synthetic data based upon the feedback; and
embedding, via the one or more processors, the additional synthetic data to generate an embedding space for training a classifier [“large language models (LLMs) (e.g., large generative language models) to generate robust and varied datasets that can be used to effectively train MLMs” ¶17; “machine-learning model (MLM)” ¶2].
It would have been obvious to one having ordinary skill in the art, having the teachings of Mannino and Nagaraju before him before the effective filing date of the claimed invention, to modify the method of Mannino to incorporate the LLM for synthetic data generation of Nagaraju.
Given the advantage of efficiency and scalability, one having ordinary skill in the art would have been motivated to make this obvious modification.
However, Mannino fails to explicitly disclose embedding, via the one or more processors, the additional synthetic data to generate an embedding space for training a classifier.
Tang discloses embedding, via the one or more processors, the additional synthetic data to generate an embedding space [“to obtain embeddings for both the original and synthetic data” §7 ¶1; Fig. 4] for training a classifier.
It would have been obvious to one having ordinary skill in the art, having the teachings of Mannino, Nagaraju, and Tang before him before the effective filing date of the claimed invention, to modify the combination to incorporate the explicit embedding of Tang.
Given the advantage of changing raw data into interpretable data in machine learning, one having ordinary skill in the art would have been motivated to make this obvious modification.
Regarding Claim 2, Mannino, Nagaraju, and Tang disclose the method of claim 1. Mannino further discloses wherein the one or more characteristics include one or more of a sentiment conveyed by the synthetic data, a topic referenced by the synthetic data, a format for the synthetic data, or a domain associated with the synthetic data [“users can visually and declaratively specify properties of the dataset they wish to generate. Such properties include the domain, and statistical distribution of each field, and relationships between fields.” Abstract].
Regarding Claim 3, Mannino, Nagaraju, and Tang disclose the method of claim 1.
However, Mannino fails to explicitly disclose wherein the request indicates a number of examples to include in the example synthetic data.
Nagaraju discloses wherein the request indicates a number of examples to include in the example synthetic data [“process may continue until a target condition is satisfied, e.g., a certain number of NL queries have been generated, a certain number of task data and/or template queries have been selected, a certain amount of time has elapsed since starting the dataset generation process, or the like.” ¶50].
It would have been obvious to one having ordinary skill in the art, having the teachings of Mannino, Nagaraju, and Tang before him before the effective filing date of the claimed invention, to modify the combination to incorporate the number of examples of Nagaraju.
Given the advantage of customizing the generation a desired number of synthetic data examples, one having ordinary skill in the art would have been motivated to make this obvious modification.
Regarding Claim 4, Mannino, Nagaraju, and Tang disclose the method of claim 1. Mannino further discloses wherein presenting the example synthetic data comprises:
presenting, via the user interface, one or more elements that enable the user to provide the feedback as to a responsiveness of the example synthetic data to the one or more characteristics indicated by the request [“provides instant feedback on every user interaction by visualizing a preview of the generated data” Abstract; “Every user interaction that changes the properties of the data set automatically updates a declarative data generation script.” §2.1 ¶1; “The engine returns a sample of the data to the front-end user interface for preview as well as different summary statistics” §2.1 ¶2; Fig. 1].
Regarding Claim 5, Mannino, Nagaraju, and Tang disclose the method of claim 1.
However, Mannino fails to explicitly disclose wherein the classifier is (i) configured to identify documents most closely related to an inquiry, or (ii) configured to segment the embedding space into two or more segments.
Nagaraju discloses wherein the classifier is (i) configured to identify documents most closely related to an inquiry, or (ii) configured to segment the embedding space into two or more segments [“For example, training a natural language processing MLM specialized in a particular conversational domain (e.g., retail sales dialogues) typically involves many examples of relevant conversations in the retail sales space.” ¶2].
It would have been obvious to one having ordinary skill in the art, having the teachings of Mannino, Nagaraju, and Tang before him before the effective filing date of the claimed invention, to modify the combination to incorporate the classifier using the synthetic data of Nagaraju.
Given the advantage of training a useful model based on the synthetic data, one having ordinary skill in the art would have been motivated to make this obvious modification.
Regarding Claim 9, Mannino, Nagaraju, and Tang disclose the method of claim 1.
However, Mannino fails to explicitly disclose further comprising:
validating, via the one or more processors, the classifier against one or more validation criteria;
determining, via the one or more processors, that the classifier does not satisfy the one or more validation criteria; and
causing, via the one or more processors, the LLM to generate further additional synthetic data.
Nagaraju discloses further comprising:
validating, via the one or more processors, the classifier against one or more validation criteria [“once validated by system 1000 (e.g., for accuracy, etc.)” ¶86];
determining, via the one or more processors, that the classifier does not satisfy the one or more validation criteria [“may be repeated until the output error for a given training input satisfies a predetermined condition (e.g., falls below a predetermined value)” ¶31]; and
causing, via the one or more processors, the LLM to generate further additional synthetic data [“a different training input may be selected, a new output generated, and a new series of adjustments implemented, until conversational MLM 192 is trained to (e.g., converges to) a target degree of accuracy” ¶31].
It would have been obvious to one having ordinary skill in the art, having the teachings of Mannino, Nagaraju, and Tang before him before the effective filing date of the claimed invention, to modify the combination to incorporate validation process of Nagaraju.
Given the advantage of ensuring the classifier is accurate, one having ordinary skill in the art would have been motivated to make this obvious modification.
Regarding Claim 10, Mannino, Nagaraju, and Tang disclose the method of claim 1. Mannino further discloses further comprising:
presenting, via the one or more processors, a second user interface coupled to the LLM to a second user [“through a user-friendly UI” §1 ¶2; “the main interface features” §2 ¶1; “Users can directly enter a few values to refine Synner’s data generation suggestions” Fig. 1];
detecting, via the second user interface, a second request to generate synthetic data, the second request indicating one or more second characteristics of the synthetic data [“To generate her data, Marcia simply begins by adding a new column. As soon as she labels it Country, Synner suggests using a Country domain generator in the details pane.” §2 ¶3; “Synner provides a preview of the generated data in the spreadsheet cells, a histogram for each column as well as a few statistics. Users can directly enter a few values to refine Synner’s data generation suggestions.” Fig. 1];
inputting, via the one or more processors, the request into the LLM to generate second synthetic data having the one or more second characteristics [“To generate her data, Marcia simply begins by adding a new column. As soon as she labels it Country, Synner suggests using a Country domain generator in the details pane.” §2 ¶3; “Synner provides a preview of the generated data in the spreadsheet cells, a histogram for each column as well as a few statistics. Users can directly enter a few values to refine Synner’s data generation suggestions.” Fig. 1].
However, Mannino fails to explicitly disclose presenting, via the one or more processors, a second user interface coupled to the LLM to a second user;
inputting, via the one or more processors, the request into the LLM to generate second synthetic data having the one or more second characteristics; and
embedding, via the one or more processors, the second synthetic data to incorporate the embedded second synthetic data into the embedding space for tuning the classifier.
Nagaraju discloses presenting, via the one or more processors, a second user interface coupled to the LLM [“LLM may then generate a conversational (e.g., natural language) query in response to the prompt that can be included in the generated dataset” §19; Figs. 1, 4B] to a second user;
inputting, via the one or more processors, the request into the LLM [“LLM may then generate a conversational (e.g., natural language) query in response to the prompt that can be included in the generated dataset” §19; Figs. 1, 4B] to generate second synthetic data having the one or more second characteristics; and
embedding, via the one or more processors, the second synthetic data to incorporate the embedded second synthetic data into the embedding space for tuning the classifier [“LLM may then generate a conversational (e.g., natural language) query in response to the prompt that can be included in the generated dataset” §19; Figs. 1, 4B].
It would have been obvious to one having ordinary skill in the art, having the teachings of Mannino, Nagaraju, and Tang before him before the effective filing date of the claimed invention, to modify the combination to incorporate the LLM for synthetic data generation of Nagaraju.
Given the advantage of efficiency and scalability, one having ordinary skill in the art would have been motivated to make this obvious modification.
However, Mannino fails to explicitly disclose embedding, via the one or more processors, the second synthetic data to incorporate the embedded second synthetic data into the embedding space for tuning the classifier.
Tang discloses embedding, via the one or more processors, the second synthetic data to incorporate the embedded second synthetic data into the embedding space [“to obtain embeddings for both the original and synthetic data” §7 ¶1; Fig. 4] for tuning the classifier.
It would have been obvious to one having ordinary skill in the art, having the teachings of Mannino, Nagaraju, and Tang before him before the effective filing date of the claimed invention, to modify the combination to incorporate the explicit embedding of Tang.
Given the advantage of changing raw data into interpretable data in machine learning, one having ordinary skill in the art would have been motivated to make this obvious modification.
Claims 11-15 are rejected on the same grounds as claims 1-5 respectively.
Claims 18-20 are rejected on the same grounds as claims 9-10 and 1 respectively.
Claim(s) 6-8, 16-17 is/are rejected under 35 U.S.C. 103 as being unpatentable over Mannino, Nagaraju, and Tang, further in view of Hosseinzadeh et al. (hereinafter Hosseinzadeh), Logistic regression projection-based feature representation for visual domain adaptation.
Regarding Claim 6, Mannino, Nagaraju, and Tang the method of claim 1.
However, Mannino fails to explicitly disclose tuning, via the one or more processors, the classifier based upon the second embedding space.
Tang discloses tuning, via the one or more processors, the classifier [“fine-tuning a local model” Abstract] based upon the second embedding space.
It would have been obvious to one having ordinary skill in the art, having the teachings of Mannino, Nagaraju, and Tang before him before the effective filing date of the claimed invention, to modify the combination to incorporate tuning of Tang.
Given the advantage of tuning a model to increase its accuracy, one having ordinary skill in the art would have been motivated to make this obvious modification.
However, Mannino fails to explicitly disclose further comprising:
mapping, via the one or more processors, the embedding space into a second embedding space generated for a corpus of documents; and
tuning, via the one or more processors, the classifier based upon the second embedding space.
Hosseinzadeh discloses further comprising:
mapping, via the one or more processors, the embedding space into a second embedding space generated [“feature-based domain adaptation methods have aimed at linking the source and target data distributions by feature space transformation.” §1 ¶5; “learn projection functions in order to project the feature spaces of source and target domains to another space which has common characteristics with the source and target domains” §1 ¶6] for a corpus of documents [“unlabeled data from the target domain” §1 ¶2]; and
tuning, via the one or more processors, the classifier based upon the second embedding space [“feature-based domain adaptation methods have aimed at linking the source and target data distributions by feature space transformation.” §1 ¶5; “learn projection functions in order to project the feature spaces of source and target domains to another space which has common characteristics with the source and target domains” §1 ¶6].
It would have been obvious to one having ordinary skill in the art, having the teachings of Mannino, Nagaraju, Tang, and Hosseinzadeh before him before the effective filing date of the claimed invention, to modify the combination ensure the embedding space between the source data and the target data are similar of Hosseinzadeh.
Given the advantage of ensuring the synthetic data is similar to the actual data, one having ordinary skill in the art would have been motivated to make this obvious modification.
Regarding Claim 7, Mannino, Nagaraju, Tang, and Hosseinzadeh disclose the method of claim 6. Mannino further discloses wherein the corpus of documents includes privileged and/or confidential information [“such as private medical or financial records” §1 ¶1].
Regarding Claim 8, Mannino, Nagaraju, Tang, and Hosseinzadeh disclose the method of claim 6.
However, Mannino fails to explicitly disclose wherein mapping the embedding space into the second embedding space comprises:
applying, via the one or more processors, a logistic regression.
Hosseinzadeh discloses wherein mapping the embedding space into the second embedding space comprises:
applying, via the one or more processors, a logistic regression [“a novel method for unsupervised domain adaptation, called logistic regression projection-based feature representation. The proposed method performs the semi-supervised learning method on both the source and the target domains to predict the pseudo-label values for unlabeled target data. We incorporate the predicted target data with the source training dataset in order to learn feature representation which can be compensated for the distribution mismatch between source and target data.” Abstract].
It would have been obvious to one having ordinary skill in the art, having the teachings of Mannino, Nagaraju, Tang, and Hosseinzadeh before him before the effective filing date of the claimed invention, to modify the combination using logistic regression to ensure the embedding space between the source data and the target data are similar of Hosseinzadeh.
Given the advantage of ensuring the synthetic data is similar to the actual data, one having ordinary skill in the art would have been motivated to make this obvious modification.
Claims 16-17 are rejected on the same grounds as claims 6 and 8 respectively.
Examiner’s Note
The Examiner respectfully requests of the Applicant in preparing responses, to fully consider the entirety of the reference(s) as potentially teaching all or part of the claimed invention. It is noted, REFERENCES ARE RELEVANT AS PRIOR ART FOR ALL THEY CONTAIN. “The use of patents as references is not limited to what the patentees describe as their own inventions or to the problems with which they are concerned. They are part of the literature of the art, relevant for all they contain.” In re Heck, 699 F.2d 1331, 1332-33, 216 USPQ 1038, 1039 (Fed. Cir. 1983) (quoting In re Lemelson, 397 F.2d 1006, 1009, 158 USPQ 275, 277 (CCPA 1968)). A reference may be relied upon for all that it would have reasonably suggested to one having ordinary skill in the art, including non-preferred embodiments (see MPEP 2123). The Examiner has cited particular locations in the reference(s) as applied to the claim(s) above for the convenience of the Applicant. Although the specified citations are representative of the teachings of the art and are applied to the specific limitations within the individual claim(s), typically other passages and figures will apply as well.
Additionally, any claim amendments for any reason should include remarks indicating clear support in the originally filed specification.
Conclusion
Any prior art made of record and not relied upon is considered pertinent to Applicant's disclosure. Applicant is reminded that in amending in response to a rejection of claims, the patentable novelty must be clearly shown in view of the state of the art disclosed by the references cited and the objections made. Applicant must also show how the amendments avoid such references and objections. See 37 CFR §1.111(c). Additionally when amending, in their remarks Applicant should particularly cite to the supporting paragraphs in the original disclosure for the amendments.
The following references were found during the examination of this patent application and were found to be relevant to patentability. Applicant is advised to review these references prior to responding to this Office action.
Lu et al. (Machine Learning for Synthetic Data Generation: A Review) discloses various machine learning techniques for synthetic data generation.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ROBERT H BEJCEK II whose telephone number is (571)270-3610. The examiner can normally be reached Monday - Friday: 9:00am - 5:00pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Michelle T. Bechtold can be reached at (571) 431-0762. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/R.B./ Examiner, Art Unit 2148
/MICHELLE T BECHTOLD/ Supervisory Patent Examiner, Art Unit 2148