DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 1-20 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claim 1 recites the limitations "generating a second weak training label for the content item", “determining a probabilistic training label for the content item”, and “training a lightweight generative machine-learning model using the content item” on lines 7, 11, and 13. There is insufficient antecedent basis for these limitations in the claim. It is unclear whether “the content item” is referring to specifically the “content item not having a training label” or another content item. For the purposes of examination, the Examiner has interpreted these instances and all subsequent instances in the dependent claims as a content item without a training label.
Claims 2-7 are rejected for at least the same reasons as Claim 1 since they depend on Claim 1.
Claim 8 recites the limitations "generating a set of weak training labels for the content item", “determining the training label for the content item”, and “training a lightweight generative machine-learning model using the content item” on lines 4, 7, and 9. There is insufficient antecedent basis for this limitation in the claim. It is unclear whether “the content item” is referring to specifically the “content item not having a training label” or another content item. For the purposes of examination, the Examiner has interpreted these instances and all subsequent instances in the dependent claims as a content item without a training label.
Claims 9-15 are rejected for at least the same reasons as claim 8 since they depend on claim 8.
Claim 18 recites the limitations "generating a set of weak training labels for the content item", “determining the training label for the content item”, and “training a lightweight generative machine-learning model using the content item” on lines 8, 10, and 12. There is insufficient antecedent basis for this limitation in the claim. It is unclear whether “the content item” is referring to specifically the “content item not having a training label” or another content item. For the purposes of examination, the Examiner has interpreted these instances and all subsequent instances in the dependent claims as a content item without a training label.
Claims 19-20 are rejected for at least the same reasons as claim 18 since they depend on claim 18.
Claim 10 recites the limitation "the first LGM directly generates the first weak training label for the content item based on a first input prompt that includes context for the dataset of subjective content items” on lines 1-3. There is insufficient antecedent basis for this limitation in the claim. Is it unclear whether “the dataset of subjective content items” is referring to specifically the “subjective data” in claim 8. For the purposes of examination, the Examiner has interpreted these instances and all subsequent instances in the dependent claims as the subjective data defined in claim 8.
Claim 13 is rejected for at least the same reasons as claim 10 since it ultimately depends on claim 10.
Claim 15 recites the limitation " weak training labels include noisy, inaccurate, or incomplete annotations of the subjective content items” on lines 4-5. There is insufficient antecedent basis for this limitation in the claim. It is unclear whether “the subjective content items” is referring to the “subjective data” in claim 8. For the purposes of examination, the Examiner has interpreted these instances as the subjective data defined in claim 8.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claim(s) 1-20 are rejected under 35 U.S.C. 101 because the claims are directed towards an abstract idea without significantly more.
Regarding Claim 1:
Subject Matter Eligibility Analysis Step 1:
Claim 1 recites a method and is thus a process, one of the four statutory categories of patentable subject matter.
Subject Matter Eligibility Analysis Step 2A Prong 1:
Claim 1 recites:
Generating a first weak training label for the content item (This limitation is a mental process as it encompasses a human mentally generating a label for an item and is thus an evaluation.)
Generating a second weak training label for the content item based on a second LGM that uses a second input prompt format (This limitation is a mental process as it encompasses a human mentally generating a label for an item and is thus an evaluation.)
wherein the first LGM generates a first output format that differs from a second output format generated by the second LGM (This limitation is a mental process as it encompasses a human mentally generating an output format and is thus an evaluation.)
Determining a probabilistic training label for the content item by utilizing a label ensembler function with the first weak training label and the second weak training label (This limitation is a mental process as it encompasses a human mentally determining a label by combining two labels and is thus an evaluation.)
Therefore, Claim 1 recites an abstract idea.
Subject Matter Eligibility Analysis Step 2A Prong 2:
Claim 1 further recites additional elements of:
A computer-implemented method for generating accurate training data sets from subjective data (This element does not integrate the abstract idea into a practical application because it amounts to mere "apply it on a computer" (see MPEP 2106.05(f)).)
obtaining a dataset of subjective content items including a content item not having a training label (This element does not integrate the abstract idea into a practical application because it recites the insignificant extra-solution activity of obtaining a dataset (see MPEP 2106.05(g)).)
utilizing a first large generative machine-learning model (a first LGM) that uses a first input prompt format (This element does not integrate the abstract idea into a practical application because it amounts to mere “apply it on a computer” (see MPEP 2106.05(f)).)
and wherein the first LGM operates separately from the second LGM (This element does not integrate the abstract idea into a practical application because it amounts to mere “apply it on a computer” (see MPEP 2106.05(f)).
And training a lightweight generative machine-learning model using the content item and the probabilistic training label to classify non-training content items in real time (This element does not integrate the abstract idea into a practical application because it recites insignificant extra-solution activity of iteratively training a model (see MPEP 2106.05(g)).)
Therefore, claim 1 is not integrated into a practical application.
Subject Matter Eligibility Analysis Step 2B:
The additional elements of Claim 1 do not provide significantly more than the abstract idea itself, taken alone and in combination because:
A computer-implemented method for generating accurate training data sets from subjective data uses a computer as a tool to perform the abstract idea and cannot provide significantly more (see MPEP 2106.05(f)).
obtaining a dataset of subjective content items including a content item not having a training label is the well understood, routine, and conventional activity of "transmitting or receiving data over a network" (see MPEP 2106.05(d)(II); OIP Techs., Inc., V. Amazon.com, Inc., 788 F.3d 1359, 1363, 115 USPQ2d 1090, 1093 (Fed. Cir. 2015) (sending messages over a network)).
utilizing a first large generative machine-learning model (a first LGM) that uses a first input prompt format uses a computer as a tool to perform the abstract idea and cannot provide significantly more (see MPEP 2106.05(f)).
and wherein the first LGM operates separately from the second LGM uses a computer as a tool to perform the abstract idea and cannot provide significantly more (see MPEP 2106.05(f)).
And training a lightweight generative machine-learning model using the content item and the probabilistic training label to classify non-training content items in real time is the well understood, routine, and conventional activity of training a machine learning model (US 2021/0125108 A1, Metzler et al. page 11, paragraph 0052, “a classifier is trained using a conventional iterative machine learning training process that determines weights for each result list position”).
Therefore, claim 1 is subject-matter ineligible.
Regarding Claim 2:
Subject Matter Eligibility Analysis Step 2A Prong 1:
Claim 2 recites:
wherein the first LGM directly generates the first weak training label for the content item of the first output format (This limitation is a mental process as it encompasses a human mentally generating a weak training label and is thus an evaluation.)
Therefore, claim 2 recites an abstract idea.
Subject Matter Eligibility Analysis Step 2A Prong 2:
Claim 2 does not further recite any additional elements. Therefore, claim 2 is not
integrated into a practical application.
Subject Matter Eligibility Analysis Step 2B:
Since there are no additional elements, claim 2 does not provide significantly more than
the abstract idea itself, taken alone and in combination. Therefore, claim 2 is subject-matter
ineligible.
Regarding Claim 3:
Subject Matter Eligibility Analysis Step 2A Prong 1:
Claim 3 recites:
wherein the second LGM is used to generate the second weak training label for the content item by (This limitation is a mental process as it encompasses a human mentally generating a label and is thus an evaluation.)
generating the second output format that includes a set of synthetic training data and corresponding synthetic training labels based on the second input prompt format to the second LGM (This limitation is a mental process as it encompasses a human mentally generating the second output format and is thus an evaluation.)
generating the second weak training label for the content item (This limitation is a mental process as it encompasses a human mentally generating the second weak training label and is thus an evaluation.)
Therefore, claim 3 recites an abstract idea.
Subject Matter Eligibility Analysis Step 2A Prong 2:
Claim 3 further recites additional elements of:
training a lightweight classifier model based on the set of synthetic training data and the corresponding synthetic training labels (This element does not integrate the abstract idea into a practical application because it recites insignificant extra-solution activity of training a classifier model (see MPEP 2106.05(g)).)
utilizing the lightweight classifier model (This element does not integrate the abstract idea into a practical application because it amounts to mere “apply it on a computer” (see MPEP 2106.05(f)).)
Therefore, claim 3 is not integrated into a practical application.
Subject Matter Eligibility Analysis Step 2B:
The additional elements of claim 3 do not provide significantly more than the abstract
idea itself, taken alone and in combination because
Training a lightweight generative machine-learning model is the well understood, routine, and conventional activity of iteratively training a model (US 2021/0125108 A1, Metzler et al. page 11, paragraph 0052, “a classifier is trained using a conventional iterative machine learning training process that determines weights for each result list position”).
utilizing the lightweight classifier model uses a computer as a tool to perform the abstract idea and cannot provide significantly more (see MPEP 2106.05(f)).
Therefore, claim 3 is subject-matter ineligible.
Regarding Claim 4:
Subject Matter Eligibility Analysis Step 2A Prong 1:
Claim 4 recites the same abstract idea as claim 1. Therefore, claim 4 recites an abstract idea.
Subject Matter Eligibility Analysis Step 2A Prong 2:
Claim 4 further recites additional elements of:
wherein the lightweight generative machine-learning model is an order of magnitude smaller than the first LGM or the second LGM while achieving comparable output accuracy. (This element does not integrate the abstract idea into a practical application because it amounts to mere “apply it on a computer” (see MPEP 2106.05(f)).)
Therefore, claim 4 is not integrated into a practical application.
Subject Matter Eligibility Analysis Step 2B:
The additional elements of claim 4 do not provide significantly more than the abstract
idea itself, taken alone and in combination because
wherein the lightweight generative machine-learning model is an order of magnitude smaller than the first LGM or the second LGM while achieving comparable output accuracy uses a computer as a tool to perform the abstract idea and cannot provide significantly more (see MPEP 2106.05(f)).)
Therefore, claim 4 is subject-matter ineligible.
Regarding Claim 5:
Subject Matter Eligibility Analysis Step 2A Prong 1:
Claim 5 recites:
determining the probabilistic training label is further based on additional weak training labels generated for the content item by additional labeling functions processing the content item (This limitation is a mental process as it encompasses a human mentally determining a label and is thus an evaluation.)
Therefore, claim 5 recites an abstract idea.
Subject Matter Eligibility Analysis Step 2A Prong 2:
Claim 5 does not further recite any additional elements. Therefore, claim 5 is not
integrated into a practical application.
Subject Matter Eligibility Analysis Step 2B:
Since there are no additional elements, claim 5 does not provide significantly more than
the abstract idea itself, taken alone and in combination. Therefore, claim 5 is subject-matter
ineligible.
Regarding Claim 6:
Subject Matter Eligibility Analysis Step 2A Prong 1:
Claim 6 recites:
generating multiple separate instances of the first weak training label for the content item (This limitation is a mental process as it encompasses a human mentally generating multiple instances of the first weak training label and is thus an evaluation.)
Therefore, claim 6 recites an abstract idea.
Subject Matter Eligibility Analysis Step 2A Prong 2:
Claim 6 further recites:
by providing different instances of the first input prompt format to the first LGM (This element does not integrate the abstract idea into a practical application because it amounts to mere “apply it on a computer” (see MPEP 2106.05(f)).)
wherein the different instances of the first input prompt format correspond to different perspectives of the dataset of the subjective content items (This element does not integrate the abstract idea into a practical application because it amounts to mere “apply it on a computer” (see MPEP 2106.05(f)).)
Therefore, claim 6 is not integrated into a practical application.
Subject Matter Eligibility Analysis Step 2B:
The additional elements of claim 6 do not provide significantly more than the abstract
idea itself, taken alone and in combination because
by providing different instances of the first input prompt format to the first LGM uses a computer as a tool to perform the abstract idea and cannot provide significantly more (see MPEP 2106.05(f)).
wherein the different instances of the first input prompt format correspond to different perspectives of the dataset of the subjective content items uses a computer as a tool to perform the abstract idea and cannot provide significantly more (see MPEP 2106.05(f)).)
Therefore, claim 6 is subject-matter ineligible.
Regarding Claim 7:
Subject Matter Eligibility Analysis Step 2A Prong 1:
Claim 7 recites the same abstract idea as claim 1. Therefore, claim 7 recites an abstract idea.
Subject Matter Eligibility Analysis Step 2A Prong 2:
Claim 7 further recites additional elements of:
wherein the first LGM and the second LGM are different instances of a same LGM (This element does not integrate the abstract idea into a practical application because it amounts to mere “apply it on a computer” (see MPEP 2106.05(f)).)
Therefore, claim 7 is not integrated into a practical application.
Subject Matter Eligibility Analysis Step 2B:
The additional elements of claim 7 do not provide significantly more than the abstract
idea itself, taken alone and in combination because
wherein the first LGM and the second LGM are different instances of a same LGM uses a computer as a tool to perform the abstract idea and cannot provide significantly more (see MPEP 2106.05(f)).)
Therefore, claim 7 is subject-matter ineligible.
Regarding Claim 8:
Subject Matter Eligibility Analysis Step 1:
Claim 8 recites a method and is thus a process, one of the four statutory categories of patentable subject matter.
Subject Matter Eligibility Analysis Step 2A Prong 1:
Claim 8 recites:
generating a set of weak training labels for the content item (This limitation is a mental process as it encompasses a human mentally generating a label for an item and is thus an evaluation.)
determining the training label for the content item from the set of weak training labels using a label ensembler function (This limitation is a mental process as it encompasses a human mentally determining a label by combining two labels and is thus an evaluation.)
Therefore, Claim 8 recites an abstract idea.
Subject Matter Eligibility Analysis Step 2A Prong 2:
Claim 8 further recites additional elements of:
A computer-implemented method for generating accurate training data sets from subjective data, comprising (This element does not integrate the abstract idea into a practical application because it amounts to mere "apply it on a computer" (see MPEP 2106.05(f)).)
obtaining a dataset of content items including a content item not having a training label (This element does not integrate the abstract idea into a practical application because it recites the insignificant extra-solution activity of obtaining a dataset (see MPEP 2106.05(g)).)
utilizing different large generative machine-learning models (different LGMs), wherein the different LGMs produce different output formats (This element does not integrate the abstract idea into a practical application because it amounts to mere “apply it on a computer” (see MPEP 2106.05(f)).)
and training a lightweight generative machine-learning model using the content item and the training label (This element does not integrate the abstract idea into a practical application because it recites the insignificant extra-solution activity of training a model (see MPEP 2106.05(g)).)
Therefore, claim 8 is not integrated into a practical application.
Subject Matter Eligibility Analysis Step 2B:
The additional elements of Claim 8 do not provide significantly more than the abstract
idea itself, taken alone and in combination because:
A computer-implemented method for generating accurate training data sets from subjective data, comprising uses a computer as a tool to perform the abstract idea and cannot provide significantly more (see MPEP 2106.05(f)).
obtaining a dataset of content items including a content item not having a training label is the well understood, routine, and conventional activity of "transmitting or receiving data over a network" (see MPEP 2106.05(d)(II); OIP Techs., Inc., V. Amazon.com, Inc., 788 F.3d 1359, 1363, 115 USPQ2d 1090, 1093 (Fed. Cir. 2015) (sending messages over a network)).
utilizing different large generative machine-learning models (different LGMs), wherein the different LGMs produce different output formats uses a computer as a tool to perform the abstract idea and cannot provide significantly more (see MPEP 2106.05(f)).
and training a lightweight generative machine-learning model using the content item and the training label is the well understood, routine, and conventional activity of iteratively training a model (US 2021/0125108 A1, Metzler et al. page 11, paragraph 0052, “a classifier is trained using a conventional iterative machine learning training process that determines weights for each result list position”).
Therefore, claim 8 is subject-matter ineligible.
Regarding Claim 9:
Subject Matter Eligibility Analysis Step 2A Prong 1:
Claim 9 recites:
wherein the different LGMs include a first LGM that directly generates a first weak training label for the content item (This limitation is a mental process as it encompasses a human mentally generating a label and is thus an evaluation.)
Therefore, claim 9 recites an abstract idea.
Subject Matter Eligibility Analysis Step 2A Prong 2:
Claim 9 does not further recite any additional elements. Therefore, claim 9 is not
integrated into a practical application.
Subject Matter Eligibility Analysis Step 2B:
Since there are no additional elements, claim 9 does not provide significantly more than
the abstract idea itself, taken alone and in combination. Therefore, claim 9 is subject-matter
ineligible.
Regarding Claim 10:
Subject Matter Eligibility Analysis Step 2A Prong 1:
Claim 10 recites:
wherein the first LGM directly generates the first weak training label for the content item based on a first input prompt that includes context for the dataset of subjective content items and rules for directly generating weak training labels (This limitation is a mental process as it encompasses a human mentally generating a training label and is thus an evaluation.)
Therefore, claim 10 recites an abstract idea.
Subject Matter Eligibility Analysis Step 2A Prong 2:
Claim 10 does not further recite any additional elements. Therefore, claim 10 is not
integrated into a practical application.
Subject Matter Eligibility Analysis Step 2B:
Since there are no additional elements, claim 10 does not provide significantly more than
the abstract idea itself, taken alone and in combination. Therefore, claim 10 is subject-matter
ineligible.
Regarding Claim 11:
Subject Matter Eligibility Analysis Step 2A Prong 1:
Claim 11 recites:
wherein the different LGMs include a second LGM that indirectly generates a second weak training label for the content item (This limitation is a mental process as it encompasses a human mentally generating a label and is thus an evaluation.)
Therefore, claim 11 recites an abstract idea.
Subject Matter Eligibility Analysis Step 2A Prong 2:
Claim 11 does not further recite any additional elements. Therefore, claim 11 is not
integrated into a practical application.
Subject Matter Eligibility Analysis Step 2B:
Since there are no additional elements, claim 11 does not provide significantly more than
the abstract idea itself, taken alone and in combination. Therefore, claim 11 is subject-matter
ineligible.
Regarding Claim 12:
Subject Matter Eligibility Analysis Step 2A Prong 1:
Claim 12 recites:
wherein the second LGM generates the second weak training label for the content item by (This limitation is a mental process as it encompasses a human mentally generating a label and is thus an evaluation.)
generating a set of synthetic training data and corresponding synthetic training labels (This limitation is a mental process as it encompasses a human mentally generating data and labels and is thus an evaluation.)
generating the second weak training label for the content item (This limitation is a mental process as it encompasses a human mentally generating a label and is thus an evaluation.)
Therefore, claim 12 recites an abstract idea.
Subject Matter Eligibility Analysis Step 2A Prong 2:
Claim 12 further recites additional elements of:
training a lightweight classifier model based on the set of synthetic training data and the corresponding synthetic training labels (This element does not integrate the abstract idea into a practical application because it recites the insignificant extra-solution activity of training a model (see MPEP 2106.05(g)).)
utilizing the lightweight classifier model (This element does not integrate the abstract idea into a practical application because it amounts to mere “apply it on a computer” (see MPEP 2106.05(f)).)
Therefore, claim 12 is not integrated into a practical application.
Subject Matter Eligibility Analysis Step 2B:
The additional elements of claim 12 do not provide significantly more than the abstract
idea itself, taken alone and in combination because
training a lightweight classifier model based on the set of synthetic training data and the corresponding synthetic training labels is the well understood, routine, and conventional activity of iteratively training a model (US 2021/0125108 A1, Metzler et al. page 11, paragraph 0052, “a classifier is trained using a conventional iterative machine learning training process that determines weights for each result list position”).
utilizing the lightweight classifier model uses a computer as a tool to perform the abstract idea and cannot provide significantly more (see MPEP 2106.05(f)).
Therefore, claim 12 is subject-matter ineligible.
Regarding Claim 13:
Subject Matter Eligibility Analysis Step 2A Prong 1:
Claim 13 recites:
wherein the second LGM generates the set of synthetic training data and the corresponding synthetic training labels based on a second input prompt that provides context for the dataset of subjective content items and rules for generating the set of synthetic training data and the corresponding synthetic training labels (This limitation is a mental process as it encompasses a human mentally generating data and labels and is thus an evaluation.)
Therefore, claim 13 recites an abstract idea.
Subject Matter Eligibility Analysis Step 2A Prong 2:
Claim 13 does not further recite any additional elements. Therefore, claim 13 is not
integrated into a practical application.
Subject Matter Eligibility Analysis Step 2B:
Since there are no additional elements, claim 13 does not provide significantly more than
the abstract idea itself, taken alone and in combination. Therefore, claim 13 is subject-matter
ineligible.
Regarding Claim 14:
Subject Matter Eligibility Analysis Step 2A Prong 1:
Claim 14 recites the same abstract idea as claim 12. Therefore, claim 14 recites an abstract idea.
Subject Matter Eligibility Analysis Step 2A Prong 2:
Claim 14 further recites additional elements of:
utilizing the lightweight generative machine-learning model to process non-training content items in real time (This element does not integrate the abstract idea into a practical application because it amounts to mere “apply it on a computer” (see MPEP 2106.05(f)).)
Therefore, claim 14 is not integrated into a practical application.
Subject Matter Eligibility Analysis Step 2B:
The additional elements of claim 14 do not provide significantly more than the abstract
idea itself, taken alone and in combination because
utilizing the lightweight generative machine-learning model to process non-training content items in real time uses a computer as a tool to perform the abstract idea and cannot provide significantly more (see MPEP 2106.05(f)).
Therefore, claim 14 is subject-matter ineligible.
Regarding Claim 15:
Subject Matter Eligibility Analysis Step 2A Prong 1:
Claim 15 recites:
weak training labels include noisy, inaccurate, or incomplete annotations of the subjective content items (This further modifies the limitation “generating a set of weak training labels for the content item” from Claim 8 and is thus a mental process as it encompasses a human mentally generating a label.)
the label ensembler function generates a probabilistic training label for the content item from the set of weak training labels generated by the different LGMs (This limitation is a mental process as it encompasses a human mentally generating a label and is thus an evaluation.)
Therefore, claim 15 recites an abstract idea.
Subject Matter Eligibility Analysis Step 2A Prong 2:
Claim 15 further recites additional elements of:
subjective content items are difficult to manually classify without an expert in a field particular to a context of the content item (This further modifies the limitation “obtaining a dataset of content items” from Claim 8 and thus recites the insignificant extra-solution activity of obtaining a dataset (see MPEP 2106.05(g)).)
Therefore, claim 15 is not integrated into a practical application.
Subject Matter Eligibility Analysis Step 2B:
The additional elements of claim 15 do not provide significantly more than the abstract
idea itself, taken alone and in combination because
subjective content items are difficult to manually classify without an expert in a field particular to a context of the content item further modifies the limitation “obtaining a dataset of content items” from Claim 8 and thus is the well understood, routine, and conventional activity of "transmitting or receiving data over a network" (see MPEP 2106.05(d)(II); OIP Techs., Inc., V. Amazon.com, Inc., 788 F.3d 1359, 1363, 115 USPQ2d 1090, 1093 (Fed. Cir. 2015) (sending messages over a network)).
Therefore, claim 15 is subject-matter ineligible.
Regarding Claim 16:
Subject Matter Eligibility Analysis Step 2A Prong 1:
Claim 16 recites the same abstract idea as claim 8. Therefore, claim 16 recites an abstract idea.
Subject Matter Eligibility Analysis Step 2A Prong 2:
Claim 16 further recites additional elements of:
wherein a LGM of the different LGMs is a large generic multi-modal generative model (This element does not integrate the abstract idea into a practical application because it recites generic computing components on which to perform the abstract idea (see MPEP 2106.05(f)).)
Therefore, claim 16 is not integrated into a practical application.
Subject Matter Eligibility Analysis Step 2B:
The additional elements of claim 16 do not provide significantly more than the abstract
idea itself, taken alone and in combination because
wherein a LGM of the different LGMs is a large generic multi-modal generative model uses a computer as a tool to perform the abstract idea and cannot provide significantly more (see MPEP 2106.05(f)).
Therefore, claim 16 is subject-matter ineligible.
Regarding Claim 17:
Subject Matter Eligibility Analysis Step 2A Prong 1:
Claim 17 recites the same abstract idea as claim 8. Therefore, claim 17 recites an abstract idea.
Subject Matter Eligibility Analysis Step 2A Prong 2:
Claim 17 further recites additional elements of:
wherein the different LGMs are different instances of a same LGM provided with different input prompt formats (This element does not integrate the abstract idea into a practical application because it recites generic computing components on which to perform the abstract idea (see MPEP 2106.05(f)).)
Therefore, claim 17 is not integrated into a practical application.
Subject Matter Eligibility Analysis Step 2B:
The additional elements of claim 17 do not provide significantly more than the abstract
idea itself, taken alone and in combination because
wherein the different LGMs are different instances of a same LGM provided with different input prompt formats uses a computer as a tool to perform the abstract idea and cannot provide significantly more (see MPEP 2106.05(f)).
Therefore, claim 17 is subject-matter ineligible.
Regarding Claim 18:
Subject Matter Eligibility Analysis Step 1:
Claim 18 recites a system and is thus a machine, one of the four statutory categories of patentable subject matter.
Subject Matter Eligibility Analysis Step 2A Prong 1:
Claim 18 recites:
A system for generating accurate training data sets from subjective data, comprising: a dataset of content items including a content item not having a training label (This limitation is a mental process as it encompasses a human mentally generating a training data set and is thus an evaluation.)
generating a set of weak training labels for the content item (This limitation is a mental process as it encompasses a human mentally generating a set of labels and is thus an evaluation.)
determining the training label for the content item from the set of weak training labels using a label ensembler function (This limitation is a mental process as it encompasses a human mentally determining a label and is thus an evaluation.)
Therefore, claim 18 recites an abstract idea.
Subject Matter Eligibility Analysis Step 2A Prong 2:
Claim 18 further recites additional elements of:
a set of different large generative machine-learning models (different LGMs) including a first LGM and a second LGM (This element does not integrate the abstract idea into a practical application because it recites generic computing components on which to perform the abstract idea (see MPEP 2106.05(f)).)
a processing system comprising a processor; and a computer memory comprising instructions that, when executed by the processing system, cause the system to perform operations (This element does not integrate the abstract idea into a practical application because it recites generic computing components on which to perform the abstract idea (see MPEP 2106.05(f)).)
utilizing the different LGMs, wherein the different LGMs produce different output formats (This element does not integrate the abstract idea into a practical application because it amounts to mere “apply it on a computer” (see MPEP 2106.05(f)).)
training a lightweight generative machine-learning model using the content item and the training label (This element does not integrate the abstract idea into a practical application because it recites the insignificant extra-solution activity of training a model (see MPEP 2106.05(g)).)
Therefore, claim 18 is not integrated into a practical application.
Subject Matter Eligibility Analysis Step 2B:
The additional elements of claim 18 do not provide significantly more than the abstract
idea itself, taken alone and in combination because
a set of different large generative machine-learning models (different LGMs) including a first LGM and a second LGM uses a computer as a tool to perform the abstract idea and cannot provide significantly more (see MPEP 2106.05(f)).
a processing system comprising a processor; and a computer memory comprising instructions that, when executed by the processing system, cause the system to perform operations uses a computer as a tool to perform the abstract idea and cannot provide significantly more (see MPEP 2106.05(f)).
utilizing the different LGMs, wherein the different LGMs produce different output formats uses a computer as a tool to perform the abstract idea and cannot provide significantly more (see MPEP 2106.05(f)).
training a lightweight generative machine-learning model using the content item and the training label is the well understood, routine, and conventional activity of iteratively training a model (US 2021/0125108 A1, Metzler et al. page 11, paragraph 0052, “a classifier is trained using a conventional iterative machine learning training process that determines weights for each result list position”).
Therefore, claim 18 is subject-matter ineligible.
Regarding Claim 19:
Subject Matter Eligibility Analysis Step 2A Prong 1:
Claim 19 recites:
the first LGM directly generates a first weak training label for the content item (This limitation is a mental process as it encompasses a human mentally generating a label and is thus an evaluation.)
and the second LGM indirectly generates a second weak training label for the content item (This limitation is a mental process as it encompasses a human mentally generating a label and is thus an evaluation.)
Therefore, claim 19 recites an abstract idea.
Subject Matter Eligibility Analysis Step 2A Prong 2:
Claim 19 does not further recite any additional elements. Therefore, claim 19 is not
integrated into a practical application.
Subject Matter Eligibility Analysis Step 2B:
Since there are no additional elements, claim 19 does not provide significantly more than
the abstract idea itself, taken alone and in combination. Therefore, claim 19 is subject-matter
ineligible.
Regarding Claim 20:
Subject Matter Eligibility Analysis Step 2A Prong 1:
Claim 20 recites:
wherein the second LGM generates the second weak training label for the content item by (This limitation is a mental process as it encompasses a human mentally generating a label and is thus an evaluation.)
generating a set of synthetic training data and corresponding synthetic training labels offline (This limitation is a mental process as it encompasses a human mentally generating data and labels and is thus an evaluation.)
and generating the second weak training label for the content item (This limitation is a mental process as it encompasses a human mentally generating a label and is thus an evaluation.)
Therefore, claim 20 recites an abstract idea.
Subject Matter Eligibility Analysis Step 2A Prong 2:
Claim 20 further recites additional elements of:
training a lightweight classifier model based on the set of synthetic training data and the corresponding synthetic training labels (This element does not integrate the abstract idea into a practical application because it recites the insignificant extra-solution activity of training a model (see MPEP 2106.05(g)).)
utilizing the lightweight classifier model (This element does not integrate the abstract idea into a practical application because it amounts to mere “apply it on a computer” (see MPEP 2106.05(f)).)
Therefore, claim 20 is not integrated into a practical application.
Subject Matter Eligibility Analysis Step 2B:
The additional elements of claim 20 do not provide significantly more than the abstract
idea itself, taken alone and in combination because
training a lightweight classifier model based on the set of synthetic training data and the corresponding synthetic training labels is the well understood, routine, and conventional activity of iteratively training a model (US 2021/0125108 A1, Metzler et al. page 11, paragraph 0052, “a classifier is trained using a conventional iterative machine learning training process that determines weights for each result list position”).
utilizing the lightweight classifier model uses a computer as a tool to perform the abstract idea and cannot provide significantly more (see MPEP 2106.05(f)).
Therefore, claim 20 is subject-matter ineligible.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claim(s) 1-2, 4-5, 7-11, 14-17 is/are rejected under 35 U.S.C. 103 as being unpatentable over Smith et al. (“Language Models in the Loop: Incorporating Prompting into Weak Supervision”) in view of Gu et al. (“Multi-train: A semi-supervised heterogeneous ensemble classifier”)
Regarding Claim 1, Smith teaches
A computer-implemented method for generating accurate training data sets from subjective data (Smith, Page 1, “we investigate how to direct the knowledge contained in pre-trained language models toward the creation of labeled training data”);
obtaining a dataset of subjective content items including a content item not having a training label (Smith, Page 2, “Large pre-trained language models are an untapped resource as a potentially complementary source of heuristic labels. In addition to the ease of specifying heuristics with natural language, we show that they can effectively capture a wide range of fuzzy concepts that can be hard to express as traditional labeling functions written in code.” Smith, Page 9, “For each dataset in our analysis, we assume the original training split is unlabeled. All labelers, here prompted labeling functions and code-based labeling functions, are applied to the unlabeled training split to generate votes for the true label of each example.” Examiner notes that the aforementioned fuzzy concepts are defined as subjective content items.);
generating a first weak training label for the content item utilizing a first large generative machine-learning model (a first LGM) that uses a first input prompt format (Smith, Page 5, Figure 2. Smith, Page 1, “Rather than apply the model in a typical zero-shot or few-shot fashion, we treat the model as the basis for labeling functions in a weak supervision framework. To create a classifier, we first prompt the model to answer multiple distinct queries about an example and define how the possible responses should be mapped to votes for labels and abstentions.” Smith, Page 9, “All prompts are evaluated using two different language model families: GPT-3 and T0++. We use the InstructGPT [37] family of GPT-3 engines, evaluating Ada, Babbage, and Curie since different engines are claimed to be better suited to specific tasks.” Examiner notes that the types of models used in Smith are Large Language Models, a type of Large Generative Model. Examiner further notes that ‘weak supervision’ is defined as the practice of leveraging limited, noisy, or imprecise labeled (“weak”) data to train models efficiently. Examiner further notes that the first weak training label is Label 1 in Figure 2, and the first input prompt format is Prompt 1 in Figure 2.);
generating a second weak training label for the content item based on a second LGM that uses a second input prompt format (Smith, Page 5, Figure 2. Examiner notes that the second weak training label is Label 2 and the second input prompt format is Prompt 2.);
determining a probabilistic training label for the content item by utilizing a label ensembler function with the first weak training label and the second weak training label (Smith, Page 9, “All labeler votes, unless otherwise noted, are combined and denoised using the FlyingSquid [22] label model to estimate a single, probabilistic consensus label per example.”);
and training a lightweight generative machine-learning model using the content item and the probabilistic training label to classify non-training content items in real time (Smith, Page 9, “The resulting labels are used to train a RoBERTa [32] end model, which provides a smaller, more servable classification model tailored to our task of interest.”)
Smith does not teach “wherein the first LGM generates a first output format that differs from a second output format generated by the second LGM” and “wherein the first LGM operates separately from the second LGM”. However, Gu teaches
wherein the first LGM generates a first output format that differs from a second output format generated by the second LGM (Gu, Page 205, “the proposed algorithm requires the user to pair up the feature manipulation methods and the learning model, each pair representing a base classifier. The corresponding classifier is initially trained with the specified learning algorithm with features manipulated by the pre-specified feature manipulation method. Therefore, a pair of feature manipulation method and a model represents a “view” to the data.” Gu, page 206, “We use three methods to create different views from the same training data, namely, use of different learning models, use differently manipulated features, or a combination of the above.” Examiner notes that utilizing different output formats is the feature manipulation method. Examiner further notes that a view comprises a feature manipulation method and a model, and different views are used meaning different feature manipulation methods are used for different models.)
and wherein the first LGM operates separately from the second LGM (Gu, Page 202, “we propose a novel semi-supervised ensemble learning algorithm, termed Multi-Train, which generates a number of heterogeneous classifiers that use different classification models and/or different features.”)
Smith and Gu are considered analogous to the claimed invention because they involve labeling data using an ensemble of models. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Smith to instead utilize separately operating LGM models and different output formats in order to increase the diversity of the classifier. Doing so is advantageous because the better performance of Multi-Train is attributed to “the multiple views generated using different models as well as different feature manipulation methods in contrast to the original single-view data” (Gu, 210). Gu’s findings show that “Heterogeneous ensembles have proved to be more effective in achieving diversity [23,26,27], which is also empirically confirmed by statistically better results than those of the tri-training algorithm on the 12 datasets used in this work” (Gu, Page 203).
Regarding Claim 2, Smith and Gu teach
The computer-implemented method of claim 1, wherein the first LGM directly generates the first weak training label for the content item of the first output format (Smith, Page 5, Figure 2. Examiner notes that the first weak training label is Label 1.)
Regarding Claim 4, Smith and Gu teach
the computer-implemented method of claim 1, wherein the lightweight generative machine-learning model is an order of magnitude smaller than the first LGM or the second LGM while achieving comparable output accuracy (Smith, Page 9, “The resulting labels are used to train a RoBERTa [32] end model, which provides a smaller, more servable classification model tailored to our task of interest.” Page 20, “The RoBERTa end model provides consistent improvements.”)
Regarding Claim 5, Smith and Gu teach
The computer-implemented method of claim 1, wherein determining the probabilistic training label is further based on additional weak training labels generated for the content item by additional labeling functions processing the content item. (Smith, Page 2, “We model prompts as labeling functions by adding additional metadata that maps possible completions to target labels or abstentions. For example, if a task is to classify spam comments, a prompt could be “Does the following comment ask the user to click a link?” If the language model responds positively, then this is an indication that the comment is spam.” Smith, Page 5, Figure 2. Examiner notes that prompts are modeled as labeling functions. Examiner further notes that the additional labeling function is Prompt 3, and the additional weak training label is Label 3).
Regarding Claim 7, Smith does not teach “wherein the first LGM and the second LGM are different instances of a same LGM”. However, Gu teaches
The computer-implemented method of claim 1, wherein the first LGM and the second LGM are different instances of a same LGM (Gu, Page 202, “we propose a novel semi-supervised ensemble learning algorithm, termed Multi-Train, which generates a number of heterogeneous classifiers that use different classification models and/or different features.” Examiner notes that the different classification models can include different instances of a same type of LGM model.)
Smith and Gu are considered analogous to the claimed invention because they involve labeling data using an ensemble of models. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Smith to instead utilize different instances of a same LGM to increase diversity. Doing so is advantageous because the better performance of Multi-Train is attributed to “the multiple views generated using different models as well as different feature manipulation methods in contrast to the original single-view data” (Gu, 210). Gu’s findings show that “Heterogeneous ensembles have proved to be more effective in achieving diversity [23,26,27], which is also empirically confirmed by statistically better results than those of the tri-training algorithm on the 12 datasets used in this work” (Gu, Page 203).
Regarding Claim 8, Smith teaches
A computer-implemented method for generating accurate training data sets from subjective data (Smith, Page 1, “we investigate how to direct the knowledge contained in pre-trained language models toward the creation of labeled training data”);
obtaining a dataset of content items including a content item not having a training label (Smith, Page 9, “For each dataset in our analysis, we assume the original training split is unlabeled. All labelers, here prompted labeling functions and code-based labeling functions, are applied to the unlabeled training split to generate votes for the true label of each example.”);
determining the training label for the content item from the set of weak training labels using a label ensembler function (Smith, Page 9, “All labeler votes, unless otherwise noted, are combined and denoised using the FlyingSquid [22] label model to estimate a single, probabilistic consensus label per example.”);
and training a lightweight generative machine-learning model using the content item and the training label (Smith, Page 9, “The resulting labels are used to train a RoBERTa [32] end model, which provides a smaller, more servable classification model tailored to our task of interest.”)
Smith does not “generating a set of weak training labels for the content item utilizing different large generative machine-learning models (different LGMs)” and “wherein the different LGMs produce different output formats”. However, Gu teaches
generating a set of weak training labels for the content item utilizing different large generative machine-learning models (different LGMs) (Gu, Page 202, “we propose a novel semi-supervised ensemble learning algorithm, termed Multi-Train, which generates a number of heterogeneous classifiers that use different classification models and/or different features.”)
wherein the different LGMs produce different output formats (Gu, Page 205, “the proposed algorithm requires the user to pair up the feature manipulation methods and the learning model, each pair representing a base classifier. The corresponding classifier is initially trained with the specified learning algorithm with features manipulated by the pre-specified feature manipulation method. Therefore, a pair of feature manipulation method and a model represents a “view” to the data.” Gu, page 206, “We use three methods to create different views from the same training data, namely, use of different learning models, use differently manipulated features, or a combination of the above.” Examiner notes that utilizing different output formats is the feature manipulation method. Examiner further notes that a view comprises a feature manipulation method and a model, and different views are used meaning different feature manipulation methods are used for different models.)
Smith and Gu are considered analogous to the claimed invention because they involve labeling data using an ensemble of models. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Smith to instead utilize different LGM models and different output formats in order to increase the diversity of the classifier. Doing so is advantageous because the better performance of Multi-Train is attributed to “the multiple views generated using different models as well as different feature manipulation methods in contrast to the original single-view data” (Gu, 210). Gu’s findings show that “Heterogeneous ensembles have proved to be more effective in achieving diversity [23,26,27], which is also empirically confirmed by statistically better results than those of the tri-training algorithm on the 12 datasets used in this work” (Gu, Page 203).
Claim 9 recites substantially similar limitations to claim 2, and is therefore rejected under the same analysis.
Regarding Claim 10, Smith and Gu teach
The computer-implemented method of claim 9, wherein the first LGM directly generates the first weak training label for the content item based on a first input prompt that includes context for the dataset of subjective content items and rules for directly generating weak training labels (Smith, Page 2, “we propose a framework for incorporating prompting into programmatic weak supervision, in order to address the above challenges and realize potential benefits from pre-trained language models (Figure 1). We model prompts as labeling functions by adding additional metadata that maps possible completions to target labels or abstentions. For example, if a task is to classify spam comments, a prompt could be “Does the following comment ask the user to click a link?” If the language model responds positively, then this is an indication that the comment is spam. On the other hand, if the model responds negatively then that might be mapped to an abstention because both spam and non-spam comments can lack that property.” Smith, Page 5, Figure 2. Examiner notes that the first weak training label is Label 1 and the first input prompt is Prompt 1)
Regarding Claim 11:
Smith does not teach “wherein the different LGMs include a second LGM that indirectly generates a second weak training label for the content item”. However, Gu teaches
The computer-implemented method of claim 9, wherein the different LGMs include a second LGM that indirectly generates a second weak training label for the content item. (Gu, Page 202, “we propose a novel semi-supervised ensemble learning algorithm, termed Multi-Train, which generates a number of heterogeneous classifiers that use different classification models and/or different features.”)
Smith and Gu are considered analogous to the claimed invention because they involve labeling data using an ensemble of models. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Smith to instead utilize a second LGM model in order to increase the diversity of the classifier. Doing so is advantageous because the better performance of Multi-Train is attributed to “the multiple views generated using different models as well as different feature manipulation methods in contrast to the original single-view data” (Gu, 210). Gu’s findings show that “Heterogeneous ensembles have proved to be more effective in achieving diversity [23,26,27], which is also empirically confirmed by statistically better results than those of the tri-training algorithm on the 12 datasets used in this work” (Gu, Page 203).
Regarding Claim 14, Smith and Gu teach
The computer-implemented method of claim 12, further comprising utilizing the lightweight generative machine-learning model to process non-training content items in real time (Smith, Page 9, “The resulting labels are used to train a RoBERTa [32] end model, which provides a smaller, more servable classification model tailored to our task of interest.”)
Regarding Claim 15, Smith and Gu teach
The computer-implemented method of claim 8, wherein: subjective content items are difficult to manually classify without an expert in a field particular to a context of the content item (Smith, Page 2, “In addition to the ease of specifying heuristics with natural language, we show that they can effectively capture a wide range of fuzzy concepts that can be hard to express as traditional labeling functions written in code.”)
weak training labels include noisy, inaccurate, or incomplete annotations of the subjective content items (Smith, Page 3, “Weak supervision refers to a broad family of techniques that attempts to learn from data that is noisily or less precisely labeled than usual. Our focus is on programmatic weak supervision, in which the sources of supervision are heuristic labelers, often called labeling functions that vote on the true labels of unlabeled examples [65].”)
and the label ensembler function generates a probabilistic training label for the content item from the set of weak training labels generated by the different LGMs (Smith, Page 9, “All labeler votes, unless otherwise noted, are combined and denoised using the FlyingSquid [22] label model to estimate a single, probabilistic consensus label per example.”);
Regarding Claim 16, Smith and Gu teach
The computer-implemented method of claim 8, wherein a LGM of the different LGMs is a large generic multi-modal generative model. (Smith, Page 9, “All prompts are evaluated using two different language model families: GPT-3 and T0++. We use the InstructGPT [37] family of GPT-3 engines, evaluating Ada, Babbage, and Curie since different engines are claimed to be better suited to specific tasks.” Examiner notes that GPT-3 is a large generative multimodal model.)
Regarding Claim 17, Smith and Gu teach The computer-implemented method of claim 8.
Smith does not teach a method in which the different LGMs are different instances of a same LGM provided with different input prompt formats. However, Gu teaches
wherein the different LGMs are different instances of a same LGM provided with different input prompt formats (Gu, Page 202, “we propose a novel semi-supervised ensemble learning algorithm, termed Multi-Train, which generates a number of heterogeneous classifiers that use different classification models and/or different features.” Gu, Page 205, “the proposed algorithm requires the user to pair up the feature manipulation methods and the learning model, each pair representing a base classifier. The corresponding classifier is initially trained with the specified learning algorithm with features manipulated by the pre-specified feature manipulation method. Therefore, a pair of feature manipulation method and a model represents a “view” to the data.” Gu, page 206, “We use three methods to create different views from the same training data, namely, use of different learning models, use differently manipulated features, or a combination of the above.” Examiner notes that utilizing different input formats is the feature manipulation method. Examiner further notes that a view comprises a feature manipulation method and a model, and different views are used meaning different feature manipulation methods are used for different models.)
Smith and Gu are considered analogous to the claimed invention because they involve labeling data using an ensemble of models. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Smith to utilize different instances of a same LGM provided with different input prompt formats in order to increase the diversity of the classifier. Doing so is advantageous because the better performance of Multi-Train is attributed to “the multiple views generated using different models as well as different feature manipulation methods in contrast to the original single-view data” (Gu, 210). Gu’s findings show that “Heterogeneous ensembles have proved to be more effective in achieving diversity [23,26,27], which is also empirically confirmed by statistically better results than those of the tri-training algorithm on the 12 datasets used in this work” (Gu, Page 203).
Claim(s) 3, 12-13, and 18-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Smith et al. (“Language Models in the Loop: Incorporating Prompting into Weak Supervision”) in view of Gu et al. (“Multi-train: A semi-supervised heterogeneous ensemble classifier”) and Manikani et al. (US 2024/0428138 A1)
Regarding Claim 3, Smith and Gu teach the computer-implemented method of claim 2. Smith further teaches
based on the second input prompt format to the second LGM (Smith, Page 5, Figure 2. Examiner notes that the second input prompt format is Prompt 2.)
Smith and Gu do not teach generating the second output format that includes a set of synthetic training data and corresponding synthetic training labels based on the second input prompt format to the second LGM nor training a lightweight classifier model based on the set of synthetic training data and the corresponding synthetic training labels nor generating the second weak training label for the content item utilizing the lightweight classifier model. However, Manikani teaches:
wherein the second LGM is used to generate the second weak training label for the content item by: generating the second output format that includes a set of synthetic training data and corresponding synthetic training labels (Manikani, Page 1, Abstract, “The method includes generating synthetic data by a large language model (LLM) by prompting the LLM with emission classes and few shot examples. The synthetic data includes multiple synthetic data instances and corresponding instance labels. A training dataset is obtained from the synthetic data.”)
training a lightweight classifier model based on the set of synthetic training data and the corresponding synthetic training labels (Manikani, 0007, “The method further includes training a field machine learning (ML) model with training instances which are synthetic data instances from the training dataset and corresponding training labels.”)
and generating the second weak training label for the content item utilizing the lightweight classifier model (Manikani, 0007, “The field ML model generates a predicted probability distribution of a training output class corresponding to a training instance.”)
Smith, Gu, and Manikani are all analogous to the claimed invention because they utilize large generative models to label data. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Smith and Gu to utilize the second LGM to generate a synthetic training dataset in order to train a model to produce the second weak training label. Doing so is advantageous because the trained field ML model would be “specialized to leverage the distilled knowledge from the teaching LLM for accurate classification” (Manikani, 0085).
Claim 12 recites substantially similar limitations to claim 3, and is therefore rejected under the same analysis.
Regarding Claim 13, Smith Gu, and Manikani teach
The computer-implemented method of claim 12, wherein the second LGM generates the set of synthetic training data and the corresponding synthetic training labels based on a second input prompt that provides context for the dataset of subjective content items and rules for generating the set of synthetic training data and the corresponding synthetic training labels. (Smith, Page 2, “we propose a framework for incorporating prompting into programmatic weak supervision, in order to address the above challenges and realize potential benefits from pre-trained language models (Figure 1). We model prompts as labeling functions by adding additional metadata that maps possible completions to target labels or abstentions. For example, if a task is to classify spam comments, a prompt could be “Does the following comment ask the user to click a link?” If the language model responds positively, then this is an indication that the comment is spam. On the other hand, if the model responds negatively then that might be mapped to an abstention because both spam and non-spam comments can lack that property.” Smith, Page 5, Figure 2. Examiner notes that the second input prompt is Prompt 2.)
Regarding Claim 18, Smith teaches
A system for generating accurate training data sets from subjective data (Smith, Page 1, “we investigate how to direct the knowledge contained in pre-trained language models toward the creation of labeled training data”);
comprising: a dataset of content items including a content item not having a training label (Smith, Page 9, “For each dataset in our analysis, we assume the original training split is unlabeled. All labelers, here prompted labeling functions and code-based labeling functions, are applied to the unlabeled training split to generate votes for the true label of each example.”);
generating a set of weak training labels for the content item utilizing the different LGMs (Smith, Page 5, Figure 2. Examiner notes that the set of weak training labels are Labels 1-3.)
determining the training label for the content item from the set of weak training labels using a label ensembler function (Smith, Page 9, “All labeler votes, unless otherwise noted, are combined and denoised using the FlyingSquid [22] label model to estimate a single, probabilistic consensus label per example.”);
and training a lightweight generative machine-learning model using the content item and the training label (Smith, Page 9, “The resulting labels are used to train a RoBERTa [32] end model, which provides a smaller, more servable classification model tailored to our task of interest.”);
Smith does not teach “a set of different large generative machine-learning models (different LGMs) including a first LGM and a second LGM” or “wherein the different LGMs produce different output formats”. However, Gu teaches
a set of different large generative machine-learning models (different LGMs) including a first LGM and a second LGM (Gu, Page 202, “we propose a novel semi-supervised ensemble learning algorithm, termed Multi-Train, which generates a number of heterogeneous classifiers that use different classification models and/or different features.”)
wherein the different LGMs produce different output formats (Gu, Page 205, “the proposed algorithm requires the user to pair up the feature manipulation methods and the learning model, each pair representing a base classifier. The corresponding classifier is initially trained with the specified learning algorithm with features manipulated by the pre-specified feature manipulation method. Therefore, a pair of feature manipulation method and a model represents a “view” to the data.” Gu, page 206, “We use three methods to create different views from the same training data, namely, use of different learning models, use differently manipulated features, or a combination of the above.” Examiner notes that utilizing different output formats is the feature manipulation method. Examiner further notes that a view comprises a feature manipulation method and a model, and different views are used meaning different feature manipulation methods are used for different models.)
Smith and Gu are considered analogous to the claimed invention because they involve labeling data using an ensemble of models. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Smith to instead utilize different LGM models and different output formats in order to increase the diversity of the classifier. Doing so is advantageous because the better performance of Multi-Train is attributed to “the multiple views generated using different models as well as different feature manipulation methods in contrast to the original single-view data” (Gu, 210). Gu’s findings show that “Heterogeneous ensembles have proved to be more effective in achieving diversity [23,26,27], which is also empirically confirmed by statistically better results than those of the tri-training algorithm on the 12 datasets used in this work” (Gu, Page 203).
Smith does not explicitly teach “a processing system comprising a processor; and a computer memory comprising instructions”. However, Manikani teaches
a processing system comprising a processor; and a computer memory comprising instructions that, when executed by the processing system, cause the system to perform operations comprising
(Manikani, 0038, “FIG. 2 shows a flowchart 200 of a method for training a field ML model with synthetic data, in accordance with one or more embodiments. The method of FIG. 2 may be implemented using the components and computing systems shown in the system of FIG. 1 and one or more of the steps may be performed on or received at one or more computer processors.” Manikani, 0090, “Software instructions in the form of computer readable program code to perform embodiments may be stored, in whole or in part, temporarily or permanently, on a non-transitory computer readable medium such as a solid state drive (SSD), compact disk (CD), digital video disk (DVD), storage device, a diskette, a tape, flash memory, physical memory, or any other computer readable storage medium. Specifically, the software instructions may correspond to computer readable program code that, when executed by the computer processor(s) (602), is configured to perform one or more embodiments, which may include transmitting, receiving, presenting, and displaying data and messages described in the other figures of the disclosure.”)
Smith teaches a method comprising an ensemble of LGM models to generate a synthetic training dataset. Manikani teaches a computer medium with instructions to run a method comprising an LGM model used to generate an ensemble of labels to create a synthetic dataset. Thus, Smith and Manikani each teach an ensemble using LGM models for generating synthetic data and thus are both considered analogous to the claimed invention. It would have been obvious to one with ordinary skill in the art to modify Smith to run on a computer medium as Manikani did as this would be applying a known technique to a known device to yield the predictable result of running software instructions on a computer medium successfully (see MPEP 2143(d)).
Regarding Claim 19, Smith, Gu, and Manikani teach
The system of claim 18, wherein: the first LGM directly generates a first weak training label for the content item; and the second LGM indirectly generates a second weak training label for the content item (Smith, Page 5, Figure 2. Examiner notes that the first weak training label is Label 1 and the second weak training label is Label 2.)
Claim 20 recites substantially similar limitations to claim 3, and is therefore rejected under the same analysis.
Claim(s) 6 is/are rejected under 35 U.S.C. 103 as being unpatentable over Smith in view of Gu in further view of Bang et al. (“AISOCRATES: Towards Answering Ethical Quandary Questions”).
Regarding Claim 6, Smith and Gu teach the computer-implemented method of claim 1.
Smith and Gu do not teach further comprising generating multiple separate instances of the first weak training label for the content item by providing different instances of the first input prompt format to the first LGM. However, Bang teaches
further comprising generating multiple separate instances of the first weak training label for the content item by providing different instances of the first input prompt format to the first LGM (Bang, page 3, Figure 1. Examiner notes that the multiple separate instances of the first weak training label are A1 and A2, and the first input prompt format is Prompt 1. Examiner further notes that Prompt 1 is used to generate both A1 and A2, which are separate from each other.)
wherein the different instances of the first input prompt format correspond to different perspectives of the dataset of the subjective content items (Bang, page 3, “Then, the final answer A is obtained by two consecutive generations with the previously selected principles, <p1, p2>, so the answer contains multiple perspectives as addressing the quandary.” Examiner notes that p1 and p2 are principles with different perspectives that correspond to different outputs A1 and A2.)
Smith, Gu, and Bang are all analogous to the claimed invention because they utilize large generative models to label data. It would have been obvious to one having ordinary skill in the art prior to the effective filing data to have modified Smith and Gu to utilize generating multiple separate instances of the first weak training label by providing different instances of the first input prompt, wherein the different instances of the first input prompt format correspond to different perspectives of the dataset. Doing so is advantageous because when faced with questions or prompts that are difficult to answer, “a discussion with multiple perspectives (i.e., a manner of debate) is crucial (Talat et al., 2021; Hendrycks et al., 2020) and sophisticated logical reasoning is required to answer such questions” (Bang, page 1).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Lang et al. (“Co-training Improves Prompt-based Learning for Large Language Models”) also discloses the use of different views to improve classification accuracy for large generative models.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Zared O. Cohen whose telephone number is (571)270-0531. The examiner can normally be reached M-F, 9am to 5pm ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Michelle Bechtold can be reached at (571) 431-0762. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/Z.O.C./Examiner, Art Unit 2148 /MICHELLE T BECHTOLD/Supervisory Patent Examiner, Art Unit 2148