Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
The term “universal” in claim 12 is a relative term which renders the claim indefinite. The term “universal” is not defined by the claim, the specification does not provide a standard for ascertaining the requisite degree, and one of ordinary skill in the art would not be reasonably apprised of the scope of the invention. The term, universal, is too broad, it encompasses too many possibilities of input.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefore, subject to the conditions and requirements of this title.
Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The analysis of the claims will follow the 2019 Revised Patent Subject Matter Eligibility Guidelines (“2019 PEG”).
Step 1: Independent claims 1 (A method comprising:), 9 (A non-transitory computer readable medium storing instructions operable to cause one or more processors to perform operations comprising:), and 14 (A system, comprising:), are directed towards a method, a manufacture, and a system (machine) respectively. Therefore, these claims, as well as their dependent claims, are directed towards one of the four statutory categories (process, machine, manufacture, or composition of matter).
Claim 1
Step 2A, Prong 1: The claim recites, inter alia:
performing, …, a first inference operation to produce first output based on a user request;
This limitation recites a mental process using evaluation, judgment and opinion, with aid of pen and paper to evaluate and infer an answer to a request. See MPEP 2016.04(a)(2)(III);
determining, …, that the first output fails to meet a threshold;
This limitation recites a mental process using evaluation, judgment and opinion, with aid of pen and paper to judge that the answer is not acceptable. See MPEP 2016.04(a)(2)(III);
performing, … based on the first output failing to meet the threshold, a second inference operation to produce second output based on the user request, wherein the second machine learning model has a higher computational complexity than the first machine learning model;
This limitation recites a mental process using evaluation, judgment and opinion, with aid of pen and paper to evaluate and infer a second answer to the request after judging that the first answer is not acceptable. See MPEP 2016.04(a)(2)(III);
determining, …, that the second output meets the threshold; and
This limitation recites a mental process using evaluation, judgment and opinion, with aid of pen and paper to judge that the answer is acceptable. See MPEP 2016.04(a)(2)(III);
Step 2A, Prong 2: The additional elements recited in the claim do not integrate the judicial exception into a practical application.
Additional elements:
… using a first machine learning model of a model chain…
This limitation is recited at a high level of generality and merely invokes use of a generic machine learning model. Mere instructions to apply an exception using a generic machine learning mode cannot integrate a judicial exception into a practical application. See MPEP 2106.05(f)
… using a scoring machine learning model configured to measure quality of output produced by machine learning models of the model chain…
This limitation is recited at a high level of generality and merely invokes use of a generic machine learning model. Mere instructions to apply an exception using a generic machine learning mode cannot integrate a judicial exception into a practical application. See MPEP 2106.05(f)
… using a second machine learning model of the model chain …
This limitation is recited at a high level of generality and merely invokes use of a generic machine learning model. Mere instructions to apply an exception using a generic machine learning mode cannot integrate a judicial exception into a practical application. See MPEP 2106.05(f)
… using the scoring machine learning model…
This limitation is recited at a high level of generality and merely invokes use of a generic machine learning model. Mere instructions to apply an exception using a generic machine learning mode cannot integrate a judicial exception into a practical application. See MPEP 2106.05(f)
transmitting, based on the second output meeting the threshold, the second output in response to the user request.
This limitation is insignificant extra-solution activity of mere data gathering. See MPEP 2106.05(g)
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Additional elements:
… using a first machine learning model of a model chain…
This limitation is recited at a high level of generality and merely invokes use of a generic machine learning model. Mere instructions to apply an exception using a generic machine learning mode cannot integrate a judicial exception into a practical application. See MPEP 2106.05(f)
… using a scoring machine learning model configured to measure quality of output produced by machine learning models of the model chain…
This limitation is recited at a high level of generality and merely invokes use of a generic machine learning model. Mere instructions to apply an exception using a generic machine learning mode cannot integrate a judicial exception into a practical application. See MPEP 2106.05(f)
… using a second machine learning model of the model chain …
This limitation is recited at a high level of generality and merely invokes use of a generic machine learning model. Mere instructions to apply an exception using a generic machine learning mode cannot integrate a judicial exception into a practical application. See MPEP 2106.05(f)
… using the scoring machine learning model…
This limitation is recited at a high level of generality and merely invokes use of a generic machine learning model. Mere instructions to apply an exception using a generic machine learning mode cannot integrate a judicial exception into a practical application. See MPEP 2106.05(f)
transmitting, based on the second output meeting the threshold, the second output in response to the user request.
MPEP 2106.05(d)(II) indicates that receiving or transmitting data is a well-understood, routine, and conventional function when claimed in a merely generic manner or as insignificant extra-solution activity, as it is in this limitation.
Claim 2
Step 2A, Prong 1: The claim recites, inter alia:
producing, by the scoring machine learning model, scoring output including a score to compare against the threshold and data representing a rationalization of the score, and
This limitation recites a mental process using evaluation, judgment, and opinion, with aid of pen and paper, to evaluate quality of output, give opinion on justification of the evaluation, and judge if the evaluation meets a threshold. See MPEP 2016.04(a)(2)(III);
wherein performing the second inference operation to produce the second output based on the user request comprises: performing the second inference operation using input including the user request and the scoring output.
This limitation recites a mental process using evaluation, judgment and opinion, with aid of pen and paper to evaluate and infer a second answer to the request using a request and a score as input. See MPEP 2016.04(a)(2)(III);
Step 2A, Prong 2: There are no further additional elements in this claim.
Step 2B: There are no further additional elements in this claim.
Claim 3
Step 2A, Prong 1: The claim recites, inter alia:
performing the second inference operation using input including the user request and the first output.
This limitation recites a mental process using evaluation, judgment and opinion, with aid of pen and paper to evaluate and infer a second answer to the request using a request and a previous answer as input. See MPEP 2016.04(a)(2)(III);
Step 2A, Prong 2: There are no further additional elements in this claim.
Step 2B: There are no further additional elements in this claim.
Claim 4
Step 2A, Prong 1: The claim recites, inter alia:
wherein performing the second inference operation to produce the second output based on the user request comprises: performing the second inference operation using input including the user request and the labeling data.
This limitation recites a mental process using evaluation, judgment and opinion, with aid of pen and paper to evaluate and infer a second answer to the request using a request and labeling data as input. See MPEP 2016.04(a)(2)(III);
Step 2A, Prong 2: The additional elements recited in the claim do not integrate the judicial exception into a practical application.
Additional elements:
obtaining, based on the first output failing to meet the threshold, labeling data associated with the user request,
This limitation is an insignificant extra-solution activity of mere data gathering. See MPEP 2106.05(g);
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Additional elements:
obtaining, based on the first output failing to meet the threshold, labeling data associated with the user request,
MPEP 2106.05(d)(II) indicates that storing and retrieving information in memory is a well-understood, routine, and conventional function when claimed in a merely generic manner (as it is in the present claim);
Claim 5
Step 2A, Prong 1: The claim recites, inter alia:
selecting the second machine learning model for the model chain based on at least one of the first output or output produced by the scoring machine learning model processing the first output.
This limitation recites a mental process using evaluation, judgment and opinion, with aid of pen and paper to evaluate and select a second machine learning model based on a score or a previous answer. See MPEP 2016.04(a)(2)(III);
Step 2A, Prong 2: There are no further additional elements in this claim.
Step 2B: There are no further additional elements in this claim.
Claim 6
Step 2A, Prong 1: The claim recites, inter alia:
determining the model chain based on at least one of the user request or a user device at which the user request is initiated.
This limitation recites a mental process using evaluation, judgment, and opinion, with aid of pen and paper, to evaluate and select machine learning models based on a request or device. See MPEP 2016.04(a)(2)(III);
Step 2A, Prong 2: There are no further additional elements in this claim.
Step 2B: There are no further additional elements in this claim.
Claim 7
Step 2A, Prong 1: The claim recites, inter alia:
wherein the second output meeting the threshold prevents a performance of a third inference operation using a third machine learning model of the model chain, wherein the third machine learning model has a higher computational complexity than the second machine learning model.
This limitation recites a mental process using evaluation, judgment, and opinion, with aid of pen and paper, to judge that a third inference operation is not needed. See MPEP 2016.04(a)(2)(III);
Step 2A, Prong 2: There are no further additional elements in this claim.
Step 2B: There are no further additional elements in this claim.
Claim 8
Step 2A, Prong 1: There are no further abstract ideas in this claim.
Step 2A, Prong 2: The additional elements recited in the claim do not integrate the judicial exception into a practical application.
Additional elements:
wherein the first machine learning model and the second machine learning model are a same type of machine learning model trained to produce a same type of output.
This limitation indicates specifying that the answer output of two machine learning models must be the same type and that the machine learning model that the answer output is output from must be the same. Merely indicating a field of use or technological environment in which to apply a judicial exception cannot integrate the judicial exception into a practical application. See MPEP 2106.05(h);
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Additional elements:
wherein the first machine learning model and the second machine learning model are a same type of machine learning model trained to produce a same type of output.
This limitation indicates specifying that the answer output of two machine learning models must be the same type and that the machine learning model that the answer output is output from must be the same. Merely indicating a field of use or technological environment in which to apply a judicial exception is not significantly more than the judicial. See MPEP 2106.05(h);
Claim 9
Step 2A, Prong 1: The claim recites, inter alia:
performing, …, a first inference operation to produce first output based on a user request;
This limitation recites a mental process using evaluation, judgment and opinion, with aid of pen and paper to evaluate and infer an answer to a request. See MPEP 2016.04(a)(2)(III);
determining, …, that the first output fails to meet a threshold;
This limitation recites a mental process using evaluation, judgment and opinion, with aid of pen and paper to judge that the answer is not acceptable. See MPEP 2016.04(a)(2)(III);
performing, [[using a second machine learning model of the model chain]] based on the first output failing to meet the threshold, a second inference operation to produce second output based on the user request, wherein the second machine learning model has a higher computational complexity than the first machine learning model;
This limitation recites a mental process using evaluation, judgment and opinion, with aid of pen and paper to evaluate and infer a second answer to the request based on a judgement that the first answer is not acceptable. See MPEP 2016.04(a)(2)(III);
determining, …, that the second output meets the threshold; and
This limitation recites a mental process using evaluation, judgment and opinion, with aid of pen and paper to judge that the answer is acceptable. See MPEP 2016.04(a)(2)(III);
Step 2A, Prong 2: The additional elements recited in the claim do not integrate the judicial exception into a practical application.
Additional elements:
A non-transitory computer readable medium storing instructions operable to cause one or more processors to perform operations comprising:
This limitation is recited at a high level of generality and invokes use of general computer equipment merely as a tool to apply instructions. Mere instructions to apply an exception using general computer equipment cannot integrate a judicial exception into a practical application. See MPEP 2106.05(f)
… using a first machine learning model of a model chain…
This limitation is recited at a high level of generality and merely invokes use of a generic machine learning model. Mere instructions to apply an exception using a generic machine learning mode cannot integrate a judicial exception into a practical application. See MPEP 2106.05(f)
… using a scoring machine learning model configured to measure quality of output produced by machine learning models of the model chain…
This limitation is recited at a high level of generality and merely invokes use of a generic machine learning model. Mere instructions to apply an exception using a generic machine learning mode cannot integrate a judicial exception into a practical application. See MPEP 2106.05(f)
… using the scoring machine learning model…
This limitation is recited at a high level of generality and merely invokes use of a generic machine learning model. Mere instructions to apply an exception using a generic machine learning mode cannot integrate a judicial exception into a practical application. See MPEP 2106.05(f)
transmitting, based on the second output meeting the threshold, the second output in response to the user request.
This limitation is insignificant extra-solution activity of mere data gathering. See MPEP 2106.05(g)
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Additional elements:
A non-transitory computer readable medium storing instructions operable to cause one or more processors to perform operations comprising:
This limitation is recited at a high level of generality and invokes use of general computer equipment merely as a tool to apply instructions. Mere instructions to apply an exception using general computer equipment is not significantly more than the judicial exception. See MPEP 2106.05(f)
… using a first machine learning model of a model chain…
This limitation is recited at a high level of generality and merely invokes use of a generic machine learning model. Mere instructions to apply an exception using a generic machine learning mode is not significantly more than the judicial exception. See MPEP 2106.05(f)
… using a scoring machine learning model configured to measure quality of output produced by machine learning models of the model chain…
This limitation is recited at a high level of generality and merely invokes use of a generic machine learning model. Mere instructions to apply an exception using a generic machine learning mode is not significantly more than the judicial exception. See MPEP 2106.05(f)
… using the scoring machine learning model…
This limitation is recited at a high level of generality and merely invokes use of a generic machine learning model. Mere instructions to apply an exception using a generic machine learning mode is not significantly more than the judicial exception. See MPEP 2106.05(f)
transmitting, based on the second output meeting the threshold, the second output in response to the user request.
MPEP 2106.05(d)(II) indicates that receiving or transmitting data is a well-understood, routine, and conventional function when claimed in a merely generic manner or as insignificant extra-solution activity, as it is in this limitation.
Claim 10
Step 2A, Prong 1: The claim recites, inter alia:
wherein the second machine learning model performs the second inference operation using input including the user request and the first output.
This limitation recites a mental process using evaluation, judgment and opinion, with aid of pen and paper to evaluate and infer a second answer to the request using a request and a previous answer as input. See MPEP 2016.04(a)(2)(III);
Step 2A, Prong 2: There are no further additional elements in this claim.
Step 2B: There are no further additional elements in this claim.
Claim 11
Step 2A, Prong 1: There are no further abstract ideas in this claim.
Step 2A, Prong 2: The additional elements recited in the claim do not integrate the judicial exception into a practical application.
Additional elements:
wherein the machine learning models of the model chain are language learning models.
This limitation indicates limiting a field of use to a general class of machine learning model. Merely indicating that an exception is to be limited to a general class of machine learning models cannot integrate a judicial exception into a practical application. See MPEP 2106.05(h)
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Additional elements:
wherein the machine learning models of the model chain are language learning models.
This limitation indicates limiting a field of use to a general class of machine learning model. Merely indicating that an exception is to be limited to a general class of machine learning model is not significantly more than the judicial exception. See MPEP 2106.05(h)
Claim 12
Step 2A, Prong 1: There are no further abstract ideas in this claim.
Step 2A, Prong 2: The additional elements recited in the claim do not integrate the judicial exception into a practical application.
Additional elements:
wherein the model chain is defined for universal use with user requests.
This limitation merely indicates specifying that the model chain must be able to be used with any user request. Merely indicating a field of use or technological environment in which to apply a judicial exception does not integrate the judicial exception into a practical application. See MPEP 2106.05(h);
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Additional elements:
wherein the model chain is defined for universal use with user requests.
This limitation merely indicates specifying that the model chain must be able to be used with any user request. Merely indicating a field of use or technological environment in which to apply a judicial exception is not significantly more than the judicial. See MPEP 2106.05(h);
Claim 13
Step 2A, Prong 1: There are no further abstract ideas in this claim.
Step 2A, Prong 2: The additional elements recited in the claim do not integrate the judicial exception into a practical application.
Additional elements:
wherein the first machine learning model is implemented at a user device and the second machine learning model is implemented at a server device.
This limitation merely indicates specifying that the first and second machine learning models must be on different devices. Merely indicating a field or use or technological environment in which to apply a judicial exception does not integrate the judicial exception into a practical application. See MPEP 2106.05(h);
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Additional elements:
wherein the first machine learning model is implemented at a user device and the second machine learning model is implemented at a server device.
This limitation merely indicates specifying that the first and second machine learning models must be on different devices. Merely indicating a field or use or technological environment in which to apply a judicial exception does not integrate the judicial exception into a practical application. See MPEP 2106.05(h);
Claim 14
Step 2A, Prong 1: The claim recites, inter alia:
perform, …, a first inference operation to produce first output based on a user request;
This limitation recites a mental process using evaluation, judgment and opinion, with aid of pen and paper to evaluate and infer an answer to a request. See MPEP 2016.04(a)(2)(III);
determine, …, that the first output fails to meet a threshold;
This limitation recites a mental process using evaluation, judgment and opinion, with aid of pen and paper to judge that the answer is not acceptable. See MPEP 2016.04(a)(2)(III);
perform, [[using a second machine learning model of the model chain]] based on the first output failing to meet the threshold, a second inference operation to produce second output based on the user request, wherein the second machine learning model has a higher computational complexity than the first machine learning model;
This limitation recites a mental process using evaluation, judgment and opinion, with aid of pen and paper to evaluate and infer a second answer to the request based on a judgement that the first answer was not acceptable. See MPEP 2016.04(a)(2)(III);
determine, …, that the second output meets the threshold; and
This limitation recites a mental process using evaluation, judgment and opinion, with aid of pen and paper to judge that the answer is acceptable. See MPEP 2016.04(a)(2)(III);
Step 2A, Prong 2: The additional elements recited in the claim do not integrate the judicial exception into a practical application.
Additional elements:
A system, comprising: a memory subsystem storing instructions; and processing circuitry configured to execute the instructions
This limitation is recited at a high level of generality and merely invokes use of general computer equipment. Mere instructions to apply an exception using general computer equipment cannot integrate a judicial exception into a practical application. See MPEP 2106.05(f)
… using a first machine learning model of a model chain…
This limitation is recited at a high level of generality and merely invokes use of a generic machine learning model. Mere instructions to apply an exception using a generic machine learning mode cannot integrate a judicial exception into a practical application. See MPEP 2106.05(f)
… using a scoring machine learning model configured to measure quality of output produced by machine learning models of the model chain…
This limitation is recited at a high level of generality and merely invokes use of a generic machine learning model. Mere instructions to apply an exception using a generic machine learning mode cannot integrate a judicial exception into a practical application. See MPEP 2106.05(f)
… using the scoring machine learning model…
This limitation is recited at a high level of generality and merely invokes use of a generic machine learning model. Mere instructions to apply an exception using a generic machine learning mode cannot integrate a judicial exception into a practical application. See MPEP 2106.05(f)
transmit, based on the second output meeting the threshold, the second output in response to the user request.
This limitation is insignificant extra-solution activity of mere data gathering. See MPEP 2106.05(g)
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Additional elements:
A system, comprising: a memory subsystem storing instructions; and processing circuitry configured to execute the instructions
This limitation is recited at a high level of generality and merely invokes use of general computer equipment. Mere instructions to apply an exception using general computer equipment is not significantly more than the judicial exception. See MPEP 2106.05(f)
… using a first machine learning model of a model chain…
This limitation is recited at a high level of generality and merely invokes use of a generic machine learning model. Mere instructions to apply an exception using a generic machine learning mode is not significantly more than the judicial exception. See MPEP 2106.05(f)
… using a scoring machine learning model configured to measure quality of output produced by machine learning models of the model chain…
This limitation is recited at a high level of generality and merely invokes use of a generic machine learning model. Mere instructions to apply an exception using a generic machine learning mode is not significantly more than the judicial exception. See MPEP 2106.05(f)
… using the scoring machine learning model…
This limitation is recited at a high level of generality and merely invokes use of a generic machine learning model. Mere instructions to apply an exception using a generic machine learning mode is not significantly more than the judicial exception. See MPEP 2106.05(f)
transmitting, based on the second output meeting the threshold, the second output in response to the user request.
MPEP 2106.05(d)(II) indicates that receiving or transmitting data is a well-understood, routine, and conventional function when claimed in a merely generic manner or as insignificant extra-solution activity, as it is in this limitation.
Claim 15
Step 2A, Prong 1: The claim recites, inter alia:
determine, …, that the second output fails to meet a threshold;
This limitation recites a mental process using evaluation, judgment and opinion, with aid of pen and paper to judge that an answer is not acceptable. See MPEP 2016.04(a)(2)(III);
perform, …, multiple third inference operations in parallel based on the user request, wherein each of the third inference operations produces different third output;
This limitation recites a mental process using evaluation, judgment and opinion, with aid of pen and paper to evaluate and infer multiple answers to the request at the same time. See MPEP 2016.04(a)(2)(III);
determine, …, a third output having a highest score amongst the different third output; and
This limitation recites a mental process using evaluation, judgment and opinion, with aid of pen and paper to evaluate answer scores and judge that an answer score is highest. See MPEP 2016.04(a)(2)(III);
Step 2A, Prong 2: The additional elements recited in the claim do not integrate the judicial exception into a practical application.
Additional elements:
… using the scoring machine learning model…
This limitation is recited at a high level of generality and merely invokes use of a generic machine learning model. Mere instructions to apply an exception using a generic machine learning mode cannot integrate a judicial exception into a practical application. See MPEP 2106.05(f)
… using a last machine learning model of the model chain…
This limitation is recited at a high level of generality and merely invokes use of a generic machine learning model. Mere instructions to apply an exception using a generic machine learning mode cannot integrate a judicial exception into a practical application. See MPEP 2106.05(f)
… using the scoring machine learning model…
This limitation is recited at a high level of generality and merely invokes use of a generic machine learning model. Mere instructions to apply an exception using a generic machine learning mode cannot integrate a judicial exception into a practical application. See MPEP 2106.05(f)
transmit the third output in response to the user request.
This limitation is insignificant extra-solution activity of mere data gathering. See MPEP 2106.05(g)
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Additional elements:
… using the scoring machine learning model…
This limitation is recited at a high level of generality and merely invokes use of a generic machine learning model. Mere instructions to apply an exception using a generic machine learning mode is not significantly more than the judicial exception. See MPEP 2106.05(f)
… using a last machine learning model of the model chain…
This limitation is recited at a high level of generality and merely invokes use of a generic machine learning model. Mere instructions to apply an exception using a generic machine learning mode is not significantly more than the judicial exception. See MPEP 2106.05(f)
… using the scoring machine learning model…
This limitation is recited at a high level of generality and merely invokes use of a generic machine learning model. Mere instructions to apply an exception using a generic machine learning mode is not significantly more than the judicial exception. See MPEP 2106.05(f)
transmit the third output in response to the user request.
MPEP 2106.05(d)(II) indicates that receiving or transmitting data is a well-understood, routine, and conventional function when claimed in a merely generic manner or as insignificant extra-solution activity, as it is in this limitation.
Claim 16
Step 2A, Prong 1: The claim recites, inter alia:
wherein the second machine learning model performs the second inference operation using input including the user request, information representative of the first output, and information representative of scoring output produced by the scoring machine learning model processing the first output.
This limitation recites a mental process using evaluation, judgment and opinion, with aid of pen and paper to evaluate and infer a second answer to the request using a request, data representing a previous answer, and data representing a score as input. See MPEP 2016.04(a)(2)(III);
Step 2A, Prong 2: There are no further additional elements in this claim.
Step 2B: There are no further additional elements in this claim.
Claim 17
Step 2A, Prong 1: The claim recites, inter alia:
wherein the second machine learning model performs the second inference operation using input including the user request, information representative of the first output, and label information associated with the first output.
This limitation recites a mental process using evaluation, judgment and opinion, with aid of pen and paper to evaluate and infer a second answer to the request using a request, data representing a previous answer, and data representing a label related to the previous answer. See MPEP 2016.04(a)(2)(III);
Step 2A, Prong 2: There are no further additional elements in this claim.
Step 2B: There are no further additional elements in this claim.
Claim 18
Step 2A, Prong 1: There are no further abstract ideas in this claim.
Step 2A, Prong 2: The additional elements recited in the claim do not integrate the judicial exception into a practical application.
Additional elements:
generate the scoring machine learning model as a discriminative regression model using supervised learning.
This limitation indicates limiting the machine learning model to a specific type of machine learning model. Merely indicating a field of use or technological environment in which to apply a judicial exception cannot integrate the judicial exception into a practical application and does not amount to significantly more than the judicial exception. See MPEP 2106.05(h);
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Additional elements:
generate the scoring machine learning model as a discriminative regression model using supervised learning.
This limitation indicates limiting the machine learning model to a specific type of machine learning model. Merely indicating a field of use or technological environment in which to apply a judicial exception cannot integrate the judicial exception into a practical application and does not amount to significantly more than the judicial exception. See MPEP 2106.05(h);
Claim 19
Step 2A, Prong 1: The claim recites, inter alia:
wherein the machine learning models of the model chain are arranged in order of computational complexity in which a lowest computational complexity model of the machine learning models is first in the model chain and a highest computational complexity model of the machine learning models is last in the model chain.
This limitation recites a mental process using evaluation, judgment and opinion, with aid of pen and paper to order a series of machine learning models based on complexity. See MPEP 2016.04(a)(2)(III);
Step 2A, Prong 2: There are no further additional elements in this claim.
Step 2B: There are no further additional elements in this claim.
Claim 20
Step 2A, Prong 1: There are no further abstract ideas in this claim.
Step 2A, Prong 2: The additional elements recited in the claim do not integrate the judicial exception into a practical application.
Additional elements:
wherein the user request is initiated in connection with a software service of a unified communications as a service platform.
This limitation indicates language specifying that the user request was to be implemented using “a software service of a unified communications as a service platform” that broadly includes multiple forms of communication. Merely indicating a field of use or technological environment in which to apply a judicial exception cannot integrate the judicial exception into a practical application and does not amount to significantly more than the judicial exception. See MPEP 2106.05(h);
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Additional elements:
wherein the user request is initiated in connection with a software service of a unified communications as a service platform.
This limitation indicates language specifying that the user request was to be implemented using “a software service of a unified communications as a service platform” that broadly includes multiple forms of communication. Merely indicating a field of use or technological environment in which to apply a judicial exception cannot integrate the judicial exception into a practical application and does not amount to significantly more than the judicial exception. See MPEP 2106.05(h);
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claim(s) 1, and 4-8 is/are rejected under 35 U.S.C. 102(a)(1) as being anticipated by FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance by Chen et al., hereafter Chen.
Regarding claim 1, Chen teaches:
A method comprising:
performing, using a first machine learning model of a model chain (GPT-J), a first inference operation (GPT-J response generation) to produce first output (GPT-J answer/response) based on a user request (query); (Page 4, Figure 2e, GPT-J outputs an inference based on a query; Page 5, Section 3, subsection Strategy 3, Paragraph 1.)
determining, using a scoring machine learning model (generation scoring function) configured to measure quality of output (whether a generation is correct) produced by machine learning models of the model chain, that the first output fails (“<”) to meet a threshold (“0.5”); (Page 4, Figure 2e; Page 5, Section 3, subsection Strategy 3, paragraphs 2-3, “It returns the generation if the score is higher than a threshold, and queries the next service otherwise… The scoring function can be obtained by training a simple regression model that learns whether a generation is correct from the query and a generated answer.”)
performing, using a second machine learning model of the model chain (GPT-3) based on the first output failing (“<”) to meet the threshold (“0.5”), a second inference operation (GPT-3 response generation) to produce second output (GPT-3 answer/response) based on the user request (query), wherein the second machine learning model (GPT-3) has a higher computational complexity than the first machine learning model (GPT-J); (Page 4, Figure 2e; Page 5, Section 3, subsection Strategy 3, paragraph 1; Page 1, Section Introduction, Paragraph 4)
determining, using the scoring machine learning model, that the second output meets (not “<”) the threshold (“0.9”); and (Page 4, Figure 2e, GPT-3 does not have a score less than the new threshold; Page 5, Section 3, subsection Strategy 3, paragraphs 2-3, “It returns the generation if the score is higher than a threshold…”)
transmitting (accepting and returning), based on the second output meets (not “<”) the threshold (“0.9”), the second output (GPT-3 answer) in response to the user request (query). (Page 4, Figure 2e; Page 5, Section 3, subsection Strategy 3, paragraph 1-2)
Regarding claim 4, Chen teaches the material disclosed in claim 1, additionally Chen teaches:
obtaining (receiving the converted corresponding prompt), based on the first output failing to meet the threshold, labeling data (prompt corresponding to a query) associated with the user request (query), (Page 3, Section 2, subsection LLM marketplace, “Note that to use LLM APIs, one has to convert each query q to some corresponding prompt first.” The corresponding prompt is labeling data associated with the user request query. The function generating an answer includes the labeling data prompt and the converted query user request)
wherein performing the second inference operation to produce the second output based on the user request comprises:
performing the second inference operation (GPT-3 response generation) using input including the user request (query) and the labeling data (corresponding prompt). (Page 3, Section 2, subsection LLM marketplace, “Each fi[ ]: P->A is a function that, given a prompt p from the prompt space P, generates an answer from the answer distribution A. Note that to use LLM APIs, one has to convert each query q to some corresponding prompt first.” The corresponding prompt is labeling data associated with the user request query. The function generating an answer includes the labeling data prompt and the converted query user request)
Regarding claim 5, Chen teaches the material in claim 1, and additionally teaches:
selecting the second machine learning model (invokes ith API) for the model chain based on at least one of the first output or output produced by the scoring machine learning model (scoring function to generate a score. It returns the generation if the score is higher than a threshold and queries the next service otherwise) processing the first output. (Page 4, Figure 2e, the score of GPT-J is less than a threshold; Page 5, Section 3, subsection Strategy 3, paragraphs 2-3, “The LLM router selects m LLM APIs to include in the list. … Given a new query, it iteratively invokes the ith API in the list to obtain an answer. Then, it uses the scoring function to generate a score. It returns the generation if the score is higher than a threshold and queries the next service otherwise.”)
Regarding claim 6, Chen teaches the material in claim 1, and additionally teaches:
determining the model chain (LLM selection) based on at least one of the user request (query) or a user device at which the user request is initiated. (Page 6, Section 3, subsection Compositions, “For instance, joint prompt and LLM selection is a composition of prompt selection and LLM cascade: for a given query, it searches for the smallest prompt and most affordable LLM that achieves satisfactory task performance.”)
Regarding claim 7, Chen teaches the material in claim 1, and additionally teaches:
wherein the second output (GPT-3 generation) meeting the threshold (sufficiently reliable) prevents a performance (no further LLMS in the list are needed) of a third inference operation (GPT-4 inference) using a third machine learning model of the model chain (GPT-4), wherein the third machine learning model has a higher computational complexity than the second machine learning model. (Page 4, Figure 2e, GPT-4 inference is shaded out and is not performed because GPT-3’s score was not less than the threshold.; Page 5, Section 3, subsection Strategy 3, paragraph 1, “The remaining LLM APIs are queried only if the previous APIs' generations are deemed insufficiently reliable.” GPT-4 has higher computational complexity than GPT-3.)
Regarding claim 8, Chen teaches the material in claim 1, and additionally teaches:
wherein the first machine learning model (GPT-J) and the second machine learning model (GPT-3) are a same type of machine learning model (LLM) trained to produce a same type of output (output an answer, ex. “camouflage”). (Page 4, Figure 2e, both the first and second machine learning models (GPT-J and GPT-3) are LLM models trained to output an answer.)
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claim(s) 2 is/are rejected under 35 U.S.C. 103 as being unpatentable over Chen in view of Towards an On-device Agent for Text Rewriting by Zhu et al., hereafter Zhu.
Regarding claim 2, Chen teaches the material disclosed in claim 1, and additionally Chen teaches:
producing, by the scoring machine learning model (scoring function), scoring output (reliability score) including a score to compare against the threshold (threshold Ti)…, ((Chen) Page 5, Section 3 subsection Strategy 3, “Then, it uses the scoring function to generate a score”)and
wherein performing the second inference operation to produce the second output based on the user request comprises:
performing the second inference operation (GPT-3 generation) using input including the user request (query) and the scoring output (score g(q, f(q))). ((Chen) Page 4, Figure 2e, Reliability score output and the query are used as input to the GPT-3 model; Page 5, Section 3 subsection Strategy 3, paragraph 2)
Chen does not explicitly disclose:
… data representing a rationalization of the score
Zhu teaches:
Explanation data representing a rationalization of the score (#Explanation). ((Zhu) Page 20, Table 10, #Explanation gives rationalization for the quality grade choice.)
Zhu and Chen are analogous art because they are in the same area of invention, cascading machine learning models.
Thus, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention, having the references in front of them, to have combined the teachings of Zhu of explanation data with the scoring function taught by Chen. The motivation for this would have been to leverage LLMs capability to effectively judge the quality of responses. This combination of a scoring function, as taught by Chen, with additional explanation data, as Zhu teaches, would not change the functionality of either component and would yield the predictable result that is the invention specified in claim 2 of the instant application.
Claim(s) 3 is/are rejected under 35 U.S.C. 103 as being unpatentable over Chen in view of Boosted Prompt Ensembles for Large Language Models by Pitis et al., hereafter Pitis.
Regarding claim 3, Chen teaches the material claimed in claim 1.
Chen does not teach:
performing the second inference operation using input including the user request and the first output.
Pitis teaches:
Using first model prediction outputs (substitute ground truth answer labels with model predictions) and the user request (prompt) as input to perform a second inference operation (inference). ((Pitis) Page 3, Algorithm 1; Page 3-4, Section 3, Subsections Train-time Boosting - Test-time Boosting, “To perform inference at test time, we use our language model to generate m chain of thought answers for each of the n prompt in our boosted ensemble…In this case, we substitute ground truth answer labels with model predictions…The algorithm is otherwise the same as the train-time algorithm”)
Pitis and Chen are analogous art because they are in the same area of invention, that using iterative and successive machine learning models.
Thus, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention, having the references in front of them, to have combined the teachings of Pitis of using a previous first model prediction output as input in addition to a user request to perform a second inference operation with using a user request to perform a second inference operation taught by Chen. The motivation for this would be to adapt to the absence of ground truth answer labels. This combination of an output as an additional input, as Pitis teaches, to the second inference operation using a user request, as taught by Chen, would not change the functionality of either component and would yield the predictable result specified in claim 3 of the instant application.
Claim(s) 9, 11, 12, 14, 19, and 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Chen in view of US 20150242747 A1 by Packes et al., hereafter Packes.
Regarding claim 9, Chen teaches:
performing, using a first machine learning model of a model chain (GPT-J), a first inference operation (GPT-J response generation) to produce first output (GPT-J answer/response) based on a user request (query); (Page 4, Figure 2e, GPT-J outputs an inference based on a query; Page 5, Section 3, subsection Strategy 3, Paragraph 1.)
determining, using a scoring machine learning model (generation scoring function) configured to measure quality of output (whether a generation is correct) produced by machine learning models of the model chain, that the first output fails (“<”) to meet a threshold (“0.5”); (Page 4, Figure 2e; Page 5, Section 3, subsection Strategy 3, paragraphs 2-3, “It returns the generation if the score is higher than a threshold, and queries the next service otherwise… The scoring function can be obtained by training a simple regression model that learns whether a generation is correct from the query and a generated answer.”)
performing, using a second machine learning model of the model chain (GPT-3) based on the first output failing (“<”) to meet the threshold (“0.5”), a second inference operation (GPT-3 response generation) to produce second output (GPT-3 answer/response) based on the user request (query), wherein the second machine learning model (GPT-3) is a higher computational complexity model than the first machine learning model (GPT-J); (Page 4, Figure 2e; Page 5, Section 3, subsection Strategy 3, paragraph 1; Page 1, Section Introduction, Paragraph 4)
determining, using the scoring machine learning model, that the second output meets (not “<”) the threshold (“0.9”); and (Page 4, Figure 2e, GPT-3 does not have a score less than the new threshold; Page 5, Section 3, subsection Strategy 3, paragraphs 2-3, “It returns the generation if the score is higher than a threshold…”)
transmitting (accepting and returning), based on the second output meets (not “<”) the threshold (“0.9”), the second output (GPT-3 answer) in response to the user request (query). (Page 4, Figure 2e; Page 5, Section 3, subsection Strategy 3, paragraph 1-2)
Chen does not teach:
A non-transitory computer readable medium storing instructions operable to cause one or more processors to perform operations comprising:
Packes teaches:
A non-transitory computer-readable medium storing instructions accessible to processors to perform operations ((Packes) Paragraph [0026], “For example, in some embodiments, a non-transitory computer-readable medium may include computer-readable instructions recorded thereon for training, with a processing system”)
Packes and Chen are analogous art because they are in the same area of invention, that using cascading machine learning models.
Thus, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention, having the references in front of them, to have combined the teachings of Packes of implementing the method taught by Chen using a non-transitory computer-readable medium. The motivation for this would be to provide a physical body to perform the method. This combination of a method, as taught by Chen, with a manufacture on which to run it one, as Packes teaches, would not change the functionality of either component and would yield the predictable result specified in claim 9 of the instant application. The Examiner notes that this motivation applies to all dependent and/or otherwise subsequently addressed claims.
Regarding claim 11, Chen, in view of Packes teaches the material disclosed in claim 9, and Chen additionally teaches:
wherein the machine learning models of the model chain (LLM chain) are language learning models. ((Chen) Page 4, Figure 2e, LLM stands for large language learning model)
Regarding claim 12, Chen, in view of Packes teaches the material disclosed in claim 9, and Chen additionally teaches:
wherein the model chain is defined for universal use (data-adaptive) with user requests (query). ((Chen) Page 5, Section 3, Strategy 3, lines 1-2, “…The increasing availability of LLM APIs with heterogeneous performance and costs presents a unique opportunity for data-adaptive LLM selection”)
Regarding claim 14, Chen teaches:
perform, using a first machine learning model of a model chain (GPT-J), a first inference operation (GPT-J response generation) to produce first output (GPT-J answer/response) based on a user request (query); (Page 4, Figure 2e, GPT-J outputs an inference based on a query; Page 5, Section 3, subsection Strategy 3, Paragraph 1.)
determine, using a scoring machine learning model (generation scoring function) configured to measure quality of output (whether a generation is correct) produced by machine learning models of the model chain, that the first output fails (“<”) to meet a threshold (“0.5”); (Page 4, Figure 2e; Page 5, Section 3, subsection Strategy 3, paragraphs 2-3, “It returns the generation if the score is higher than a threshold, and queries the next service otherwise… The scoring function can be obtained by training a simple regression model that learns whether a generation is correct from the query and a generated answer.”)
perform, using a second machine learning model of the model chain (GPT-3) based on the first output failing (“<”) to meet the threshold (“0.5”), a second inference operation (GPT-3 response generation) to produce second output (GPT-3 answer/response) based on the user request (query), wherein the second machine learning model (GPT-3) has a higher computational complexity than the first machine learning model (GPT-J); (Page 4, Figure 2e; Page 5, Section 3, subsection Strategy 3, paragraph 1; Page 1, Section Introduction, Paragraph 4)
determine, using the scoring machine learning model, that the second output meets (not “<”) the threshold (“0.9”); and (Page 4, Figure 2e, GPT-3 does not have a score less than the new threshold; Page 5, Section 3, subsection Strategy 3, paragraphs 2-3, “It returns the generation if the score is higher than a threshold…”)
transmit (accepting and returning), based on the second output meets (not “<”) the threshold (“0.9”), the second output (GPT-3 answer) in response to the user request (query). (Page 4, Figure 2e; Page 5, Section 3, subsection Strategy 3, paragraph 1-2)
Chen does not teach:
A system, comprising:
a memory subsystem storing instructions; and
processing circuitry configured to execute the instructions
Packes teaches:
A processing system configured to execute instructions from a non-transitory computer-readable medium storing the instructions ((Packes) Paragraph [0026], “For example, in some embodiments, a non-transitory computer-readable medium may include computer-readable instructions recorded thereon for training, with a processing system”)
Packes and Chen are analogous art because they are in the same area of invention, that using cascading machine learning models.
Thus, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention, having the references in front of them, to have combined the teachings of Packes of implementing the method taught by Chen using a processing system. The motivation for this would be to provide a physical body to perform the method. This combination of a method, as taught by Chen, with a machine on which to run it on, as Packes teaches, would not change the functionality of either component and would yield the predictable result specified in claim 14 of the instant application. The Examiner notes that this motivation applies to all dependent and/or otherwise subsequently addressed claims.
Regarding claim 19, Chen, in view of Packes teaches the material disclosed in claim 14, and Chen additionally teaches:
wherein the machine learning models of the model chain are arranged in order of computational complexity in which a lowest computational complexity model (GPT-J) of the machine learning models is first in the model chain and a highest computational complexity model (GPT-4) of the machine learning models is last in the model chain. ((Chen) Page 4, Figure 2e, GPT-4 is more complex than GPT-3. GPT-3 is more complex than GPT-J. GPT-J is first and GPT-4 is last.)
Regarding claim 20, Chen, in view of Packes teaches the material disclosed in claim 14.
Chen does not teach:
wherein the user request is initiated in connection with a software service of a unified communications as a service platform.
Packes teaches:
Wherein the user request is initiated in connection with a software service of a unified communications as a service platform (messaging subcomponent). ((Packes) Paragraph [0114], “The messaging subcomponent may facilitate message communications such as email, instant messaging, Voice over IP (VoIP), video conferencing, Short Message Service (SMS), web chat, in-app messaging (e.g., alerts, notifications), and/or the like.”)
Packes and Chen are analogous art because they are in the same area of invention, that using cascading machine learning models.
Thus, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention, having the references in front of them, to have combined the teachings of Packes of a messaging subcomponent with unified message communication capability to initiate a user request taught by Chen. The motivation for this would be to allow increased diversity of possible inputs. This application of initiating a user request, as taught by Chen, using a messaging subcomponent, such as Packes teaches, would be applicable by a person having ordinary skill in the art and would yield the predictable result specified in claim 20 of the instant application.
Claim(s) 10, 16, and 17 is/are rejected under 35 U.S.C. 103 as being unpatentable over Chen, in view of Packes, in further view of Pitis.
Regarding claim 10, Chen, in view of Packes, teaches the material disclosed in claim 9.
Chen, in view of Packes, does not teach:
wherein the second machine learning model performs the second inference operation using input including the user request and the first output.
Pitis teaches:
Using first output (model predictions) and the query user request as input to perform a second inference. ((Pitis) Page 3, Algorithm 1; Page 4, Section 3, Subsection Test-time Boosting, line 5-23, “In this case, we substitute ground truth answer labels with model predictions…The algorithm is otherwise the same as the train-time algorithm”)
Pitis and Chen are analogous art because they are in the same area of invention, that using iterative and successive machine learning models.
Thus, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention, having the references in front of them, to have combined the teachings of Pitis of using a previous first model prediction output as input in addition to a user request to perform a second inference operation with using a user request to perform a second inference operation taught by Chen. The motivation for this would be to adapt to the absence of ground truth answer labels. This combination of an output as an additional input, as Pitis teaches, to the second inference operation using a user request, as taught by Chen, would not change the functionality of either component and would yield the predictable result specified in claim 10 of the instant application.
Regarding claim 16, Chen, in view of Packes, teaches the material disclosed in claim 14, and Chen additionally teaches:
wherein the second machine learning model (GPT-3) performs the second inference operation (GPT-3 answer/response/generation) using input including the user request (query), …, and information representative of scoring output produced by the scoring machine learning model processing the first output (output of “score<0.5”). ((Chen) Page 4, Figure 2e.)
Chen, in view of Packes, does not teach:
Wherein the second machine learning model performs the second inference operation using input including … information representative of the first output …
Pitis teaches:
An inference using first output (model predictions) and the user request (prompt) as input. ((Pitis) Page 3, Algorithm 1; Page 4, Section 3, Subsection Test-time Boosting, line 5-23, “In this case, we substitute ground truth answer labels with model predictions…The algorithm is otherwise the same as the train-time algorithm”)
Pitis and Chen are analogous art because they are in the same area of invention, that using iterative and successive machine learning models.
Thus, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention, having the references in front of them, to have combined the teachings of Pitis of using a previous first model prediction output as input in addition to a user request to perform a second inference operation with using a user request and information representative of scoring output to perform a second inference operation taught by Chen. The motivation for this would be to adapt to the absence of ground truth answer labels. This combination of an output as an additional input, as Pitis teaches, to the second inference operation using a user request and information representative of scoring output, as taught by Chen, would not change the functionality of either component and would yield the predictable result specified in claim 16 of the instant application.
Regarding claim 17, Chen, in view of Packes, teaches the material disclosed in claim 14, and Chen additionally teaches:
wherein the second machine learning model (GPT-3) performs the second inference operation (GPT-3 answer/response/generation) using input including the user request (query)… ((Chen) Page 4, Figure 2e; Page 5, Section 3, subsection Strategy 3, paragraph 1)
Chen, in view of Packes, does not teach:
wherein the second machine learning model performs the second inference operation using input including … information representative of the first output, and label information associated with the first output.
Pitis teaches:
An inference being performed using first output (model predictions), label information associated with the first output (predictions being treated (labeled) as correct), and the user request (prompt) as input. ((Pitis) Page 3, Algorithm 1; Page 4, Section 3, Subsection Test-time Boosting, line 5-23, “In this case, we substitute ground truth answer labels with model predictions, whereby predictions with “sufficient agreement” are treated as correct.…The algorithm is otherwise the same as the train-time algorithm”)
Pitis and Chen are analogous art because they are in the same area of invention, that using iterative and successive machine learning models.
Thus, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention, having the references in front of them, to have combined the teachings of Pitis of using a previous first model prediction output and label information associated with the first model prediction output as input in addition to a user request to perform a second inference operation with using a user request associated with the first output to perform a second inference operation taught by Chen. The motivation for this would be to adapt to the absence of ground truth answer labels. This combination of an output and label information associated with the output as additional inputs, as Pitis teaches, to the second inference operation using a user, as taught by Chen, would not change the functionality of either component and would yield the predictable result specified in claim 17 of the instant application.
Claim(s) 13, and 18 is/are rejected under 35 U.S.C. 103 as being unpatentable over Chen, in view of Packes, in further view of Zhu.
Regarding claim 13, Chen, in view of Packes, teaches the material disclosed in claim 9.
Chen, in view of Packes, does not teach:
wherein the first machine learning model is implemented at a user device and the second machine learning model is implemented at a server device.
Zhu teaches:
Implementing a first machine learning model at a user device (on-device model) and implementing a second machine learning model (server model). ((Zhu) Page 2, lines 20-40, “To further close the gap between the server-side giant LLMs and their smaller on-device counterparts, we propose a cascading approach to chain our on-device model with the more powerful server model. The system follows a simple yet effective principle: the server side will only be used when the on-device language model fails to provide a good response.”)
Zhu and Chen are analogous art because they are in the same area of invention: cascading machine learning models.
Thus, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention, having the references in front of them, to have combined the teachings of Zhu of implementing a smaller machine learning model on-device, while additionally implementing a more powerful model on a server, with the smaller machine learning model and more powerful model taught by Chen. The motivation for this would be to form a bridge between large server machine learning models that are too big to fit on-device and smaller on-device machine learning models so that the disadvantages of both smaller and too large models are mitigated. This application of a known technique of implementing and chaining models that are on-device and server side would yield the predictable result that is specified in claim 13 of the instant application.
Regarding claim 18, Chen, in view of Packes, teaches the material disclosed in claim 14, and Chen additionally teaches:
wherein the processing circuitry is configured to: generate (training) the scoring machine learning model as a discriminative regression model (simple regression model)... ((Chen) Page 5, Strategy 3 Paragraph 3, “The scoring function can be obtained by training a simple regression model that learns whether a generation is correct from the query and a generated answer.” Learning whether a generation is correct makes a regression model discriminative.)
Chen, in view of Packes, does not teach:
…using supervised learning.
Zhu teaches:
Training using supervised learning (supervised fine-tuning). ((Zhu) Page 2, lines 10-12, “we train our model using a combination of supervised fine-tuning (SFT) and reinforcement learning (RL).”
Zhu and Chen are analogous art because they are in the same area of invention, cascading machine learning models.
Thus, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention, having the references in front of them, to have combined the teachings of Zhu of training using a supervised fine-tuning learning method with a discriminative regression model scoring machine learning function as taught by Chen. The motivation for this would be to more effectively map outputs to appropriate scoring labels. This application of supervised learning as Zhu teaches to the discriminative regression model scoring function as taught by Chen would yield the predictable result that is specified in claim 18 of the instant application.
Claim(s) 15 is/are rejected under 35 U.S.C. 103 as being unpatentable over Chen, in view of Packes, in further view of LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative Fusion by Jiang et al., hereafter Jiang
Regarding claim 15, Chen, in view of Packes, teaches the material disclosed in claim 14, and Chen additionally teaches:
wherein the processing circuitry is configured to execute the instructions to:
determine, using the scoring machine learning model (generation scoring function), that the second output (GPT-3 answer/response) fails (“<”) to meet the threshold (“0.9”); ((Chen) Page 4, Figure 2e, the figure shows that if GPT-3 model fails to reach a score threshold of 0.9 it executes a GPT-4 model)
transmit (accept/return)[a] third output (GPT-4 answer) in response to the user request (query); ((Chen) Page 4, Figure 2e, the figure shows that GPT-4 may accept (transmit) an answer in response to the query.)
Chen, in view of Packes, does not teach:
wherein the processing circuitry is configured to execute the instructions to:
perform, using a last machine learning model of the model chain ((Chen) GPT-4), multiple third inference operations in parallel based on the user request, wherein each of the third inference operations produces different third output;
determine, using the scoring machine learning model, a third output having a highest score amongst the different third output;
Jiang teaches:
Multiple inference operations in parallel (N LLMs) based on a user request (input x) wherein each of the inference operations produces a different output (each output a candidate output); and ((Jiang) Page 3, Figure 2, Input: x is input to N LLMs that each output a candidate output.)
Determining, using a scoring machine learning model (PairRanker), an output having a highest score amongst the different output (Candidate at the top K=1). ((Jiang) Page 3, Figure 2, PairRanker ranks all candidates and takes the top K of them.)
Jiang and Chen are analogous art because they are in the same area of invention, that being deployment and scoring of LLM machine learning models.
Thus, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention, having the references in front of them, to have combined the teachings of Jiang of the multiple parallel inference operations and ranking function with the determination of the failure of the second machine learning model, performance of a third inference operation by a third machine learning model, and transmitting an output generated by the third machine learning model as taught by Chen. The motivation for this would be to achieve consistently superior performance by using the outputs of multiple LLMs. The simple substitution of the performance of multiple inference operations in parallel and determining which of the parallel inference operations has a highest score as Jiang teaches with the performance of a third inference operation as taught by Chen, would have been obvious to do for a person having ordinary skill in the art and would obtain the predictable result that is specified in claim 15 of the instant application.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Patents and/or related publications are cited in the Notice of References Cited (Form PTO-892) attached to this action to further show the state of the art with respect to cascading neural networks, ensemble neural networks, and machine learning model chains.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to DYLAN H LAI whose telephone number is (571)272-8628. The examiner can normally be reached Monday - Friday 7:30am-5:00pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Tamara Kyle can be reached at 5712524241. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
D. H. L.
Examiner
Art Unit 2144
/TAMARA T KYLE/ Supervisory Patent Examiner, Art Unit 2144