Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statements (IDSs) submitted on January 27, 2025, and April 30, 2025, are in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statements are being considered by the examiner.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 1-20 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre--AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
The term “computationally expensive” in claims 1, 4, 10-11, 15, and 20 is a relative term which renders the claims indefinite. The term “computationally expensive” is not defined by the claim, the specification does not provide a standard for ascertaining the requisite degree, and one of ordinary skill in the art would not be reasonably apprised of the scope of the invention. The term “expensive” renders the computation indefinite.
Claims 2-3, 5-9, 12-14, and 16-19 are further rejected under 35 U.S.C. 112b for dependence, either directly or indirectly, on claims 1 and 11.
Examiner’s Note: For the purposes of Examination, a “computationally expensive” will be interpreted as a number of parameters.
Claim Rejections – 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Regarding Claim 1,
Claim 1 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1 Analysis: Claim 1 is directed to a computer-implemented method for routing inputs to machine learning models, which is directed to a process, one of the statutory categories.
Step 2A Prong One Analysis: The limitations:
“computing one or more metric values based on an input”
“determining, for at least one trained machine learning model included in a plurality of trained machine learning models, a corresponding output quality degradation based on the one or more metric values, wherein the corresponding output quality degradation is relative to a most computationally expensive trained machine learning model included in the plurality of trained machine learning models”
“selecting a first trained machine learning model included in the plurality of trained machine learning models based on the corresponding output quality degradations”
As drafted, under their broadest reasonable interpretations, cover mental processes, i.e., concepts performed in the human mind (including an observation, evaluation, judgement, opinion). The above limitations in the context of this claim correspond to mental processes, e.g., evaluation and judgement with assistance of pen and paper.
Step 2A Prong Two Analysis: The judicial exceptions are not integrated into a practical application. In particular, the claim recited additional elements that are mere instructions to apply an exception (See MPEP 2106.05(f)) and insignificant extra-solution activity (See MPEP 2106.05(g)).
The limitations:
“A computer-implemented method for routing inputs to machine learning models for execution”
As drafted, are additional elements that amount to no more than mere instructions to apply an exception for the abstract ideas. See MPEP 2106.05(f).
The limitations:
“transmitting the input to the first trained machine learning model for execution”
As drafted, are additional elements that amount to no more than insignificant extra-solution activity. See MPEP 2106.05(g).
Therefore, the additional elements do not integrate the abstract ideas into a practical application.
Step 2B Analysis: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to the integration of the abstract ideas into a practical application, all of the additional elements are “mere instructions to apply” and “insignificant extra-solution activity”. Further, the transmitting limitation recites the well-understood, routine, and conventional activity of receiving or transmitting data over a network. MPEP 2106.05(d)(II); OIP Techs., Inc., v. Amazon.com, Inc., 788 F.3d 1359, 1363, 115 USPQ2d 1090, 1093 (Fed. Cir. 2015) (sending messages over a network). Mere instructions to apply an exception and insignificant extra-solution activity cannot provide an inventive concept. The claim is not patent eligible.
Regarding Claim 2,
Claim 2 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1 Analysis: Claim 2 is directed to a computer-implemented method for routing inputs to machine learning models, which is directed to a process, one of the statutory categories.
Step 2A Prong One Analysis: The limitations:
“for each trained machine learning model included in the at least one trained machine learning model: for each metric value included in the one or more metric values, determining a corresponding intermediate output quality degradation”
“selecting a largest output quality degradation from the corresponding intermediate output quality degradations as the corresponding output quality degradation”
As drafted, under their broadest reasonable interpretations, cover mental processes, i.e., concepts performed in the human mind (including an observation, evaluation, judgement, opinion). The above limitations in the context of this claim correspond to mental processes, e.g., evaluation and judgement with assistance of pen and paper.
Step 2A Prong Two Analysis: See corresponding analysis of claim 1.
Step 2B Analysis: See corresponding analysis of claim 1.
Regarding Claim 3,
Claim 3 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1 Analysis: Claim 3 is directed to a computer-implemented method for routing inputs to machine learning models, which is directed to a process, one of the statutory categories.
Step 2A Prong One Analysis: The limitations:
“determining a bucket included in a plurality of buckets to which the metric value belongs”
“determining the corresponding intermediate output quality degradation based on the bucket”
As drafted, under their broadest reasonable interpretations, cover mental processes, i.e., concepts performed in the human mind (including an observation, evaluation, judgement, opinion). The above limitations in the context of this claim correspond to mental processes, e.g., evaluation and judgement with assistance of pen and paper.
Step 2A Prong Two Analysis: See corresponding analysis of claim 2.
Step 2B Analysis: See corresponding analysis of claim 2.
Regarding Claim 4,
Claim 4 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1 Analysis: Claim 4 is directed to a computer-implemented method for routing inputs to machine learning models, which is directed to a process, one of the statutory categories.
Step 2A Prong One Analysis: See corresponding analysis of claim 1.
Step 2A Prong Two Analysis: The judicial exceptions are not integrated into a practical application. In particular, the claim recited additional elements that are additional details that do not apply the exception in a meaningful way (See MPEP 2106.05(e)).
The limitations:
“wherein the first trained machine learning model comprises a least computationally expensive trained machine learning model included in the plurality of machine learning models having a corresponding output quality degradation that satisfies a predefined condition”
As drafted, are additional elements that do not apply an exception for the abstract ideas in a meaningful way. See MPEP 2106.05(e).
Therefore, the additional elements do not integrate the abstract ideas into a practical application.
Step 2B Analysis: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to the integration of the abstract ideas into a practical application, all of the additional elements do not apply the exception in a meaningful way. The claim is not patent eligible.
Regarding Claim 5,
Claim 5 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1 Analysis: Claim 5 is directed to a computer-implemented method for routing inputs to machine learning models, which is directed to a process, one of the statutory categories.
Step 2A Prong One Analysis: See corresponding analysis of claim 1.
Step 2A Prong Two Analysis: The judicial exceptions are not integrated into a practical application. In particular, the claim recited additional elements that are additional details that do not apply the exception in a meaningful way (See MPEP 2106.05(e)).
The limitations:
“wherein the one or more metric values include a count of a number of words included in the input”
As drafted, are additional elements that do not apply an exception for the abstract ideas in a meaningful way. See MPEP 2106.05(e).
Therefore, the additional elements do not integrate the abstract ideas into a practical application.
Step 2B Analysis: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to the integration of the abstract ideas into a practical application, all of the additional elements do not apply the exception in a meaningful way. The claim is not patent eligible.
Regarding Claim 6,
Claim 6 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1 Analysis: Claim 6 is directed to a computer-implemented method for routing inputs to machine learning models, which is directed to a process, one of the statutory categories.
Step 2A Prong One Analysis: See corresponding analysis of claim 1.
Step 2A Prong Two Analysis: The judicial exceptions are not integrated into a practical application. In particular, the claim recited additional elements that are additional details that do not apply the exception in a meaningful way (See MPEP 2106.05(e)).
The limitations:
“wherein the one or more metric values include a count of a number of nouns included in the input”
As drafted, are additional elements that do not apply an exception for the abstract ideas in a meaningful way. See MPEP 2106.05(e).
Therefore, the additional elements do not integrate the abstract ideas into a practical application.
Step 2B Analysis: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to the integration of the abstract ideas into a practical application, all of the additional elements do not apply the exception in a meaningful way. The claim is not patent eligible.
Regarding Claim 7,
Claim 7 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1 Analysis: Claim 7 is directed to a computer-implemented method for routing inputs to machine learning models, which is directed to a process, one of the statutory categories.
Step 2A Prong One Analysis: See corresponding analysis of claim 1.
Step 2A Prong Two Analysis: The judicial exceptions are not integrated into a practical application. In particular, the claim recited additional elements that are additional details that do not apply the exception in a meaningful way (See MPEP 2106.05(e)).
The limitations:
“wherein the one or more metric values include a reading time duration associated with the input”
As drafted, are additional elements that do not apply an exception for the abstract ideas in a meaningful way. See MPEP 2106.05(e).
Therefore, the additional elements do not integrate the abstract ideas into a practical application.
Step 2B Analysis: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to the integration of the abstract ideas into a practical application, all of the additional elements do not apply the exception in a meaningful way. The claim is not patent eligible.
Regarding Claim 8,
Claim 8 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1 Analysis: Claim 8 is directed to a computer-implemented method for routing inputs to machine learning models, which is directed to a process, one of the statutory categories.
Step 2A Prong One Analysis: See corresponding analysis of claim 1.
Step 2A Prong Two Analysis: The judicial exceptions are not integrated into a practical application. In particular, the claim recited additional elements that are additional details that do not apply the exception in a meaningful way (See MPEP 2106.05(e)).
The limitations:
“wherein each trained machine learning model included in the plurality of trained machine learning models comprises a language model”
As drafted, are additional elements that do not apply an exception for the abstract ideas in a meaningful way. See MPEP 2106.05(e).
Therefore, the additional elements do not integrate the abstract ideas into a practical application.
Step 2B Analysis: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to the integration of the abstract ideas into a practical application, all of the additional elements do not apply the exception in a meaningful way. The claim is not patent eligible.
Regarding Claim 9,
Claim 9 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1 Analysis: Claim 9 is directed to a computer-implemented method for routing inputs to machine learning models, which is directed to a process, one of the statutory categories.
Step 2A Prong One Analysis: The limitations:
“computing a plurality of additional metric values based on a plurality of example inputs”
“determining, for each trained machine learning model included in the one or more trained machine learning models, an additional corresponding output quality degradation for each example input included in the plurality of example inputs”
As drafted, under their broadest reasonable interpretations, cover mental processes, i.e., concepts performed in the human mind (including an observation, evaluation, judgement, opinion). The above limitations in the context of this claim correspond to mental processes, e.g., evaluation and judgement with assistance of pen and paper.
Step 2A Prong Two Analysis: The judicial exceptions are not integrated into a practical application. In particular, the claim recited additional elements that are insignificant extra-solution activity (See MPEP 2106.05(g)).
The limitations:
“storing one or more associations between the plurality of additional metric values and the additional corresponding output quality degradations”
As drafted, are additional elements that amount to no more than insignificant extra-solution activity. See MPEP 2106.05(g).
Therefore, the additional elements do not integrate the abstract ideas into a practical application.
Step 2B Analysis: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to the integration of the abstract ideas into a practical application, all of the additional elements are “insignificant extra-solution activity”. Additionally, the storing limitation recites the well-understood, routine, and conventional activity of storing and retrieving information in memory. MPEP 2106.05(d)(II); Versata Dev. Group, Inc. v. SAP Am., Inc., 793 F.3d 1306, 1334, 115 USPQ2d 1681, 1701 (Fed. Cir. 2015). Insignificant extra-solution activity cannot provide an inventive concept. The claim is not patent eligible.
Regarding Claim 10,
Claim 10 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1 Analysis: Claim 10 is directed to a computer-implemented method for routing inputs to machine learning models, which is directed to a process, one of the statutory categories.
Step 2A Prong One Analysis: The limitations:
“computing a difference between an accuracy of the trained machine learning model for the example input and an accuracy of the most computationally expensive trained machine learning model included in the plurality of trained machine learning models for the example input”
As drafted, under their broadest reasonable interpretations, cover mental processes, i.e., concepts performed in the human mind (including an observation, evaluation, judgement, opinion). The above limitations in the context of this claim correspond to mental processes, e.g., evaluation and judgement with assistance of pen and paper.
Step 2A Prong Two Analysis: See corresponding analysis of claim 9.
Step 2B Analysis: See corresponding analysis of claim 9.
Regarding Claim 11,
Claim 11 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1 Analysis: Claim 11 is directed to one or more non-transitory computer-readable media storing instructions, which is directed to a machine, one of the statutory categories.
Step 2A Prong One Analysis: The limitations:
“computing one or more metric values based on an input”
“determining, for at least one trained machine learning model included in a plurality of trained machine learning models, a corresponding output quality degradation based on the one or more metric values, wherein the corresponding output quality degradation is relative to a most computationally expensive trained machine learning model included in the plurality of trained machine learning models”
“selecting a first trained machine learning model included in the plurality of trained machine learning models based on the corresponding output quality degradations”
As drafted, under their broadest reasonable interpretations, cover mental processes, i.e., concepts performed in the human mind (including an observation, evaluation, judgement, opinion). The above limitations in the context of this claim correspond to mental processes, e.g., evaluation and judgement with assistance of pen and paper.
Step 2A Prong Two Analysis: The judicial exceptions are not integrated into a practical application. In particular, the claim recited additional elements that are mere instructions to apply an exception (See MPEP 2106.05(f)) and insignificant extra-solution activity (See MPEP 2106.05(g)).
The limitations:
“One or more non-transitory computer-readable media storing instructions that, when executed by at least one processor, cause the at least one processor to perform steps”
As drafted, are additional elements that amount to no more than mere instructions to apply an exception for the abstract ideas. See MPEP 2106.05(f).
The limitations:
“transmitting the input to the first trained machine learning model for execution”
As drafted, are additional elements that amount to no more than insignificant extra-solution activity. See MPEP 2106.05(g).
Therefore, the additional elements do not integrate the abstract ideas into a practical application.
Step 2B Analysis: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to the integration of the abstract ideas into a practical application, all of the additional elements are “mere instructions to apply” and “insignificant extra-solution activity”. Further, the transmitting limitation recites the well-understood, routine, and conventional activity of receiving or transmitting data over a network. MPEP 2106.05(d)(II); OIP Techs., Inc., v. Amazon.com, Inc., 788 F.3d 1359, 1363, 115 USPQ2d 1090, 1093 (Fed. Cir. 2015) (sending messages over a network). Mere instructions to apply an exception and insignificant extra-solution activity cannot provide an inventive concept. The claim is not patent eligible.
Regarding Claim 12,
Claim 12 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1 Analysis: Claim 12 is directed to one or more non-transitory computer-readable media storing instructions, which is directed to a machine, one of the statutory categories.
Step 2A Prong One Analysis: The limitations:
“for each trained machine learning model included in the one or more trained machine learning models: for each metric value included in the one or more metric values, determining a corresponding intermediate output quality degradation”
“selecting a largest output quality degradation from the corresponding intermediate output quality degradations as the corresponding output quality degradation”
As drafted, under their broadest reasonable interpretations, cover mental processes, i.e., concepts performed in the human mind (including an observation, evaluation, judgement, opinion). The above limitations in the context of this claim correspond to mental processes, e.g., evaluation and judgement with assistance of pen and paper.
Step 2A Prong Two Analysis: See corresponding analysis of claim 11.
Step 2B Analysis: See corresponding analysis of claim 11.
Regarding Claim 13,
Claim 13 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1 Analysis: Claim 13 is directed to one or more non-transitory computer-readable media storing instructions, which is directed to a machine, one of the statutory categories.
Step 2A Prong One Analysis: The limitations:
“determining a bucket included in a plurality of buckets to which the metric value belongs”
“determining the corresponding intermediate output quality degradation based on the bucket”
As drafted, under their broadest reasonable interpretations, cover mental processes, i.e., concepts performed in the human mind (including an observation, evaluation, judgement, opinion). The above limitations in the context of this claim correspond to mental processes, e.g., evaluation and judgement with assistance of pen and paper.
Step 2A Prong Two Analysis: See corresponding analysis of claim 12.
Step 2B Analysis: See corresponding analysis of claim 12.
Regarding Claim 14,
Claim 14 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1 Analysis: Claim 14 is directed to one or more non-transitory computer-readable media storing instructions, which is directed to a machine, one of the statutory categories.
Step 2A Prong One Analysis: See corresponding analysis of claim 13.
Step 2A Prong Two Analysis: The judicial exceptions are not integrated into a practical application. In particular, the claim recited additional elements that are additional details that do not apply the exception in a meaningful way (See MPEP 2106.05(e)).
The limitations:
“wherein the plurality of buckets include a plurality of quantile buckets”
As drafted, are additional elements that do not apply an exception for the abstract ideas in a meaningful way. See MPEP 2106.05(e).
Therefore, the additional elements do not integrate the abstract ideas into a practical application.
Step 2B Analysis: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to the integration of the abstract ideas into a practical application, all of the additional elements do not apply the exception in a meaningful way. The claim is not patent eligible.
Regarding Claim 15,
Claim 15 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1 Analysis: Claim 15 is directed to one or more non-transitory computer-readable media storing instructions, which is directed to a machine, one of the statutory categories.
Step 2A Prong One Analysis: See corresponding analysis of claim 11.
Step 2A Prong Two Analysis: The judicial exceptions are not integrated into a practical application. In particular, the claim recited additional elements that are additional details that do not apply the exception in a meaningful way (See MPEP 2106.05(e)).
The limitations:
“wherein the first trained machine learning model comprises a least computationally expensive trained machine learning model included in the plurality of machine learning models having a corresponding output quality degradation that satisfies a predefined condition”
As drafted, are additional elements that do not apply an exception for the abstract ideas in a meaningful way. See MPEP 2106.05(e).
Therefore, the additional elements do not integrate the abstract ideas into a practical application.
Step 2B Analysis: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to the integration of the abstract ideas into a practical application, all of the additional elements do not apply the exception in a meaningful way. The claim is not patent eligible.
Regarding Claim 16,
Claim 16 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1 Analysis: Claim 16 is directed to one or more non-transitory computer-readable media storing instructions, which is directed to a machine, one of the statutory categories.
Step 2A Prong One Analysis: See corresponding analysis of claim 11.
Step 2A Prong Two Analysis: The judicial exceptions are not integrated into a practical application. In particular, the claim recited additional elements that are additional details that do not apply the exception in a meaningful way (See MPEP 2106.05(e)).
The limitations:
“wherein the one or more metric values include at least one of a count of a number of words included in the input, a count of a number of nouns included in the input, or a reading time duration associated with the input”
As drafted, are additional elements that do not apply an exception for the abstract ideas in a meaningful way. See MPEP 2106.05(e).
Therefore, the additional elements do not integrate the abstract ideas into a practical application.
Step 2B Analysis: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to the integration of the abstract ideas into a practical application, all of the additional elements do not apply the exception in a meaningful way. The claim is not patent eligible.
Regarding Claim 17,
Claim 17 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1 Analysis: Claim 17 is directed to one or more non-transitory computer-readable media storing instructions, which is directed to a machine, one of the statutory categories.
Step 2A Prong One Analysis: See corresponding analysis of claim 11.
Step 2A Prong Two Analysis: The judicial exceptions are not integrated into a practical application. In particular, the claim recited additional elements that are additional details that do not apply the exception in a meaningful way (See MPEP 2106.05(e)).
The limitations:
“wherein each trained machine learning model included in the plurality of trained machine learning models comprises a language model”
As drafted, are additional elements that do not apply an exception for the abstract ideas in a meaningful way. See MPEP 2106.05(e).
Therefore, the additional elements do not integrate the abstract ideas into a practical application.
Step 2B Analysis: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to the integration of the abstract ideas into a practical application, all of the additional elements do not apply the exception in a meaningful way. The claim is not patent eligible.
Regarding Claim 18,
Claim 18 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1 Analysis: Claim 18 is directed to one or more non-transitory computer-readable media storing instructions, which is directed to a machine, one of the statutory categories.
Step 2A Prong One Analysis: The limitations:
“computing a plurality of additional metric values based on a plurality of example inputs”
“determining, for each trained machine learning model included in the one or more trained machine learning models, an additional corresponding output quality degradation for each example input included in the plurality of example inputs”
As drafted, under their broadest reasonable interpretations, cover mental processes, i.e., concepts performed in the human mind (including an observation, evaluation, judgement, opinion). The above limitations in the context of this claim correspond to mental processes, e.g., evaluation and judgement with assistance of pen and paper.
Step 2A Prong Two Analysis: The judicial exceptions are not integrated into a practical application. In particular, the claim recited additional elements that are insignificant extra-solution activity (See MPEP 2106.05(g)).
The limitations:
“storing one or more associations between the plurality of additional metric values and the additional corresponding output quality degradations”
As drafted, are additional elements that amount to no more than insignificant extra-solution activity. See MPEP 2106.05(g).
Therefore, the additional elements do not integrate the abstract ideas into a practical application.
Step 2B Analysis: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to the integration of the abstract ideas into a practical application, all of the additional elements are “insignificant extra-solution activity”. Additionally, the storing limitation recites the well-understood, routine, and conventional activity of storing and retrieving information in memory. MPEP 2106.05(d)(II); Versata Dev. Group, Inc. v. SAP Am., Inc., 793 F.3d 1306, 1334, 115 USPQ2d 1681, 1701 (Fed. Cir. 2015). Insignificant extra-solution activity cannot provide an inventive concept. The claim is not patent eligible.
Regarding Claim 19,
Claim 19 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1 Analysis: Claim 19 is directed to one or more non-transitory computer-readable media storing instructions, which is directed to a machine, one of the statutory categories.
Step 2A Prong One Analysis: The limitations:
“determining, for each additional metric value included in the plurality of additional metric values, a bucket included in a plurality of buckets that the additional metric value belongs to”
As drafted, under their broadest reasonable interpretations, cover mental processes, i.e., concepts performed in the human mind (including an observation, evaluation, judgement, opinion). The above limitations in the context of this claim correspond to mental processes, e.g., evaluation and judgement with assistance of pen and paper.
Step 2A Prong Two Analysis: The judicial exceptions are not integrated into a practical application. In particular, the claim recited additional elements that are insignificant extra-solution activity (See MPEP 2106.05(g)).
The limitations:
“storing an association between each bucket included in the plurality of buckets and an average of the additional corresponding output quality degradations that are determined for example inputs whose additional metric values belong to the bucket”
As drafted, are additional elements that amount to no more than insignificant extra-solution activity. See MPEP 2106.05(g).
Therefore, the additional elements do not integrate the abstract ideas into a practical application.
Step 2B Analysis: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to the integration of the abstract ideas into a practical application, all of the additional elements are “insignificant extra-solution activity”. Additionally, the storing limitation recites the well-understood, routine, and conventional activity of storing and retrieving information in memory. MPEP 2106.05(d)(II); Versata Dev. Group, Inc. v. SAP Am., Inc., 793 F.3d 1306, 1334, 115 USPQ2d 1681, 1701 (Fed. Cir. 2015). Insignificant extra-solution activity cannot provide an inventive concept. The claim is not patent eligible.
Regarding Claim 20,
Claim 20 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1 Analysis: Claim 20 is directed to a system, comprising: one or more memories storing instructions; and one or more processors coupled to the one or more memories, which is directed to a machine, one of the statutory categories.
Step 2A Prong One Analysis: The limitations:
“computing one or more metric values based on an input”
“determining, for at least one trained machine learning models included in a plurality of trained machine learning models, a corresponding output quality degradation based on the one or more metric values, wherein the corresponding output quality degradation is relative to a most computationally expensive trained machine learning model included in the plurality of trained machine learning models”
“selecting a first trained machine learning model included in the plurality of trained machine learning models based on the corresponding output quality degradations”
As drafted, under their broadest reasonable interpretations, cover mental processes, i.e., concepts performed in the human mind (including an observation, evaluation, judgement, opinion). The above limitations in the context of this claim correspond to mental processes, e.g., evaluation and judgement with assistance of pen and paper.
Step 2A Prong Two Analysis: The judicial exceptions are not integrated into a practical application. In particular, the claim recited additional elements that are mere instructions to apply an exception (See MPEP 2106.05(f)) and insignificant extra-solution activity (See MPEP 2106.05(g)).
The limitations:
“A system, comprising: one or more memories storing instructions; and one or more processors coupled to the one or more memories that, when executing the instructions, perform the steps of:”
As drafted, are additional elements that amount to no more than mere instructions to apply an exception for the abstract ideas. See MPEP 2106.05(f).
The limitations:
“transmitting the input to the first trained machine learning model for execution”
As drafted, are additional elements that amount to no more than insignificant extra-solution activity. See MPEP 2106.05(g).
Therefore, the additional elements do not integrate the abstract ideas into a practical application.
Step 2B Analysis: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to the integration of the abstract ideas into a practical application, all of the additional elements are “mere instructions to apply” and “insignificant extra-solution activity”. Further, the transmitting limitation recites the well-understood, routine, and conventional activity of receiving or transmitting data over a network. MPEP 2106.05(d)(II); OIP Techs., Inc., v. Amazon.com, Inc., 788 F.3d 1359, 1363, 115 USPQ2d 1090, 1093 (Fed. Cir. 2015) (sending messages over a network). Mere instructions to apply an exception and insignificant extra-solution activity cannot provide an inventive concept. The claim is not patent eligible.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 1, 4, 8-11, 15, 17-18, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Shnitzer et al. (Large Language Model Routing with Benchmark Datasets) (“Shnitzer”) in view of Chen et al. (FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance) (“Chen”).
Regarding claim 1, Shnitzer teaches a computer-implemented method for routing inputs to machine learning models for execution (Shnitzer Section 1 Introduction “Our contributions are summarized below: We formalize the problem of learning the strengths and weaknesses of LLMs for downstream routing, i.e., selecting the best model, as a collection of binary classification problems. The goal of each classification problem is to predict whether a given LLM will be “correct” on an input.” Shnitzer teaches a method for routing inputs to Large Language Models, corresponding to routing inputs to machine learning models for execution.), the method comprising: computing one or more metric values based on an input (Shnitzer Section 3 Learning from Benchmarks “We start by introducing notation to describe the majority of NLP benchmarks. Let {xd 1,...,xd nd }D d=1 be a collection of inputs across D tasks. Each input text xd i corresponds to a reference answer rd i, i.e., an ideal generation for the corresponding input. Finally, there is a metric Fd(x,o,r) that can be task-dependent and measures how well a response o for an input x corresponds to the reference r.” Shnitzer teaches computing a metric Fd(x,o,r) that measures how well a response o for an input x corresponds to the reference r, corresponding to a computed metric value based on an input.). determining, for at least one trained machine learning model included in a plurality of trained machine learning models, a corresponding output quality degradation based on the one or more metric values (Shnitzer Section 3 Learning from Benchmarks “Our goal is to learn a simple routing function gm(x) for each LLM, m=1,...,M, that can predict {fd′ im}nd′ i=1, i.e., the performance of the corresponding LLM on a new task d′. Then it is trivial to select the best LLM for this task. For efficiency at test time, we restrict the routers {gm}M m=1 to only depend on the input x… To complete the problem formulation, we denote the “correctness” of model m on an input x by y(x,m) ∈ {0,1}. Correctness is evaluated as follows: generate a response with LLM m on input xd i, compare it to the corresponding reference rd i, and output 1 if the model’s response is good enough, i.e., fd im > ηd, and 0 otherwise, where ηd is some threshold that can be task and/or metric specific.” Shnitzer teaches predicting a correctness of an LLM on a given input by comparing metric values to a threshold, wherein predicting correctness of a given LLM on a given input corresponds to determining an output quality degradation, which is achieved by outputting a 1 if the models response is good enough, and 0 otherwise.), wherein the corresponding output quality degradation is relative to a most computationally expensive trained machine learning model included in the plurality of trained machine learning models (Shnitzer Section 5.1 “Models We evaluate 18 open-source models ranging in size from 3B to 70B, including base and chat variations of Llama 2 in different sizes. All models are summarized in Table 4.” Shnitzer teaches a 70B parameter model, as shown in Table 4, corresponding to a most computationally expensive trained machine learning model.); selecting a first trained machine learning model included in the plurality of trained machine learning models based on the corresponding output quality degradations (Shnitzer Section 3 Learning from Benchmarks “Our goal is to learn a simple routing function gm(x) for each LLM, m=1,...,M, that can predict {fd′ im}nd′ i=1, i.e., the performance of the corresponding LLM on a new task d′. Then it is trivial to select the best LLM for this task… To complete the problem formulation, we denote the “correctness” of model m on an input x by y(x,m) ∈ {0,1}. Correctness is evaluated as follows: generate a response od im with LLM m on input xd i, compare it to the corresponding reference rd i, and output 1 if the model’s response is good enough, i.e., fd im > ηd, and 0 otherwise, where ηd is some threshold that can be task and/or metric specific.”; Section 5 Experiments “In Table1 we report averages across experiments for the performance of the selected model (Acc.)” Shnitzer teaches selecting a first LLM when the calculated correctness equals 1 (i.e., the models response is good enough for the input), wherein the results of the selection are shown in Table 1, corresponding to selecting a first model based on the output quality degradations, wherein correctness corresponds to the output quality degradations.);
Shnitzer fails to explicitly teach transmitting the input to the first trained machine learning model for execution.
However, Chen teaches transmitting the input to the first trained machine learning model for execution (Chen Section 3; Strategy 3: LLM cascade “Consequently, appropriately selecting which LLMs to use can provide both cost reduction and performance improvements. LLM cascade, as illustrated in Figure 2 (e), is one such example. LLM cascade sends a query to a list of LLM APIs sequentially. If one LLM API’s response is reliable, then its response is returned, and no further LLMs in the list are needed.” Chen teaches sending a query to a first LLM and returning the response if its reliable, corresponding to transmitting the input to a first machine learning model for execution.).
Shnitzer and Chen are both considered to be analogous to the claimed invention because they are in the same field of artificial intelligence and more specifically applied to LLM routing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Shnitzer with the above teachings of Chen. Doing so may reduce query cost (Chen Section 3; Strategy 3: LLM cascade “Query cost is significantly reduced if the first few APIs are relatively inexpensive and produce reliable generations.”).
Regarding claim 4, Shnitzer in view of Chen teaches wherein the first trained machine learning model comprises a least computationally expensive trained machine learning model included in the plurality of machine learning models (Shnitzer Section 5.1 “Models We evaluate 18 open-source models ranging in size from 3B to 70B, including base and chat variations of Llama 2 in different sizes. All models are summarized in Table 4.” Shnitzer teaches a 3B parameter model, as shown in Table 4, corresponding to a least computationally expensive trained machine learning model.) having a corresponding output quality degradation that satisfies a predefined condition (Shnitzer Section 3 Learning from Benchmarks “To complete the problem formulation, we denote the “correctness” of model m on an input x by y(x,m) ∈ {0,1}. Correctness is evaluated as follows: generate a response od im with LLM m on input xd i, compare it to the corresponding reference rd i, and output 1 if the model’s response is good enough, i.e., fd im > ηd, and 0 otherwise, where ηd is some threshold that can be task and/or metric specific.” Shnitzer teaches a threshold that can be task and/or metric specific for model output quality, corresponding to a predefined condition.).
It would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Shnitzer in view of Chen for the same reasons disclosed above in the rejection of claim 1.
Regarding claim 8, Shnitzer in view of Chen teaches wherein each trained machine learning model included in the plurality of trained machine learning models comprises a language model (Shnitzer Section 2 “In this paper, instead, we use data from benchmarks to learn the strengths and weaknesses of LLMs across tasks and domains. The resulting model router requires generating outputs only with the chosen LLM at test time.” Shnitzer teaches Large Language Models (LLMs) corresponding to a language model.).
It would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Shnitzer in view of Chen for the same reasons disclosed above in the rejection of claim 1.
Regarding claim 9, Shnitzer in view of Chen teaches further comprising: computing a plurality of additional metric values based on a plurality of example inputs (Shnitzer Section 3 Learning from Benchmarks “In Section 5.2, we also present results with raw metrics instead of correctness.”; Section 5.2 “The dataset is composed of instruction following tasks, divided into train/validation/test sets of100K/5K/5Ksamples, and includes evaluations of N=11 open-source LLMs using common metrics, e.g. BERT Score (Zhang et al., 2020), BART Score (Yuan et al., 2021), and BLEURT (Sellam et al., 2020). In Jiang et al. (2023), this benchmark was used to compare different LLM ranking methods in per-instance model selection. We follow the same setting and apply our score S1 (m,d′) to the test set, per instance, where we use the 100K-sample train set as the benchmark data for training our LLM router.” Shnitzer teaches predicting model performance by training the router and instead using raw/common metrics provided from an additional dataset, corresponding to additional metric values.); determining, for each trained machine learning model included in the one or more trained machine learning models, an additional corresponding output quality degradation for each example input included in the plurality of example inputs (Shnitzer Section 5.2 “The dataset is composed of instruction following tasks, divided into train/validation/test sets of100K/5K/5Ksamples, and includes evaluations of N=11 open-source LLMs using common metrics, e.g. BERT Score (Zhang et al., 2020), BART Score (Yuan et al., 2021), and BLEURT (Sellam et al., 2020). In Jiang et al. (2023), this benchmark was used to compare different LLM ranking methods in per-instance model selection. We follow the same setting and apply our score S1 (m,d′) to the test set, per instance, where we use the 100K-sample train set as the benchmark data for training our LLM router. See Appendices A and C for details on the score computation and the experiment parameters, respectively.”; Appendix A “Finally, we compute the per-model correctness predictors, S1(m,d′) and S2(m,d′), for the new task d′, according to equation 3 and equation 4, respectively.” Shnitzer teaches predicting per-model correctness based on the additional metrics, corresponding to the additional corresponding output quality degradations.); and storing one or more associations between the plurality of additional metric values and the additional corresponding output quality degradations (Shnitzer Section 5.2 “Additionally, we present the metrics for the best models on average (BMA), Open-Assistant (LAION-AI,2023) and Vicuna (Chiang et al.,2023). We report the results of BERT Score, BART Score and BLEURT in Table 2, along with the number of model calls per instance (MCPI) performed during inference time.” Shnitzer teaches Table 2 for storing the associations between additional metrics and model performance.).
It would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Shnitzer in view of Chen for the same reasons disclosed above in the rejection of claim 1.
Regarding claim 10, Shnitzer in view of Chen teaches wherein determining, for each trained machine learning model included in the one or more trained machine learning models, an additional corresponding output quality degradation for each example input (Shnitzer Section 5.2 “The dataset is composed of instruction following tasks, divided into train/validation/test sets f100K/5K/5K samples ,and includes evaluations of N=11 open-source LLMs using common metrics, e.g. BERT Score (Zhang et al., 2020), BART Score (Yuan et al., 2021), and BLEURT (Sellam et al., 2020). In Jiang et al. (2023), this benchmark was used to compare different LLM ranking methods in per-instance model selection. We follow the same setting and apply our score S1 (m,d′) to the test set, per instance, where we use the 100K-sample train set as the benchmark data for training our LLM router. See Appendices A and C for details on the score computation and the experiment parameters, respectively.”; Appendix A “Finally, we compute the per-model correctness predictors, S1(m,d′) and S2(m,d′), for the new task d′, according to equation 3 and equation 4, respectively.” Shnitzer teaches predicting per-model correctness based on the additional/raw metrics, corresponding to the additional corresponding output quality degradations.) comprises computing a difference between an accuracy of the trained machine learning model for the example input and an accuracy of the most computationally expensive trained machine learning model included in the plurality of trained machine learning models for the example input (Shnitzer Table 4 “Average Accuracy on the 29 HELM tasks”; Section 5.1 Model routing on Helm Models” We evaluate 18 open-source models ranging in size from 3B to 70B, including base and chat variations of Llama2 in different sizes. All models are summarized in Table 4” Shnitzer teaches comparing different model accuracies by size of model, as shown in Table 4, corresponding to computing a difference between an accuracy a trained machine learning model and the accuracy of a most computationally expensive machine learning model (i.e., the model with a 70 billion parameter size).).
It would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Shnitzer in view of Chen for the same reasons disclosed above in the rejection of claim 9.
Regarding claim 11, it is the one or more non-transitory computer-readable media storing instructions embodiment of claim 1 with similar limitations to claim 1 and is rejected using the same reasoning found above in the rejection of claim 1. Further, Chen teaches one or more non-transitory computer-readable media storing instructions that, when executed by at least one processor, cause the at least one processor to perform steps (Chen Section 3 “The concept of LLM approximation is quite simple: if an LLM API is too costly to utilize, one can approximate it using more affordable models or infrastructures. One example is the completion cache: as depicted in Figure 2 (c), the fundamental idea involves storing the response locally in a cache (e.g., a database) when submitting a query to an LLM API.”).
Shnitzer and Chen are both considered to be analogous to the claimed invention because they are in the same field of artificial intelligence and more specifically applied to LLM routing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Shnitzer with the above teachings of Chen. Doing so may reduce query cost (Chen Section 3; Strategy 3: LLM cascade “Query cost is significantly reduced if the first few APIs are relatively inexpensive and produce reliable generations.”).
Regarding claim 15, the rejection of claim 11 is incorporated herein. Further, the limitations in this claim are taught by Shnitzer in view of Chen for the same reasons disclosed above in the rejection of claim 4.
Regarding claim 17, the rejection of claim 11 is incorporated herein. Further, the limitations in this claim are taught by Shnitzer in view of Chen for the same reasons disclosed above in the rejection of claim 8.
Regarding claim 18, the rejection of claim 11 is incorporated herein. Further, the limitations in this claim are taught by Shnitzer in view of Chen for the same reasons disclosed above in the rejection of claim 9.
Regarding claim 20, it is the system embodiment of claim 1 with similar limitations to claim 1 and is rejected using the same reasoning found above in the rejection of claim 1. Further, Chen teaches a system, comprising: one or more memories storing instructions; and one or more processors coupled to the one or more memories (Chen Section 3 “The concept of LLM approximation is quite simple: if an LLM API is too costly to utilize, one can approximate it using more affordable models or infrastructures. One example is the completion cache: as depicted in Figure 2 (c), the fundamental idea involves storing the response locally in a cache (e.g., a database) when submitting a query to an LLM API.”).
Shnitzer and Chen are both considered to be analogous to the claimed invention because they are in the same field of artificial intelligence and more specifically applied to LLM routing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Shnitzer with the above teachings of Chen. Doing so may reduce query cost (Chen Section 3; Strategy 3: LLM cascade “Query cost is significantly reduced if the first few APIs are relatively inexpensive and produce reliable generations.”).
Claims 2-3, 12-14 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Shnitzer et al. (Large Language Model Routing with Benchmark Datasets) (“Shnitzer”) in view of Chen et al. (FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance) (“Chen”) in further view of Šakota et al. (Fly-Swat or Cannon? Cost-Effective Language Model Choice via Meta-Modeling) (“Šakota”).
Regarding claim 2, Shnitzer in view of Chen teaches the computer-implemented method of claim 1, as discussed above in the rejection of claim 1, but fails to explicitly teach wherein determining, for the at least one trained machine learning model included in the plurality of machine learning models, a corresponding output quality degradation comprises, for each trained machine learning model included in the at least one trained machine learning model: for each metric value included in the one or more metric values, determining a corresponding intermediate output quality degradation; and selecting a largest output quality degradation from the corresponding intermediate output quality degradations as the corresponding output quality degradation.
However, Šakota teaches wherein determining, for the at least one trained machine learning model included in the plurality of machine learning models, a corresponding output quality degradation (Šakota Section 3.2 Framework setting “Then, using meta-model, we predict the performance 𝑝𝑖𝑗 of each LM 𝑙𝑖 when running the query𝑞𝑗.” Šakota teaches determining/predicting model output performance for each of a plurality of language models (i.e., machine learning models), corresponding to determining an output quality degradation.) comprises, for each trained machine learning model included in the at least one trained machine learning model (Šakota Section 3.2 Framework setting “Then, using meta-model, we predict the performance 𝑝𝑖𝑗 of each LM 𝑙𝑖 when running the query𝑞𝑗.” Šakota teaching the following method for each language model (i.e., each trained machine learning model).): for each metric value included in the one or more metric values, determining a corresponding intermediate output quality degradation (Šakota Section 3.2 Framework setting; Meta-model and cost estimation “In order to know which LM to use for a certain query, we first have to be able to predict the performance metric 𝑝𝑖𝑗 an LM 𝑙𝑖 achieves when solving this query 𝑞𝑗. To do that, we train a meta-model. In our case, the meta-model is a binary classifier. During the training, we send a query𝑞𝑗 to which we append token representing LM 𝑙𝑖 as input to the meta-model, while the targets are 1 or 0, depending on whether LM𝑙𝑖 solves the query 𝑞𝑗 or not… Along with the measure of performance, we have to estimate the cost 𝑐𝑖𝑗 of running the query 𝑞𝑗 with each model 𝑙𝑖.” Šakota teaches predicting performance and cost for each query using a meta-model, corresponding to determining a corresponding output quality degradation for each metric (query).); and selecting a largest output quality degradation from the corresponding intermediate output quality degradations as the corresponding output quality degradation (Šakota Section 3.2 Framework setting; Assignment strategies “Performance-maximizing strategy: This strategy is based on the outputs of the meta-model. For each sample, we choose the LM that, according to the meta-model, is predicted to achieve the highest performance.” Šakota teaches selecting the LM with the highest predicted performance, corresponding to selecting a largest output quality degradation.).
Shnitzer, Chen, and Šakota are all considered to be analogous to the claimed invention because they are in the same field of artificial intelligence and more specifically applied to LLM routing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Shnitzer in view of Chen with the above teachings of Šakota. Doing so may reduce costs while maintaining the same performance as a biggest LLM (Šakota Section 6 Conclusion “By employing CELMOC to the union of 14 datasets, we are able to reduce costs by 62.79% while maintaining the same performance as the biggest LM in our pool.”).
Regarding claim 3, Shnitzer in view of Chen in further view of Šakota teaches wherein for each metric value included in the one or more metric values, determining a corresponding intermediate output quality degradation comprises: determining a bucket included in a plurality of buckets to which the metric value belongs (Šakota Section 3.2 Framework setting “Thresholding strategy: This strategy is also based on the outputs of the meta-model. The user has to specify an acceptable performance threshold that defines whether a task is solved or not. Outputs are binarized according to that threshold. A concrete example where this strategy might be useful are tasks that are evaluated with binary metrics such as accuracy.” Šakota teaches thresholding corresponds to buckets.); and determining the corresponding intermediate output quality degradation based on the bucket (Šakota Section 3.2 Framework setting “Thresholding strategy: This strategy is also based on the outputs of the meta-model. The user has to specify an acceptable performance threshold that defines whether a task is solved or not. Outputs are binarized according to that threshold. A concrete example where this strategy might be useful are tasks that are evaluated with binary metrics such as accuracy.” Šakota teaches binarized outputs based on the performance threshold that defines whether a task is solved or not, corresponding to determining the output quality degradation based on the bucket.).
Shnitzer, Chen, and Šakota are all considered to be analogous to the claimed invention because they are in the same field of artificial intelligence and more specifically applied to LLM routing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Shnitzer in view of Chen in further view of Šakota with the above teachings of Šakota. Doing so may reduce costs while maintaining the same performance as a biggest LLM (Šakota Section 6 Conclusion “By employing CELMOC to the union of 14 datasets, we are able to reduce costs by 62.79% while maintaining the same performance as the biggest LM in our pool.”).
Regarding claim 12, the rejection of claim 11 is incorporated herein. Further, the limitations in this claim are taught by Shnitzer in view of Chen in further view of Šakota for the same reasons disclosed above in the rejection of claim 2.
Regarding claim 13, the rejection of claim 12 is incorporated herein. Further, the limitations in this claim are taught by Shnitzer in view of Chen in further view of Šakota for the same reasons disclosed above in the rejection of claim 3.
Regarding claim 14, Chen in further view of Šakota teaches wherein the plurality of buckets include a plurality of quantile buckets (Šakota Figure 3 “Figure 3: Cost-accuracy plot. Accuracy and average cost per query (in US$) achieved by assigning every query from the set to an LM. The plot shows results obtained using assignment strategies from Sec. 3.2… Two thresholding strategies for cases when none of the LMs solve the data sample are marked by a letter under them: choosing (a) the biggest and (b) the smallest LM. Error bars are 95% confidence intervals.” Šakota teaches quantiles on the x-axis of Figure 3, which are quantile because the measurements are split into equal sized groups, wherein Figure 3 includes a plot of the thresholding corresponding to the buckets.).
Shnitzer, Chen, and Šakota are all considered to be analogous to the claimed invention because they are in the same field of artificial intelligence and more specifically applied to LLM routing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Shnitzer in view of Chen in further view of Šakota with the above teachings of Šakota. Doing so may reduce costs while maintaining the same performance as a biggest LLM (Šakota Section 6 Conclusion “By employing CELMOC to the union of 14 datasets, we are able to reduce costs by 62.79% while maintaining the same performance as the biggest LM in our pool.”).
Regarding claim 19, Shnitzer in view of Chen teaches the one or more non-transitory computer-readable media of claim 18, as discussed above in the rejection of claim 18, but fails to explicitly teach wherein determining the one or more associations comprises: determining, for each additional metric value included in the plurality of additional metric values, a bucket included in a plurality of buckets that the additional metric value belongs to; and storing an association between each bucket included in the plurality of buckets and an average of the additional corresponding output quality degradations that are determined for example inputs whose additional metric values belong to the bucket.
However, Šakota teaches wherein determining the one or more associations comprises: determining, for each additional metric value included in the plurality of additional metric values, a bucket included in a plurality of buckets that the additional metric value belongs to (Šakota Section 3.2 Framework setting “Thresholding strategy: This strategy is also based on the outputs of the meta-model. The user has to specify an acceptable performance threshold that defines whether a task is solved or not. Outputs are binarized according to that threshold. A concrete example where this strategy might be useful are tasks that are evaluated with binary metrics such as accuracy.” Šakota teaches thresholding corresponding to buckets for metrics.); and storing an association between each bucket included in the plurality of buckets and an average of the additional corresponding output quality degradations that are determined for example inputs whose additional metric values belong to the bucket (Šakota Figure 3 “Figure 3: Cost-accuracy plot. Accuracy and average cost per query (in US$) achieved by assigning every query from the query set to an LM from the LM pool. The plot shows results obtained using assignment strategies from Sec. 3.2. Single model strategies for each LMs are marked by the number under them: (1) text-ada-001 (2) text-babbage-001 (3) text curie-001 (4) text-davinci-002. Two thresholding strategies for cases when none of the LMs solve the data sample are marked by a letter under them: choosing (a) the biggest and (b) the smallest LM. Error bars are 95% confidence intervals” Šakota teaches storing the results of the thresholding corresponding to the association between buckets for the metrics in figure 3 including an average output cost, corresponding to storing the associations between each bucket and an average, wherein the average is shown on the x-axis of the figure.).
Shnitzer, Chen, and Šakota are all considered to be analogous to the claimed invention because they are in the same field of artificial intelligence and more specifically applied to LLM routing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Shnitzer in view of Chen with the above teachings of Šakota. Doing so may reduce costs while maintaining the same performance as a biggest LLM (Šakota Section 6 Conclusion “By employing CELMOC to the union of 14 datasets, we are able to reduce costs by 62.79% while maintaining the same performance as the biggest LM in our pool.”).
Claims 5-6 and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Shnitzer et al. (Large Language Model Routing with Benchmark Datasets) (“Shnitzer”) in view of Chen et al. (FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance) (“Chen”) in further view of Franco et al. (U.S. Patent Publication No. 2025/0045930) (“Franco”).
Regarding claim 5, Shnitzer in view of Chen teaches the computer-implemented method of claim 1 as discussed above in the rejection of claim 1, but fails to teach wherein the one or more metric values include a count of a number of words included in the input.
However, Franco teaches wherein the one or more metric values include a count of a number of words included in the input (Franco [0058] “Whether the prompt is one provided to the description generator 160 or obtained by the description generator 160, e.g., from the captioning model 161, the prompt includes one or more words. The words can also be referred to as tokens.” [0063] “The cross-attention tensors can be represented as: Equation (7) where Q represents the number of tokenized words in the prompt.”; [0073] “Pixel accuracy (ACC) and mean intersection over union (mIoU) are used to measure segmentation performance of the example implementation.” Franco teaches counting a number of words in an input prompt and measuring the output quality of the model.).
Shnitzer, Chen, and Franco are all considered to be analogous to the claimed invention because they are in the same field of artificial intelligence and more specifically applied to language models. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Shnitzer in view of Chen with the above teachings of Franco. Doing so may make processing more efficient (faster) and more accurate (Franco [0024] “For example, a segmentation mask can be provided to a classifier, which provides a prediction of what object is represented in the portion of the image corresponding to a segmentation layer of the segmentation mask. This focusing can make the further processing more efficient (faster) and more accurate.”).
Regarding claim 6, Shnitzer in view of Chen teaches the computer-implemented method of claim 1 as discussed above in the rejection of claim 1, but fails to teach wherein the one or more metric values include a count of a number of nouns included in the input.
However, Franco teaches wherein the one or more metric values include a count of a number of nouns included in the input (Franco [0059] “The description generator 160 may extract all nouns from the prompt, generating a set of tokens from the prompt. The nouns can include noun phrases.”; [0063] “The cross-attention extractor 162 may output a cross-attention tensor A.sub.N corresponding only to the noun tokens. This may be represented as Equation (8) where L corresponds to the number of indices in the token lists 163 (e.g., where a noun appears twice in a prompt, it will correspond with two indices in the token lists 163).”; [0073] “Pixel accuracy (ACC) and mean intersection over union (mIoU) are used to measure segmentation performance of the example implementation.” Franco teaches counting a number of nouns included in the input prompt and measuring the output performance of the model.).
Shnitzer, Chen, and Franco are all considered to be analogous to the claimed invention because they are in the same field of artificial intelligence and more specifically applied to language models. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Shnitzer in view of Chen with the above teachings of Franco. Doing so may make processing more efficient (faster) and more accurate (Franco [0024] “For example, a segmentation mask can be provided to a classifier, which provides a prediction of what object is represented in the portion of the image corresponding to a segmentation layer of the segmentation mask. This focusing can make the further processing more efficient (faster) and more accurate.”).
Regarding claim 16, the rejection of claim 11 is incorporated herein. Further, the limitations in this claim are taught by Shnitzer in view of Chen in further view of Franco for the same reasons disclosed above in the rejections of claims 5 or 6.
Claim 7 is rejected under 35 U.S.C. 103 as being unpatentable over Shnitzer et al. (Large Language Model Routing with Benchmark Datasets) (“Shnitzer”) in view of Chen et al. (FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance) (“Chen”) in further view of Wu et al. (Effective Neural Modeling Leveraging Readability Features for Automated Essay Scoring) (“Wu”).
Regarding claim 7, Shnitzer in view of Chen teaches the computer-implemented method of claim 1, as discussed above in the rejection of claim 1, but fails to teach wherein the one or more metric values include a reading time duration associated with the input.
However, Wu teaches wherein the one or more metric values include a reading time duration associated with the input (Wu Section 3 Methodology “Originally intended for readability assessment tasks, the LFTK toolkit is designed for multilingual purposes but also with English applications in mind. It comprises a comprehensive set of linguistic elements spanning multiple domains:… Surface features include general text properties like word count, punctuation frequency, and traditional readability and read time metrics.”; Section 4.3 Experimental Results “From our experimental results, we observe that the inclusion of auxiliary readability-aware features can indeed augment the overall performance, particularly enhancing accuracy across subsets when combined with BERT and the prototypical network (SED). Readability-aware features, such as text complexity and sentence structure, provide additional context that complements the semantic representation captured by BERT… In essence, these findings underscore the value of incorporating readability-aware features into our model, showcasing their ability to improve performance when combined with robust models and effective distance metrics.” Wu teaches read time duration metrics associated with an essay as input and measuring the model accuracy based on the readability metrics.).
Shnitzer, Chen, and Wu are all considered to be analogous to the claimed invention because they are in the same field of artificial intelligence and more specifically applied to language models. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Shnitzer in view of Chen with the above teachings of Wu. Doing so may help alleviate data scarcity and imbalanced data distribution problems (Wu Section 1 Introduction “On the ground of the above observations, we in this paper put forward a promising neural method which leverages BERT as the backbone model in conjunction with an effective metric based learning approach for alleviating the data scarcity and imbalanced data distribution problems.”).
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to KURT NICHOLAS PRESSLY whose telephone number is (703)756-4639. The examiner can normally be reached M-F 8-4.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kamran Afshar can be reached at (571) 272-7796. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/KURT NICHOLAS PRESSLY/Examiner, Art Unit 2125
/KAMRAN AFSHAR/Supervisory Patent Examiner, Art Unit 2125