DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Status of the Claims
Claims 1-20 are pending for examination.
Claims 1, 9, 17 and 19 are independent Claims.
Claims 1-20 are rejected under 35 U.S.C. §101 (abstract idea).
Claims 1-20 are rejected under 35 U.S.C. §103.
Claim Interpretation
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked.
As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph:
(A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function;
(B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and
(C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function.
Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function.
Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function.
Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action.
Such claim limitation(s) is/are: “means for obtaining”, “means for obtaining”, “means for determining” and “means for backpropagating” in claims 19-20.
Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof. The limitations are construed as “one or more processors” as indicated in paragraph 0072 of the Applicant’s Specification.
If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to a Judicial Exception without significantly more.
Independent Claims
As Claims 1, 9, 17 and 19:
Step 1: Are the Claims to a process, machine, manufacture or composition of matter? Yes.
Step 2A: Are the Claims directed to a law of nature, a natural phenomenon (product of nature) or an abstract idea? Yes, the Claims is an abstract idea. See the analysis below.
The Claim recites:
An apparatus for training a machine learning (ML) model, comprising:
at least one memory; and
at least one processor coupled to the at least one memory and configured to:
obtain, from a teacher ML model, a first prediction based on an input;
obtain, from a student ML model, a second prediction based on the input;
determine a loss based on a difference between the second prediction from the first prediction, wherein the loss comprises a variance reduced total variation distance (TVD) loss based on an unbiased gradient estimate of the loss; and
backpropagate the loss through the student ML model to train the student ML model.
The non-emphasized limitations describe abstract processes while emphasized limitations recited additional limitation(s).
Regarding the non-emphasized limitations:
Step 2A prong 1:
“determine a loss based on a difference between the second prediction from the first prediction, wherein the loss comprises a variance reduced total variation distance (TVD) loss based on an unbiased gradient estimate of the loss; and ” is/are directed to a mental processes group of abstract idea. Mental processes are defined as concepts that can practically be performed in the human mind, or by a human using pen and paper as a physical aid. Examples of mental processes includes observations, evaluations, judgements and opinions.
These steps are considered mental processes group of abstract idea.
Step 2A prong 2:
Limitations “at least one memory; and
at least one processor coupled to the at least one memory and configured to:
obtain, from a teacher ML model, a first prediction based on an input;
obtain, from a student ML model, a second prediction based on the input;
backpropagate the loss through the student ML model to train the student ML model.” are insignificant extra solution activity. See MPEP §2106.05(g).
The Claim(s) does not recite additional elements that integrate the judicial exception into a practical application.
Step 2B: Does the Claim recite additional elements that integrate the Judicial Exception into a practical application? No.
Limitations “at least one memory; and at least one processor coupled to the at least one memory and configured to:” was considered insignificant extra solution activity in Step 2A, and thus it is reevaluated in Step 2B to determine if it is more than what is well-understood, routine, conventional activity in the field. The addition of insignificant extra-solution activity does not amount to an inventive concept, particularly when the activity is well-understood or conventional (MPEP 2106.05(d)).
Limitations “obtain, from a teacher ML model, a first prediction based on an input;
obtain, from a student ML model, a second prediction based on the input;
backpropagate the loss through the student ML model to train the student ML model” were extra-solution activity in Step 2A, and thus it is reevaluated in Step 2B to determine if it is more than what is well-understood, routine, conventional activity in the field. The addition of insignificant extra-solution activity does not amount to an inventive concept, particularly when the activity is well-understood or conventional (MPEP 2106.05(d)). This appears to be well-understood, routine, conventional as evidenced by MPEP 2106.05(d)(II)(i. Receiving or transmitting data over a network, e.g., using the Internet to gather data”).
The claim is directed to mental processes group of abstract ideas. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. The claim is not patent eligible.
Dependent Claims
As Claims 2, 10, 18 and 20, the Claims recite “wherein the variance reduced TVD loss is based on advantage normalization.”
The non-emphasized limitations describe abstract processes while emphasized limitations recited additional limitation(s).
Step 2A: Are the Claims directed to a law of nature, a natural phenomenon (product of nature) or an abstract idea? Yes, the Claims is an abstract idea. Prong 1: There are no additional abstract idea(s). Prong 2: The limitation “wherein the variance reduced TVD loss is based on advantage normalization” are field of use and technology environment. See MPEP §2106.05(h). Claim(s) does not recite additional elements that integrate the judicial exception into a practical application.
Step 2B: Does the Claim recite additional elements that amount to significantly more than the Judicial Exception? No. The Claim is not patent eligible.
As Claims 3 and 11, the Claims recite “wherein the variance reduced TVD loss is based on advantage normalization.”
The non-emphasized limitations describe abstract processes while emphasized limitations recited additional limitation(s).
Step 2A: Are the Claims directed to a law of nature, a natural phenomenon (product of nature) or an abstract idea? Yes, the Claims is an abstract idea. Prong 1: The limitation “ wherein the variance reduced TVD loss is based on advantage normalization” is directed to mathematical calculations group of abstract idea. Prong 2: There is no additional limitation(s). Claim(s) does not recite additional elements that integrate the judicial exception into a practical application.
Step 2B: Does the Claim recite additional elements that amount to significantly more than the Judicial Exception? No. There is no additional limitation(s). The Claim is not patent eligible.
As Claims 4 and 12, the Claims recite “wherein the variance reduced TVD loss includes negative values.”
The non-emphasized limitations describe abstract processes while emphasized limitations recited additional limitation(s).
Step 2A: Are the Claims directed to a law of nature, a natural phenomenon (product of nature) or an abstract idea? Yes, the Claims is an abstract idea. Prong 1: There are no additional abstract idea(s). Prong 2: The limitation “wherein the variance reduced TVD loss includes negative values” are field of use and technology environment. See MPEP §2106.05(h). Claim(s) does not recite additional elements that integrate the judicial exception into a practical application.
Step 2B: Does the Claim recite additional elements that amount to significantly more than the Judicial Exception? No. The Claim is not patent eligible.
As Claims 5 and 13, the Claims recite “wherein the teacher ML model includes more layers than the student ML model.”
The non-emphasized limitations describe abstract processes while emphasized limitations recited additional limitation(s).
Step 2A: Are the Claims directed to a law of nature, a natural phenomenon (product of nature) or an abstract idea? Yes, the Claims is an abstract idea. Prong 1: There are no additional abstract idea(s). Prong 2: The limitation “wherein the teacher ML model includes more layers than the student ML model” are field of use and technology environment. See MPEP §2106.05(h). Claim(s) does not recite additional elements that integrate the judicial exception into a practical application.
Step 2B: Does the Claim recite additional elements that amount to significantly more than the Judicial Exception? No. The Claim is not patent eligible
As Claims 6 and 14, the Claims recite “wherein the teacher ML model comprises a large language model.”
The non-emphasized limitations describe abstract processes while emphasized limitations recited additional limitation(s).
Step 2A: Are the Claims directed to a law of nature, a natural phenomenon (product of nature) or an abstract idea? Yes, the Claims is an abstract idea. Prong 1: There are no additional abstract idea(s). Prong 2: The limitation “wherein the teacher ML model comprises a large language model” are field of use and technology environment. See MPEP §2106.05(h). Claim(s) does not recite additional elements that integrate the judicial exception into a practical application.
Step 2B: Does the Claim recite additional elements that amount to significantly more than the Judicial Exception? No. The Claim is not patent eligible
As Claims 7 and 15, the Claims recite “wherein the input comprises a textual data.”
The non-emphasized limitations describe abstract processes while emphasized limitations recited additional limitation(s).
Step 2A: Are the Claims directed to a law of nature, a natural phenomenon (product of nature) or an abstract idea? Yes, the Claims is an abstract idea. Prong 1: There are no additional abstract idea(s). Prong 2: The limitation “wherein the input comprises a textual data” are field of use and technology environment. See MPEP §2106.05(h). Claim(s) does not recite additional elements that integrate the judicial exception into a practical application.
Step 2B: Does the Claim recite additional elements that amount to significantly more than the Judicial Exception? No. The Claim is not patent eligible
As Claims 8 and 16, the Claims recite “wherein the teacher ML model and student ML model are for use in an autoregressive model.”
The non-emphasized limitations describe abstract processes while emphasized limitations recited additional limitation(s).
Step 2A: Are the Claims directed to a law of nature, a natural phenomenon (product of nature) or an abstract idea? Yes, the Claims is an abstract idea. Prong 1: There are no additional abstract idea(s). Prong 2: The limitation “wherein the teacher ML model and student ML model are for use in an autoregressive model” are field of use and technology environment. See MPEP §2106.05(h). Claim(s) does not recite additional elements that integrate the judicial exception into a practical application.
Step 2B: Does the Claim recite additional elements that amount to significantly more than the Judicial Exception? No. The Claim is not patent eligible
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-2, 4-10 and 12-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Hori et al. (U.S. 2021/0082398 hereinafter Hori) in view of Dhurandhar et al. (U.S. 11,972,344 hereinafter Dhurandhar).
As Claim 1, Hori teaches a method for training a machine learning (ML) model, comprising:
obtaining, from a teacher ML model, a first prediction based on an input (Hori (¶0030 line 13-15), “the first encoder-decoder generates first output values based on the first audio-visual datasets with the first corresponding video description sentences 20”);
obtaining, from a student ML model, a second prediction based on the input (Hori (¶0030 line 15-20), “providing the first audio visual datasets excluding the first corresponding video description sentences 209 to the second multimodal encoder-decoder for dialog response generation 210, wherein the second multimodal encoder-decoder generates second output values based on the first audio-visual datasets”);
determining a loss based on a difference between the second prediction from the first prediction (Hori (¶0030 last 5 lines), “wherein an optimizer module updates second network parameters of the second multimodal encoder-decoder until errors between the first output values and the second output values are reduced into a predetermined range, wherein the errors are computed based on a loss function”), based on an unbiased gradient estimate of the loss (Hori (¶0030 last 5 lines), “wherein an optimizer module updates second network parameters of the second multimodal encoder-decoder until errors between the first output values and the second output values are reduced into a predetermined range, wherein the errors are computed based on a loss function”); and
backpropagating the loss through the student ML model to train the student ML model (Hori (¶0030 last 5 lines), “wherein an optimizer module updates second network parameters of the second multimodal encoder-decoder until errors between the first output values and the second output values are reduced into a predetermined range, wherein the errors are computed based on a loss function”).
Hori may not explicitly disclose:
wherein the loss comprises a variance reduced total variation distance (TVD) loss
Dhurandhar teaches:
wherein the loss comprises a variance reduced total variation distance (TVD) loss (Dhurandhar (col. 6 line 18-25 and 50-60), “DTV is the total variation distance between two distributions. P(x|y) are the class distributions in dataset D.” Dhurandhar (col. 5 line 10-12), “That is, equation (1) enables learning the optimal simple model S* by estimating the corresponding optimal weights W* which are used to weight examples in Ds”)
Hori discloses a system/method to optimize a student model by using loss function between student model output and teacher model output. Dhurandhar discloses a system to use TVD (total variation distance) for improving student model accuracy. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify student model of Hori instead be a student model taught by Dhurandhar, with a reasonable expectation of success. The motivation would be “to identify examples that the simple model will most likely fail on (i.e., identify truly 'hard' examples as shown in FIG. 4B). The method then informs the simple model to ignore these examples and make it expend more effort on other relatively easier examples that it could perform well on, with the eventual hope of improving generalization performance” (Dhurandhar (col. 4 line 5-11)).
As Claim 2, besides Claim 1, Hori in view of Dhurandhar teaches wherein the variance reduced TVD loss is based on advantage normalization (Dhurandhar (col. 5 line 10-12), “That is, equation (1) enables learning the optimal simple model S* by estimating the corresponding optimal weights W* which are used to weight examples in Ds”).
As Claim 4, besides Claim 2, Hori in view of Dhurandhar wherein the variance reduced TVD loss includes negative values Dhurandhar (col. 6 line 18-25 and 50-60), “DTV is the total variation distance between two distributions. P(x|y) are the class distributions in dataset D.” “The error of the Bayes optimal classifier” shows that the value could be negative).
As Claim 5, besides Claim 1, Hori in view of Dhurandhar teaches wherein the teacher ML model includes more layers than the student ML model (Dhurandhar (col. 7 line 20-21), “which is a popular method 20 for training relatively simpler neural networks.” Student neural network is construed as simpler neural networks with less layers than the teacher model).
As Claim 6, besides Claim 1, Hori in view of Dhurandhar teaches wherein the teacher ML model comprises a large language model (Hori (¶0031 line 11-17), “arranging a first multimodal encoder-decoder for video description or dialogue response generation 210 having a first input and a first output through 110, wherein the first multimodal encoder-decoder 210 has been pretrained by training audio video datasets 195 with training video description sentences 195”).
As Claim 7, besides Claim 1, Hori in view of Dhurandhar teaches wherein the input comprises a textual data (Hori (¶0031 line 6-8), “wherein the first multimodal encoder-decoder has been pretrained by training audio-video datasets using video description sentences 209,”).
As Claim 8, besides Claim 1, Hori in view of Dhurandhar teaches wherein the teacher ML model and student ML model are for use in an autoregressive model (Hori (¶0050 line 1-3), “is a user's question 208, the sentence Q = wQ,1, … , wQ, N is encoded with word embedding and BLSTM layers (autoregressive model).”).
As Claim 9, Hori teaches an apparatus for training a machine learning (ML) model, comprising:
at least one memory (Hori (¶0031 line 8), memory); and
at least one processor coupled to the at least one memory and configured to (Hori (¶0031 line 7), one or more processors):
(The rest of the limitation(s) are rejected for the same reasons as Claim 1)
As Claims 10 and 12-16, the Claims are rejected for the same reasons as Claims 2 and 4-8, respectively.
As Claims 17-18, the Claims are rejected for the same reasons as Claim 1-2, respectively.
As Claims 19-23, the Claims are rejected for the same reasons as Claim 1-2, respectively.
Claim(s) 3 and 11 is/are rejected under 35 U.S.C. 103 as being unpatentable over Hori in view of Dhurandhar in further view of Teig et al. (U.S. 12,675,635 hereinafter Teig).
As Claim 3, besides Claim 2, Hori in view of Dhurandhar may not explicitly disclose:
wherein advantage normalization comprises subtracting a mean of an advantage function and dividing by a standard deviation.
Teig teaches:
wherein advantage normalization comprises subtracting a mean of an advantage function and dividing by a standard deviation (Teig (col. 14 line 31-36), “the reference frame adjustor 512 performs this translation by subtracting the student network output mean and dividing by the student network standard deviation, then multiplying by the teacher network standard deviation and adding the teacher network mean to each student network output.”).
Hori in view of Dhurandhar discloses a system/method to optimize a student model by using loss function between student model output and teacher model output. Teig disclose a system/method for calculating mean and standard deviation for student-teacher model training. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify student model of Hori instead be a student model taught by Dhurandhar, with a reasonable expectation of success. The motivation would be “to properly compare the traditional loss function values for student network outputs and teacher network outputs” (Teig (col. 14 line 25-27)).
As Claim 11, the Claim is rejected for the same reasons as Claim 3.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Zhang et al. (U.S. 11,886,554) discloses a system/metho to improve the inference accuracy of the student model.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to NHAT HUY T NGUYEN whose telephone number is (571)270-7333. The examiner can normally be reached M-F: 12:00-8:00 EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Viker Lamardo can be reached at 571-270-5871. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/NHAT HUY T NGUYEN/Primary Examiner, Art Unit 2147