DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
This action is in response to communications filed on 06/06/2023. Claims 1-20 are pending and have been examined.
Priority
Applicant’s claim for the benefit of a prior-filed application under 35 U.S.C. 119(e) or under 35 U.S.C. 120, 121, 365(c), or 386(c) is acknowledged.
Information Disclosure Statement
The information disclosure statement (IDS) submitted was filed on 08/14/2023. The submission is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
The information disclosure statement (IDS) submitted was filed on 07/11/2025. The submission is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Objections
Claims 1, 5, 9, and 19 are objected to because of the following informalities:
As per claim 1, the phrase “executable…for” in lines 3-4 raises question as to whether the features following are limiting, or merely refer to intended use.
As per claim 5, the term “for” in line 2 raises question as to whether the features following are limiting, or merely refer to intended use. This similarly applies to claim 9.
As per claim 9, there is lack of antecedent basis for “the same values” in lines 2-3. This similarly applies to claim 19.
Appropriate correction is required.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 10 and 20 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
As per claim 10, there is lack of antecedent basis for “the respective student mask token predictions. This similarly applies to claim 20.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 1-2, 4, 11-12, and 14 are rejected under 35 U.S.C. 103 as being unpatentable over Lai et al. (US 20210182662 A1) in view of Ghazvininejad et al., ("Mask-Predict: Parallel Decoding of Conditional Masked Language Models," arXiv:1904.09324v2 [cs.CL], September 4, 2019, 10 pages as cited in IDS dated 07/11/2025).
As per independent claim 1, Lai teaches a system comprising: one or more processors and one or more non-transitory computer-readable media having instructions executable by the one or more processors (e.g. in paragraph 54, “one or more software modules configured to implement certain functionalities disclosed herein, as well as hardware configured to enable such implementation. These hardware and software components may include, among other things, a processor 132, memory 134”) for:
identifying a masked training output having two or more masked positions of a labeled output token sequence comprising a plurality of positions having respective output tokens and an associated input token sequence (e.g. in paragraphs 10, 75, and 83, “A masked token comprises one or more words that are masked or hidden in the training data… training data includes a plurality of masked tokens and a plurality of unmasked tokens… the inputs of the model 104 sequentially (or parallelly) receives various tokens of the labelled training data… the inputs of the model 106 sequentially (or parallelly) receive various tokens of the training data” and figures 4A-4B);
determining a set of teacher mask token predictions for the masked positions by applying a teacher model to the masked training output and the input token sequence, the set of teacher mask token predictions including teacher mask token predictions determined (e.g. in paragraphs 10 and 75, “Training data is input to both the student and teacher models. The training data includes a plurality of masked tokens and a plurality of unmasked tokens… the teacher model generates a third prediction and a fourth prediction… the inputs of the model 104 sequentially (or parallelly) receives various tokens of the labelled training data”);
determining a set of student mask token predictions for the masked positions by applying a student model to the masked training output and the input token sequence associated with the labeled output token sequence (e.g. in paragraphs 10 and 83, “Training data is input to both the student and teacher models. The training data includes a plurality of masked tokens and a plurality of unmasked tokens. The student model generates a first prediction and a second prediction… the inputs of the model 106 sequentially (or parallelly) receive various tokens of the training data”);
determining a training loss based on the set of teacher mask token predictions compared to the set of student mask token predictions (e.g. in paragraph 10, “student model generates a first prediction and a second prediction, and the teacher model generates a third prediction and a fourth prediction...a first loss function is generated based at least in part on a comparison of the first prediction and the third prediction (with respect to the masked token)”); and
updating parameters of the student model based on the training loss (e.g. in paragraphs 10 and 22, “student model is then trained based at least in part on the first…loss functions… model training system aims to tune parameters of the student model”),
but does not specifically teach wherein the models including non-autoregressive models and iteratively applying for a plurality of sequential iterations to determine predictions at different iterations in the plurality of sequential iterations.
However, Ghazvininejad teaches models including non-autoregressive models and iteratively applying for a plurality of sequential iterations to determine predictions at different iterations in the plurality of sequential iterations (e.g. in abstract and section 3, “approach allows for efficient iterative decoding, where we first predict all of the target words non-autoregressively, and then repeatedly mask out and regenerate the subset of words that the model is least confident about. By applying this strategy for a constant number of iterations, our model improves state-of-the-art performance levels for non-autoregressive and parallel decoding translation models… mask-predict algorithm, which decodes an entire sequence in parallel within a constant number of cycles. At each iteration, the algorithm selects a subset of tokens to mask, and then predicts them (in parallel) using an underlying CMLM” and section 5.1, “multiple mask-predict iterations alleviate the multi-modality problem”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Lai to include the teachings of Ghazvininejad because one of ordinary skill in the art would have recognized the benefit of improving performance and/or alleviating issues with multi-modality.
As per claim 2, the rejection of claim 1 is incorporated and the combination further teaches wherein the teacher mask token predictions and student mask token predictions are score distributions of output tokens and the training loss is a comparison of the score distributions for each masked token (e.g. Lai, in paragraphs 10 and 24-25, “provides a prediction in the form of a probability vector [i.e. score distribution]… generates a loss function L3 that is based on a comparison of the above discussed probability vectors Ps3 and Pt3 generated by the student and teacher models”).
As per claim 4, the rejection of claim 1 is incorporated and the combination further teaches wherein the masked training output includes unmasked tokens (e.g. Lai, in paragraph 10, “training data includes a plurality of masked tokens and a plurality of unmasked tokens”).
Claims 11-12 and 14 are the method claims corresponding to system claims 1-2 and 4 and are rejected under the same reasons set forth.
Claims 3 and 13 are rejected under 35 U.S.C. 103 as being unpatentable over Lai et al. (US 20210182662 A1) in view of Ghazvininejad et al., ("Mask-Predict: Parallel Decoding of Conditional Masked Language Models," arXiv:1904.09324v2 [cs.CL], September 4, 2019, 10 pages as cited in IDS dated 07/11/2025) as applied above, and further in view of Haidar et al. (US 20220335303 A1).
As per claim 3, the rejection of claim 2 is incorporated and the combination further teaches wherein the training loss is a function of the score distributions of the teacher mask token predictions and the student mask token predictions (e.g. Lai, in paragraphs 10 and 24-25, “provides a prediction in the form of a probability vector [i.e. score distribution]… generates a loss function L3 that is based on a comparison of the above discussed probability vectors Ps3 and Pt3 generated by the student and teacher models”), but does not specifically teach wherein the function includes KL-divergence. However, Haidar teaches KL-divergence (e.g. in paragraph 53, “L.sub.CE is therefore a function that is used to compute a Cross-Entropy (CE) loss between the ground-truth label of the labeled training data sample and the output of the student, S.sub.θ(x), and L.sub.KD is a function that is used to compute a KD loss based on the Kullback-Leibler (KL) divergence between the teacher inference data 24 (i.e. teacher prediction) and the student inference data 34 (i.e. student prediction)”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of the combination to include the teachings of Haidar because one of ordinary skill in the art would have recognized the benefit of incorporating well-known functions (also amounts a simple substitution that yields predictable results [e.g. see KSR Int'l Co v. Teleflex Inc., 550 US 398,82 USPQ2d 1385,1396 (U.S. 2007) and MPEP 2143(B)].).
Claim 13 is the method claim corresponding to system claim 3 and is rejected under the same reasons set forth.
Claims 5-7 and 15-17 are rejected under 35 U.S.C. 103 as being unpatentable over Lai et al. (US 20210182662 A1) in view of Ghazvininejad et al., ("Mask-Predict: Parallel Decoding of Conditional Masked Language Models," arXiv:1904.09324v2 [cs.CL], September 4, 2019, 10 pages as cited in IDS dated 07/11/2025) as applied above, and further in view of Liu et al. (US 20220156593 A1).
As per claim 5, the rejection of claim 1 is incorporated, but the combination does not specifically teach updating the teacher model based on the update to the student model. However, Liu teaches updating a teacher model based on an update to the student model (e.g. in paragraphs 17 and 39, “improve…learning… During training, the parameters θ.sub.t of the teacher model (e.g., the teacher models 230 and 236) are updated at each step as the exponential-moving-average of the student model (e.g., the encoder 225)'s parameters θ.sub.sn)”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of the combination to include the teachings of Liu because one of ordinary skill in the art would have recognized the benefit of improving learning.
As per claim 6, the rejection of claim 5 is incorporated and the combination further teaches wherein updating the teacher model comprises modifying parameters of the teacher model as a moving average with parameters of the student mode (e.g. Liu, in paragraph 39, “During training, the parameters θ.sub.t of the teacher model (e.g., the teacher models 230 and 236) are updated at each step as the exponential-moving-average of the student model (e.g., the encoder 225)'s parameters θ.sub.sn)”).
As per claim 7, the rejection of claim 5 is incorporated and the combination further teaches wherein updating the teacher model comprises replacing parameters of the teacher model with parameters of the student model (e.g. Liu, in paragraph 39, “During training, the parameters θ.sub.t of the teacher model (e.g., the teacher models 230 and 236) are updated at each step as the exponential-moving-average of the student model (e.g., the encoder 225)'s parameters θ.sub.sn)”).
Claims 15-17 are the method claims corresponding to system claims 5-7 and are rejected under the same reasons set forth.
Claims 8 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Lai et al. (US 20210182662 A1) in view of Ghazvininejad et al., ("Mask-Predict: Parallel Decoding of Conditional Masked Language Models," arXiv:1904.09324v2 [cs.CL], September 4, 2019, 10 pages as cited in IDS dated 07/11/2025) as applied above, and further in view of Guo et al. (US 20220129731 A1) and Evans et al. (US 20230360734 A1).
As per claim 8, the rejection of claim 1 is incorporated, but the combination does not specifically teach increasing a number of masked positions and a number of the plurality of sequential iterations of the teacher model after updating parameters of the student model.
However, Guo teaches increasing a number of a plurality of sequential iterations of a teacher model after updating parameters of a student model (e.g. in paragraph 69, “gradient is back propagated, and a parameter of the student network and a parameter of the teacher network are updated at the same time. The number of the iterations iter is increased by 1, and step 4 is repeated”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of the combination to include the teachings of Guo because one of ordinary skill in the art would have recognized the benefit of facilitating training.
but does not specifically teach increasing a number of masked positions.
However, Evans teaches increasing a number of masked positions (e.g. in paragraphs 24-25 and 75, “training system 106 increases the expected amount of data that is removed or masked from the full MSAs 114 by the reduction engine 120 at each iteration of the training procedure… can improve the performance of the structure prediction neural network”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of the combination to include the teachings of Evans because one of ordinary skill in the art would have recognized the benefit of improving performance.
Claim 18 is the method claim corresponding to system claim 8 and is rejected under the same reasons set forth.
Claims 9 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Lai et al. (US 20210182662 A1) in view of Ghazvininejad et al., ("Mask-Predict: Parallel Decoding of Conditional Masked Language Models," arXiv:1904.09324v2 [cs.CL], September 4, 2019, 10 pages as cited in IDS dated 07/11/2025) as applied above, and further in view of Wang et al. (US 20220165288 A1).
As per claim 9, the rejection of claim 1 is incorporated, but the combination does not specifically teach initializing parameters of the teacher model and the student model to the same values. However, Wang teaches initializing parameters of a teacher model and a student model to same values (e.g. in paragraph 132, “the teacher model and the student model are initialized…and parameters of the teacher model and the student model are kept the same in the first iteration process”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of the combination to include the teachings of Wang because one of ordinary skill in the art would have recognized the benefit of establishing a baseline.
Claim 19 is the method claim corresponding to system claim 9 and is rejected under the same reasons set forth.
Claims 10 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Lai et al. (US 20210182662 A1) in view of Ghazvininejad et al., ("Mask-Predict: Parallel Decoding of Conditional Masked Language Models," arXiv:1904.09324v2 [cs.CL], September 4, 2019, 10 pages as cited in IDS dated 07/11/2025) as applied above, and further in view of Romero et al. ("FitNets: Hints for Thin Deep Nets," International Conference on Learning Representations, arXiv preprint arXiv:1412.6550v4 [cs.LG], March 27, 2015, 13 pages as cited in IDS sated 07/11/2025).
As per claim 10, the rejection of claim 1 is incorporated, but the combination does not specifically teach wherein the training loss includes a hidden state loss based on one or more hidden layer values of a hidden layer of the teacher model for each teacher mask token prediction compared to hidden layer values of a hidden layer of the student model for each of the respective student mask token predictions. However, Romero teaches a training loss including a hidden state loss based on one or more hidden layer values of a hidden layer of a teacher model for each teacher mask token prediction compared to hidden layer values of a hidden layer of the student model for each of respective student mask token predictions (e.g. in abstract and sections 1 and 2.2, “generalize better or run faster… We introduce intermediate-level hints from the teacher hidden layers to guide the training process of the student, i.e., we want the student network (FitNet) to learn an intermediate representation that is predictive of the intermediate representations of the teacher network… want the guided layer to be able to predict the output of the hint layer… choose the hint to be the middle layer of the teacher network. Similarly, we choose the guided layer to be the middle layer of the student network… train the FitNet parameters from the first layer up to the guided layer as well as the regressor parameters by minimizing the following loss function”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of the combination to include the teachings of Romero because one of ordinary skill in the art would have recognized the benefit of improving performance.
Claim 20 is the method claim corresponding to system claim 10 and is rejected under the same reasons set forth.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
For example,
Yavaz et al. (US 20210375269 A1) teaches “A stochastic imputation-based teacher and student selection may be implemented by leveraging mask augmentation. For example, the teacher mask augmentation module 420a is configured to sample the input sequence of tokens according to a first probability ∈.sub.t and replace the sampled tokens with the mask token, resulting in an augmented teacher input sequence {umlaut over (x)}.sup.(t)˜z({umlaut over (x)}|x,∈.sub.t). The student mask augmentation module 420b is configured to sample the input sequence of tokens according to a second probability ∈.sub.s and replace the sampled tokens with the mask token, resulting in an augmented student input sequence” (e.g. in paragraph 47).
Any inquiry concerning this communication or earlier communications from the examiner should be directed to WILLIAM WONG whose telephone number is (571)270-1399. The examiner can normally be reached Monday-Friday 9am-5pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, TAMARA KYLE can be reached at (571)272-4241. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/W.W/Examiner, Art Unit 2144 08/22/2026
/TAMARA T KYLE/Supervisory Patent Examiner, Art Unit 2144