Prosecution Insights
Last updated: October 02, 2026
Application No. 18/733,226

GENERATING SMALL LANGUAGE MODEL VIA TWO-PHASE TRAINING

Non-Final OA §101§103
Filed
Jun 04, 2024
Priority
Sep 11, 2023 — provisional 63/537,770 +1 more
Examiner
WENG, PEI YONG
Art Unit
Tech Center
Assignee
Microsoft Technology Licensing, LLC
OA Round
1 (Non-Final)
79%
Grant Probability
Favorable
1-2
OA Rounds
9m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 79% — above average
79%
Career Allowance Rate
514 granted / 647 resolved
+19.4% vs TC avg
Strong +23% interview lift
Without
With
+22.8%
Interview Lift
resolved cases with interview
Typical timeline
3y 1m
Avg Prosecution
32 currently pending
Career history
665
Total Applications
across all art units

Statute-Specific Performance

§101
13.0%
-27.0% vs TC avg
§103
55.8%
+15.8% vs TC avg
§102
20.5%
-19.5% vs TC avg
§112
7.0%
-33.0% vs TC avg
Black line = Tech Center average estimate • Based on career data from 647 resolved cases

Office Action

§101 §103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . DETAILED ACTION This action is responsive to the following communication: Non-Provisional Application filed Jun. 4, 2024. Claims 1-20 are pending in the case. Claims 1, 9 and 16 are independent claims. Claim Rejections - 35 U.S.C. § 101 35 U.S.C. § 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-20 are rejected under 35 U.S.C. § 101 because the claimed invention is directed to an abstract idea without significantly more. As to claim 1: Step 1 Analysis: Is the claim to a process, machine, manufacture or composition of matter? See MPEP § 2106.03. Yes, the claim is to a process. Step 2A Prong One Analysis: Does the claim recite an abstract idea, law of nature, or natural phenomenon? See MPEP § 2106.04(II)(A)(1). Yes, the limitation “obtaining a general dataset, the general dataset including a plurality of general data; annotating a subset of the general dataset based on one or more classifier metrics indicative of a quality of the general dataset, the subset of the general dataset being representative of the general dataset” is the abstract idea of a mental process that can practically be performed in the human mind, with or without the use of a physical aid such as pen and paper (including an observation, evaluation, judgment, opinion). See MPEP § 2106.04(a)(2)(III). Yes, the limitation “training a classifier based on the annotated subset of the general dataset and the one or more classifier metrics; analyzing each general data of the general dataset to determine a score for each of the one or more classifier metrics associated with the respective general data using the trained classifier; generating a filtered general dataset by filtering the general dataset based on one or more filters, the one or more filters indicative of threshold scores for corresponding classifier metrics; training the small language model with the filtered general dataset” is the abstract idea of a mental process that can practically be performed in the human mind, with or without the use of a physical aid such as pen and paper (including an observation, evaluation, judgment, opinion). See MPEP § 2106.04(a)(2)(III). Yes, the limitation “generating a synthetic dataset for refining the small language model; and subsequent to training the small language model with the filtered general dataset, training the small language model with the synthetic dataset.” is the abstract idea of a mental process that can practically be performed in the human mind, with or without the use of a physical aid such as pen and paper (including an observation, evaluation, judgment, opinion). See MPEP § 2106.04(a)(2)(III). Perkins US 2022/0253647 Para [38] – “the machine learning model and/or one or more sets of inference data may be reviewed manually to evaluate the machine learning model and/or the one or more sets of inference data and/or to determine whether the one or more sets of inference data are compatible with the machine learning model.” Step 2A Prong Two Analysis: Does the claim recite additional elements that integrate the judicial exception into a practical application? See MPEP § 2106.04(d). No, the limitation “training a classifier” is an additional element that amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. See MPEP §§ 2106.04(d), 2106.05(f)(2). Step 2B Analysis: Does the claim recite additional elements that amount to significantly more than the judicial exception? See MPEP § 2106.05. No As to claim 9: Step 1 Analysis: Is the claim to a process, machine, manufacture or composition of matter? See MPEP § 2106.03. Yes, the claim is to a machine. Step 2A Prong One Analysis: Does the claim recite an abstract idea, law of nature, or natural phenomenon? See MPEP § 2106.04(II)(A)(1). Yes, the limitation “obtaining a general dataset, the general dataset including a plurality of general data; annotating a subset of the general dataset based on one or more classifier metrics indicative of a quality of the general dataset, the subset of the general dataset being representative of the general dataset” is the abstract idea of a mental process that can practically be performed in the human mind, with or without the use of a physical aid such as pen and paper (including an observation, evaluation, judgment, opinion). See MPEP § 2106.04(a)(2)(III). Yes, the limitation “training a classifier based on the annotated subset of the general dataset and the one or more classifier metrics; analyzing each general data of the general dataset to determine a score for each of the one or more classifier metrics associated with the respective general data using the trained classifier; generating a filtered general dataset by filtering the general dataset based on one or more filters, the one or more filters indicative of threshold scores for corresponding classifier metrics; training the small language model with the filtered general dataset” is the abstract idea of a mental process that can practically be performed in the human mind, with or without the use of a physical aid such as pen and paper (including an observation, evaluation, judgment, opinion). See MPEP § 2106.04(a)(2)(III). Yes, the limitation “generating a synthetic dataset for refining the small language model; and subsequent to training the small language model with the filtered general dataset, training the small language model with the synthetic dataset.” is the abstract idea of a mental process that can practically be performed in the human mind, with or without the use of a physical aid such as pen and paper (including an observation, evaluation, judgment, opinion). See MPEP § 2106.04(a)(2)(III). Perkins US 2022/0253647 Para [38] – “the machine learning model and/or one or more sets of inference data may be reviewed manually to evaluate the machine learning model and/or the one or more sets of inference data and/or to determine whether the one or more sets of inference data are compatible with the machine learning model.” Step 2A Prong Two Analysis: Does the claim recite additional elements that integrate the judicial exception into a practical application? See MPEP § 2106.04(d). No, the limitation “computing device for generating a small language model” is an additional element that generally links the use of the judicial exception to a particular technological environment or field of use. See MPEP §§ 2106.04(d), 2106.05(h). No, the limitation “training the small language model” is an additional element that amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer. See MPEP §§ 2106.04(d), 2106.05(f)(1). The additional elements, taken alone or in combination, fail to integrate the judicial exception into a practical application. Step 2B Analysis: Does the claim recite additional elements that amount to significantly more than the judicial exception? See MPEP § 2106.05. No, the limitation “computing device for generating a small language model” is an additional element that generally links the use of the judicial exception to a particular technological environment or field of use. See MPEP § 2106.05(h). No, the limitation “training the small language model” is an additional element that amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer. See MPEP § 2106.05(f)(1). The additional elements, taken alone or in combination, fail to amount to significantly more than the judicial exception. As to claim 16: Step 1 Analysis: Is the claim to a process, machine, manufacture or composition of matter? See MPEP § 2106.03. Yes, the claim is to a machine. Step 2A Prong One Analysis: Does the claim recite an abstract idea, law of nature, or natural phenomenon? See MPEP § 2106.04(II)(A)(1). Yes, the limitation “obtaining a general dataset, the general dataset including a plurality of general data; annotating a subset of the general dataset based on one or more classifier metrics indicative of a quality of the general dataset, the subset of the general dataset being representative of the general dataset” is the abstract idea of a mental process that can practically be performed in the human mind, with or without the use of a physical aid such as pen and paper (including an observation, evaluation, judgment, opinion). See MPEP § 2106.04(a)(2)(III). Yes, the limitation “training a classifier based on the annotated subset of the general dataset and the one or more classifier metrics; analyzing each general data of the general dataset to determine a score for each of the one or more classifier metrics associated with the respective general data using the trained classifier; generating a filtered general dataset by filtering the general dataset based on one or more filters, the one or more filters indicative of threshold scores for corresponding classifier metrics; training the small language model with the filtered general dataset” is the abstract idea of a mental process that can practically be performed in the human mind, with or without the use of a physical aid such as pen and paper (including an observation, evaluation, judgment, opinion). See MPEP § 2106.04(a)(2)(III). Yes, the limitation “generating a synthetic dataset for refining the small language model; and subsequent to training the small language model with the filtered general dataset, training the small language model with the synthetic dataset.” is the abstract idea of a mental process that can practically be performed in the human mind, with or without the use of a physical aid such as pen and paper (including an observation, evaluation, judgment, opinion). See MPEP § 2106.04(a)(2)(III). Perkins US 2022/0253647 Para [38] – “the machine learning model and/or one or more sets of inference data may be reviewed manually to evaluate the machine learning model and/or the one or more sets of inference data and/or to determine whether the one or more sets of inference data are compatible with the machine learning model.” Step 2A Prong Two Analysis: Does the claim recite additional elements that integrate the judicial exception into a practical application? See MPEP § 2106.04(d). No, the limitation “computer storage medium storing computer-executable instructions that when executed cause at least one processor to perform operations” is an additional element that generally links the use of the judicial exception to a particular technological environment or field of use. See MPEP §§ 2106.04(d), 2106.05(h). No, the limitation “training the small language model” is an additional element that amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer. See MPEP §§ 2106.04(d), 2106.05(f)(1). The additional elements, taken alone or in combination, fail to integrate the judicial exception into a practical application. Step 2B Analysis: Does the claim recite additional elements that amount to significantly more than the judicial exception? See MPEP § 2106.05. No, the limitation “computer storage medium storing computer-executable instructions that when executed cause at least one processor to perform operations” is an additional element that generally links the use of the judicial exception to a particular technological environment or field of use. See MPEP § 2106.05(h). No, the limitation “ training the small language model” is an additional element that amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer. See MPEP § 2106.05(f)(1). The additional elements, taken alone or in combination, fail to amount to significantly more than the judicial exception. Claims 2-8, 10-15 and 17-20 are dependent claims. The claims recite additional limitations directed to specific data processing, but do not otherwise add any meaningful limits beyond the abstract idea. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-20 are rejected under 35 U.S.C. 103 as being unpatentable over Liu et al. (hereinafter Liu) U.S. Patent Publication No. 2024/0346244 in view of Merler et al. (hereinafter Merler) U.S. Patent Publication No. 2023/0259716. With respect to independent claim 1, Liu teaches a method for generating a model (see e.g., Para [4]-[6] [26] – “Additionally or alternatively, the enhanced dataset 112 may be used to train the machine learning model 114.”), the method comprising: obtaining a general dataset, the general dataset including a plurality of general data (see e.g., Para [4]-[6] [26] – “ the operations may include obtaining a dataset including one or more data subsets. The operations may additionally include training a language model to determine relationships between data in the data subsets in the dataset using one or more question answer pairs … extracting a value and a title from each of at least two data subsets in the dataset, and determining a question based on the titles, the values, and/or a target variable inferred from data included in the dataset”); annotating a subset of the general dataset based on one or more classifier metrics indicative of a quality of the general dataset, the subset of the general dataset being representative of the general dataset (see e.g., Para [134] [162]-[165]– “the one or more comparison metrics may include mutual information analyses, the Pearson correlation coefficient, chi2, Analysis of Variance (ANOVA), F-Value, and other comparison metrics that may be configured to compare data included in the data subsets stored in the enhanced dataset.” ”the one or more question answer pairs may be generated by generating one or more semantic similarity distributions that may correspond to information between data subsets in the dataset. Further, one or more domains for the data subsets may be determined based on the one or more semantic similarity distributions satisfying a threshold. In some embodiments, one or more question answer pairs may be generated corresponding to the one or more domains, where the questions in the question answer pairs may compare data subsets within the same domain.”); training a classifier based on the annotated subset of the general dataset and the one or more classifier metrics (see e.g., Para [162]-[166]); analyzing each general data of the general dataset to determine a score for each of the one or more classifier metrics associated with the respective general data using the trained classifier (see e.g., Para [144] [155] – “the data subset pairs may be selected based on the one or more similarity analyses—e.g., based on a particular similarity score. In some embodiments, a data subset pair may be selected based on a corresponding data similarity score being above a percentage. For example, the data subset pair may be selected based on the corresponding similarity score being in the top 5, 10, 20, 25, 30, 35 or some other percent of similarity scores compared to similarity scores corresponding to each of the other data subset pairs identified at block 302. “ “a data distribution similarity may be determined between data stored in the numerical data subset pair. In some embodiments, one or more data distribution similarity metrics and/or scores may be used to determine a similarity between data included in the numerical data subset pair.”); generating a filtered general dataset by filtering the general dataset based on one or more filters, the one or more filters indicative of threshold scores for corresponding classifier metrics; training the small language model with the filtered general dataset (see e.g., Para [136] [150]-[155] – “ In some embodiments, it may be determined that one or more new data subsets may be determined up to a threshold percentage (e.g., 60, 70, 80, 90, 100% or some other percent) of the data subsets in the original dataset. For example, the original dataset may include 400 original data subsets, it may be determined that no more than 75% of the number of data subsets included in the original dataset should be generated as new data subsets (e.g., no more than 300 new data subsets). In some embodiments, the data subset pairs that may not have satisfied the particular threshold may be filtered out of the enhanced dataset.”); generating a synthetic dataset for refining the small language model (see e.g., Para [18]-[26]). Liu does not expressly show that the model is a small language model and subsequent to training the model with the filtered general dataset, training the model with the synthetic dataset. However, Merler teaches similar feature (see e.g. Para [8][38][96] – “The invention further improves the performance of the selected model by employing pretraining techniques and/or utilizing data augmentation methods.”). Both Liu and Merler are directed to improving machine learning performance. Accordingly, it would have been obvious to the skilled artisan before the effective filing date of the claimed invention having Liu and Merler in front of them to modify the system of Liu to include the above feature. The motivation to combine Liu and Merler comes from Merler. Merler discloses the motivation to improve model efficiency and performance by selecting and training an small student language model (see e.g. see e.g. Para [8][38][96]). This motivation for combination also applies to the remaining claims which depend on this combination. With respect to dependent claim 2, the modified Liu teaches each general data of the general dataset is associated with a score for each of the one or more classifier metrics (see e.g., Para [83][135] – “ the one or more generated answers 108 may include one or more vectors, matrices, arrays, tensors and/or other collection of values that may indicate probability distributions corresponding to one or more answers 108 to the generated and/or synthesized questions. “ “each data subset pair may be assigned a value that may indicate an amount of information overlapping between the data included in the two data subsets. In some embodiments, the one or more pairs of data subsets that may have been compared using one or more mutual information analyses may be classified, rank ordered, or otherwise labeled according to mutual information scores corresponding to the one or more pairs of data subsets.”). With respect to dependent claim 3, the modified Liu teaches the one or more classifier metrics comprise factual knowledge, everyday knowledge, scientific knowledge, human behavior, toxicity, completeness, obscenity, obscurity, commonality, reasoning, promotional content, and/or unwanted content (see e.g., Para [83][134] – “ In some embodiments, the one or more comparison metrics may be configured to determine how related data included in one or more data subsets may be to data included in one or more other data subsets. In some embodiments, the one or more comparison metrics may include mutual information analyses, the Pearson correlation coefficient, chi2, Analysis of Variance (ANOVA), F-Value, and other comparison metrics that may be configured to compare data included in the data subsets stored in the enhanced dataset.”). With respect to dependent claim 4, the modified Liu teaches generating the filtered general dataset by filtering the general dataset based on the one or more filters comprises: generating the one or more filters for the one or more classifier metrics, each filter corresponding to a respective classifier metric and indicative of a threshold score assigned for the respective classifier metric; and filtering the general dataset based on the one or more filters (see e.g., Para [136] – “one or more data subset pairs may be selected based on the one or more pre-processing operations. In some embodiments, one or more data subset pairs may be selected based on a comparison metric score satisfying a particular threshold. In some embodiments, the threshold may include selecting one or more data subset pairs based on a percentile corresponding to the comparison metric scores. In some embodiments, a percentage of data subset pairs may be selected in response to the comparison metric scores being in the top 1, 5, 10, 15, 20, 25, 30, or some other percent of comparison metric scores as compared to all of the data subset pairs “). With respect to dependent claim 5, the modified Liu teaches generating a synthetic dataset for refining the small language model comprises: identifying one or more deficit skills in the small language model; determining one or more data formats to address the one or more deficit skills; generating the one or more prompts for generating the one or more data formats; injecting sources of randomization and diversity in the one or more prompts; and generating the synthetic dataset based on the one or more prompts using a generative transformer, the synthetic dataset including the one or more data formats (see e.g., Para [80]-“ the language model 106 may include a pre-trained large language model which may include language models such as Generative Pre-Trained Transformer 3 (“GPT3”), Bidirectional Encoder Representations from Transformers (“BERT”), Robustly Optimized Bidirectional Encoder Representations from Transformers (“ROBERTa”), Text-to-Text Transfer Transformer (“T5”), and other language models designed to receive the question synthesized and provide an answer 108 that may include one or more operations that may be performed using data stored in data subsets in the dataset 102.” The examiner notes that Liu does not expressly show the above process for generating synthetic dataset. However, Liu does not limit any process. The above process for generating synthetic data are well known in the art). With respect to dependent claim 6, the modified Liu teaches the one or more deficit skills include any skill or topic for boosting the capability of the small language model (see e.g., Para [24][49] – “one or more other automated machine learning algorithms and/or systems may have difficulty identifying relevant features, comparing those relevant features, and generating and/or synthesizing features that may improve the dataset via feature engineering.” “the target variable may include one or more tasks that one or more machine learning models (e.g., machine learning model 114) may be trained to perform. “). With respect to dependent claim 7, the modified Liu teaches the generative transformer is a multimodal large language model (see e.g., Para [80] – “the language model 106 may include a pre-trained large language model which may include language models such as Generative Pre-Trained Transformer 3 (“GPT3”), Bidirectional Encoder Representations from Transformers (“BERT”), Robustly Optimized Bidirectional Encoder Representations from Transformers (“ROBERTa”), Text-to-Text Transfer Transformer (“T5”), and other language models designed to receive the question synthesized and provide an answer 108 that may include one or more operations that may be performed using data stored in data subsets in the dataset 102. “). With respect to dependent claim 8, the modified Liu teaches prior to training the small language model with the filtered general dataset, performing a warm start by copying weights from an existing trained model into the small language model (see e.g., Merler Para [4][38][90] – “By pretraining the NAS-selected student architecture, and applying data augmentation techniques, the invention achieves at least a 1% absolute improvement in the accuracy of the pre-trained, hand-crafted student presented in the conventional techniques, and a 15% improvement in accuracy compared to the randomly initialized base student of conventional techniques. “ Pretrain the NAS-selected student model before distillation). Claim 9 is rejected for the similar reasons discussed above with respect to claim 1. Claim 10 is rejected for the similar reasons discussed above with respect to claim 2. Claim 11 is rejected for the similar reasons discussed above with respect to claim 3. Claim 12 is rejected for the similar reasons discussed above with respect to claim 4. Claim 13 is rejected for the similar reasons discussed above with respect to claim 5. Claim 14 is rejected for the similar reasons discussed above with respect to claim 6. Claim 15 is rejected for the similar reasons discussed above with respect to claim 8. Claim 16 is rejected for the similar reasons discussed above with respect to claim 1. Claim 17 is rejected for the similar reasons discussed above with respect to claim 2. Claim 18 is rejected for the similar reasons discussed above with respect to claim 4. Claim 19 is rejected for the similar reasons discussed above with respect to claim 6. Claim 20 is rejected for the similar reasons discussed above with respect to claim 8. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to PEIYONG WENG whose telephone number is (571)270-1660. The examiner can normally be reached on Mon.-Fri. 8 am to 5 pm. If attempts to reach the examiner by telephone are unsuccessful, the examiner's supervisor, Matthew Ell, can be reached on (571) 270-3264. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://portal.uspto.gov/external/portal. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). /PEI YONG WENG/Primary Examiner, Art Unit 2141
Read full office action

Prosecution Timeline

Jun 04, 2024
Application Filed
Aug 12, 2026
Non-Final Rejection mailed — §101, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12722290
CROSS-DOMAIN IMITATION LEARNING USING GOAL CONDITIONED POLICIES
3y 5m to grant Granted Sep 01, 2026
Patent 12718006
METHODS, APPARATUS AND SYSTEMS FOR ANNOTATION OF TEXT DOCUMENTS
2y 3m to grant Granted Aug 25, 2026
Patent 12717464
Systems, methods, and user interfaces for editing digital assets
1y 4m to grant Granted Aug 25, 2026
Patent 12711424
SYSTEMS AND METHODS FOR DETERMINATION, DESCRIPTION, AND USE OF FEATURE SETS FOR MACHINE LEARNING CLASSIFICATION SYSTEMS, INCLUDING ELECTRONIC MESSAGING SYSTEMS EMPLOYING MACHINE LEARNING CLASSIFICATION
3y 6m to grant Granted Aug 18, 2026
Patent 12699878
GENERATING IMPLICIT PLANS FOR ACCOMPLISHING GOALS IN AN ENVIRONMENT USING ATTENTION OPERATIONS OVER PLANNING EMBEDDINGS
4y 0m to grant Granted Aug 04, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
79%
Grant Probability
99%
With Interview (+22.8%)
3y 1m (~9m remaining)
Median Time to Grant
Low
PTA Risk
Based on 647 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month