Prosecution Insights
Last updated: October 04, 2026
Application No. 18/302,180

SYSTEM AND METHOD FOR UPDATING LANGUAGE MODELS

Final Rejection §103
Filed
Apr 18, 2023
Priority
Jan 17, 2023 — CN 202310070814.1
Examiner
HUTCHESON, CODY DOUGLAS
Art Unit
2659
Tech Center
2600 — Communications
Assignee
Via Technologies Inc.
OA Round
4 (Final)
62%
Grant Probability
Moderate
5-6
OA Rounds
0m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 62% of resolved cases
62%
Career Allowance Rate
20 granted / 32 resolved
+0.5% vs TC avg
Strong +38% interview lift
Without
With
+37.5%
Interview Lift
resolved cases with interview
Typical timeline
2y 9m
Avg Prosecution
28 currently pending
Career history
66
Total Applications
across all art units

Statute-Specific Performance

§101
31.4%
-8.6% vs TC avg
§103
45.3%
+5.3% vs TC avg
§102
14.1%
-25.9% vs TC avg
§112
5.4%
-34.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 32 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Arguments 1. Regarding the rejection under 35 U.S.C. § 101, Applicant has amended the independent claims to incorporate language which integrate the abstract ideas into a practical application. Accordingly, the rejection is withdrawn. 2. Regarding the rejections under 35 U.S.C. § 103, Applicant’s arguments have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. 3. Claims 1-2, 4-5, 7-9, 12-13, 15-16, and 18-20 are rejected under 35 U.S.C. 103 as being unpatentable over Rao et al. (US 2014/0214419 A1, hereinafter Rao) in view of Burger et al. (US 2020/0265301 A1, hereinafter Burger) and further in view of Itoh et al. (US 2015/0294665 A1, hereinafter Itoh). Regarding claim 1, Rao discloses a data-storage module, implemented using a memory (Fig. 12, “Classifying process module 1218” in “Memory 1206”), configured to store multiple pieces of corpus data corresponding to multiple categories (Fig. 6, Output of “Classifying processing module 602” is classified corpuses corresponding to category 1-m; para. 0073 “Classifying processing module 602, configured to carry out the corpus classification calculation for the raw corpus so as to obtain a different categories of more than one classifying corpus.”; Corpus data is stored as “Classifying process module” resides in memory (Fig. 12, 1216)); a data-update module, implemented using a processing device (Fig. 10, “Classifying process module 1216” in “Memory 1206” run on CPU 1202 ), configured to store a piece of new corpus data into the data-storage module (Fig. 6, “Classifying processing module 602” takes as input new corpus data “Raw corpus” and classifies the data first so that it can then be stored in one of the classifying corpuses; Fig. 2, Step 201; para. 0037 “Step 201, carry out the corpus classification calculation for the raw corpus so as to obtain different categories of more than one classifying corpus.”), wherein the piece of new corpus data corresponds to one of the categories (New corpus data “Raw corpus” is corresponded into one of m categories; para. 0037 “…For example, the mentioned classifying corpus can be divided into many types, such as person name, place name, computer term, medical terminology, etc. For example, "isatis root" belongs to the classification of medical terminology. A term may belong to multi-classification.”); a model-building module, implemented using the processing device (Fig. 12, system contains a model-building module comprised of two modules: “Classifying language model training module 1250” and “Primary language model training module 1216” run on CPU 1202), configured to train a plurality of classified language models (Fig. 6 “Classifying language model training module 603” takes as input the m different classified corpuses, and builds m different “Classifying language model(s)”; Fig. 2, step 202; para. 0038 “Step 202, carry out a language model training calculation for every mentioned classifying corpus to obtain more than one corresponding classifying language models.”), and update one of the classified language models based on the piece of new corpus data stored in the data-storage module (Fig. 6, “Classifying language model” 1-m are updated using new corpus data that has been stored (using “Raw corpus” data that has been stored in a “Classifying corpus” 1-m))… wherein the classified language model updated corresponds to the category that corresponds to the piece of new corpus data (Each “Classifying language model” 1-m is trained using the corresponding “Classifying corpus” 1-m; Fig. 2, step 202; para. 0038 “Step 202, carry out a language model training calculation for every mentioned classifying corpus to obtain more than one corresponding classifying language models.”). Rao does not specifically disclose a data collection module, implemented by a client device, configured to record sentences that are unrecognizable by the client device through speech recognition technologies, and convert the sentences into a piece of new corpus data, wherein in response to an amount of new corpus data accumulated exceeding a threshold, the data-collection module uploads the accumulated new corpus data to a backend server for updating the classified language model, wherein responsive to the classified language model being updated, the client device downloads the updated classified language model from the backend server; wherein the data-collection module is executed by the client device, and the data-update module, the data-storage module, and the model-building module are executed by the backend server… Burger teaches a data collection module, implemented by a client device (Fig. 1, 120), configured to record sentences that are unrecognizable by the client device through speech recognition technologies (para. 0033 “The input data to the machine learning tool 170 can be spoken speech, images, time-series data such as temperatures from a temperature sensor, and so forth.”; para. 0048 “Examples of suitable applications for such neural network implementations include, but are not limited to: performing image recognition, performing speech recognition….”; para. 0080 “The quality analyzer 535 can use various techniques to determine the quality of the results from the machine learning tool 510 for a given set of input data. As one example, a misclassification by the machine learning tool 510 can indicate that the quality of the results is poor…The quality analyzer 535 can mark the input data as data that was misclassified and the misclassified input data can be uploaded (with or without a correct label) using the upload logic 540 and server interface 542 to the server computer 110. The uploaded input data can be used to incrementally train the machine learning tool so that the model parameters 534 can be adjusted based on the new training data.”), and convert the sentences into a piece of new corpus data (para. 0033 “The input data to the machine learning tool 170 can be spoken speech, images, time-series data such as temperatures from a temperature sensor, and so forth.”; para. 0080 “The uploaded input data can be used to incrementally train the machine learning tool so that the model parameters 534 can be adjusted based on the new training data.”), wherein in response to an amount of new corpus data accumulated exceeding a threshold (para. 0083 “As another example, the upload logic 540 can upload the data from the collected training data set 550 after a given amount of data has been collected…”), the data-collection module uploads the accumulated new corpus data to a backend server for updating the classified language model (para. 0080 “The quality analyzer 535 can mark the input data as data that was misclassified and the misclassified input data can be uploaded (with or without a correct label) using the upload logic 540 and server interface 542 to the server computer 110…”; Rao teaches classified language model (see above claim mapping)), wherein responsive to the classified language model being updated, the client device downloads the updated classified language model from the backend server (para. 0080 “For example, the server computer can perform the incremental training using the uploaded input data to generate updated model parameters that can be redistributed to the client computing device 500.”; para. 0081 “After incremental training, the updated model parameters can be downloaded to the model parameters 534. The updated model parameters 534 can improve the accuracy of the machine learning tool 510.”; Rao teaches classified language model (see above claim mapping)); wherein the data-collection module is executed by the client device (Fig. 5, client device 500 collected training data 550; para. 0083 “When low quality results (e.g., the input data was misclassified or the output(s) have high perplexity) are identified, the input data corresponding to the low-quality results can optionally be stored in the collected training data set 550. The collected training data set 550 can be stored in the memory 526 and/or storage 528. The upload logic 540 can manage uploading data from the collected training data set 550.”), and the data-update module, the data-storage module, and the model-building module are executed by the backend server…(server updates models (see Fig. 4, 450); server stores data (see Fig. 4, 442); server builds models via training (see Fig. 4, 440)). Rao and Burger are considered to be analogous to the claimed invention as they both are in the same field of updating machine learning models. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Rao to incorporate the teachings of Burger in order to have a data collection module, implemented by a client device, configured to record sentences that are unrecognizable by the client device through speech recognition technologies, and convert the sentences into a piece of new corpus data, wherein in response to an amount of new corpus data accumulated exceeding a threshold, the data-collection module uploads the accumulated new corpus data to a backend server for updating the classified language model, wherein responsive to the classified language model being updated, the client device downloads the updated classified language model from the backend server; wherein the data-collection module is executed by the client device, and the data-update module, the data-storage module and the model-building module are executed by the backend server. Doing so would beneficial, as this would collect new data which can be used to further update the model to improve accuracy with respect to the new data while also reducing retraining costs and communication workloads between edge and server devices (Burger, para. 0023-0024). Rao in view of Burger does not specifically disclose wherein the model-building model updates the classified language model by updating probability scores only between words that appear in the piece of new corpus data, and does not update probability scores between words that do not appear in the piece of new corpus data, wherein the probability scores are calculated based on occurrence frequencies of word sequences in the piece of new corpus data. Itoh teaches wherein the model-building model (unsupervised training system 202) updates the classified language model (baseline model: para. 0047 “The corpus A302 is a baseline corpus used to build a part that forms the basis for the N-gram language model. As an example, the corpus A302 may be a corpus having a domain and a style consistent with a target application.”) by updating probability scores only between words that appear in the piece of new corpus data, and does not update probability scores between words that do not appear in the piece of new corpus data (model updated with recognition results of speech data: para. 0048 “The corpus B304 is a corpus composed of recognition results of automatic speech recognition of speech data without manual intervention.”; para. 0054 “The language model training section 308 uses both the corpus A302 and the corpus B304 as training text to build an N-gram language model.”; may use only new corpus data B304 to add on to existing n-gram language model by only adding probabilities of new corpus data n-grams to the original model, instead of building from scratch (updating all probabilities): para. 0057 “Note that the use of the corpus A302 is optional, and N-gram entries may be selected without using the corpus A302.”; para. 0062 “The language model training section 308 may add, to the base N-gram language model, the selected one or more N-gram entries and their probabilities obtained as a result of the training, or build the N-gram language model from scratch.”), wherein the probability scores are calculated based on occurrence frequencies of word sequences in the piece of new corpus data (probabilities for N-grams based on occurrence frequencies of word sequences: para. 0061 “The probability calculation section 312 uses all the recognition results included in the corpus A302 and the corpus B304 to train the N-gram language model about one or more N-gram entries selected by the selection section 310, or when only the corpus B is used as training data, the probability calculation section 312 uses all the recognition results included in the corpus B to train the N-gram language model about the selected one or more N-gram entries.”; para. 0040 “The conditional probabiliites of the above mathematical expression (2) can be determined by using maximum likelihood estimation from the number of appearances of a string of N words and a string of (N-1) words appears in a corpus…”, see Eqs. (1)-(3)). Rao, Burger, and Itoh are considered to be analogous to the claimed invention as they are all in the same field of updating machine learning models. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Rao in view of Burger to incorporate the teachings of Itoh in order to have the model building model update the classified language model by updating probability scores only between words that appear in the piece of new corpus data, and does not update probability scores between words that do not appear in the piece of new corpus data, wherein the probability scores are calculated based on occurrence frequencies of word sequences in the piece of new corpus data. Doing so would beneficial, as this would allow for necessary updates to language models to be obtained without the need to necessarily train from scratch (Itoh, para. 0004, para. 0062). Regarding claim 2, Rao in view of Burger in further view of Itoh discloses wherein the classified language model uses n-grams to calculate probabilities between words (Itoh, Equations (1)-(3); para. 0061 “The probability calculation method is as described in the outline of the N-gram language model.”). Rao, Burger, and Itoh are considered to be analogous to the claimed invention as they are all in the same field of updating machine learning models models. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Rao in view of Burger to incorporate the teachings of Itoh in order to use n-grams to calculate probabilities between words. Doing so would beneficial, as n-gram language models are simple yet effectively, commonly used in large vocabulary speech recognition tasks (Itoh, para. 0037). Regarding claim 4, Rao in view of Burger in further view of Itoh discloses wherein the model-building module further updates a generic language model based on the piece of new corpus data stored in the data-storage module (Rao, Fig. 10, “Primary Language Model Training Module 1216”; Fig. 6, a generic language model is trained via “Primary language model training module 601” to obtain “Primary language model” using pieces of new corpus data “Raw corpus”; para. 0043 “Step 301, carry out a calculation of the language model training according to the raw corpus to obtain the primary language model. Here, the language model training is the conventional regular language model training.”). Regarding claim 5, Rao in view of Burger in further view of Itoh discloses further comprising a corpus-classification module, implemented using the processing device (Rao, Fig. 12, “Classifying process module 1218” in “Memory 1206” run on CPU 1202), configured to use a classification model to determine the category that corresponds to the piece of new corpus data (Rao, Fig. 7, A category (which “Classifying corpus 1-m” the piece of corpus corresponds to) is determined for each piece of “Raw corpus”, using a classification model “Classifier 704”; para. 0083 “Classifier 704, configured to put the term characteristic after the processing of dimensionality reduction into the classifier for training, output different categories of more than one classifying corpus. In a preferred embodiment, the mentioned classifier is a SVM classifier.”). Regarding claim 7, Rao in view of Burger in further view of Itoh discloses wherein the corpus-classification module extracts a feature vector of the piece of new corpus data (Rao, Fig. 7 discloses structure of the corpus-classification module; “Characteristic extracting module 702” extracts a feature vector of the new corpus data “Raw corpus”, which is then reduced by dimensionality by “Dimensionality reduction module 703”; para. 0081 “Characteristic extracting module 702, configured to use TF-IDF method to extract the term characteristic from the raw corpus”; para. 0082 “Dimensionality reduction module 703, configured that according to the mentioned affinity matrix, use the dimensionality reduction method to process dimension reduction for the extracted term characteristic. In a preferred embodiment, the mentioned dimensionality reduction module is PCA dimensionality reduction module.”), inputs the feature vector into the classification model (Rao, Feature vectors (output of “Dimensionality reduction module 703”) is fed into “Classifier 704”), and determines the category that corresponds to the piece of new corpus data according to a result output by the classification model (Rao, “Classifier 704“ determines a category corresponds to the piece of new corpus data (a category 1-m) according to a result output by classifier (output of SVM); para. 0083 “Classifier 704, configured to put the term characteristic after the processing of dimensionality reduction into the classifier for training, output different categories of more than one classifying corpus. In a preferred embodiment, the mentioned classifier is a SVM classifier.”). Regarding claim 8, Rao in view of Burger in further view of Itoh discloses wherein the corpus-classification module uses a term frequency-inverse document frequency (tf-idf) approach to extract the feature vector from the piece of new corpus data (Rao, “Characteristic extracting module 702” extracts features using tf-idf; para. 0081 “Characteristic extracting module 702, configured to use TF-IDF method to extract the term characteristic from the raw corpus.”). Regarding claim 9, Rao in view of Burger in further view of Itoh discloses wherein the data-storage module further stores the corpus data that correspond to the category as a classified corpus (Rao, Data storage module (“Fig. 12, “Classifying process module 1218” in “Memory 1206”) classifies the data into a plurality of classified corpuses 1-m, which are stored as they reside in memory 1206). Regarding claim 12, Rao discloses A method for updating language models, for user in a computer system (Rao, Abstract), the method comprising: storing the piece of new corpus data into a data-storage module of the computer system (Rao, Fig. 12, “Classifying process module 1218” in “Memory 1206”), wherein the data storage module is used for storing multiple pieces of corpus data corresponding to multiple categories (Rao, Fig. 6, Output of “Classifying processing module 602” is classified corpuses corresponding to category 1-m; para. 0073 “Classifying processing module 602, configured to carry out the corpus classification calculation for the raw corpus so as to obtain a different categories of more than one classifying corpus.”; Corpus data is stored as “Classifying process module” resides in memory (Fig. 12, 1216)), and the piece of new corpus data corresponds to the categories (New corpus data “Raw corpus” is corresponded into one of m categories; para. 0037 “…For example, the mentioned classifying corpus can be divided into many types, such as person name, place name, computer term, medical terminology, etc. For example, "isatis root" belongs to the classification of medical terminology. A term may belong to multi-classification.”); training and updating one of a plurality of classified language models based on the piece of new corpus data stored in the data-storage module (Fig. 6, “Classifying language model” 1-m are updated and trained using new corpus data that has been stored (using “Raw corpus” data that has been stored in a “Classifying corpus” 1-m)), wherein the classified language model updated corresponds to the category that corresponds to the piece of new corpus data (Each “Classifying language model” 1-m is trained using the corresponding “Classifying corpus” 1-m; Fig. 2, step 202; para. 0038 “Step 202, carry out a language model training calculation for every mentioned classifying corpus to obtain more than one corresponding classifying language models.”)…. Rao does not specifically disclose recording sentences that are unrecognizable by a client device through speech recognition technologies, and converting each sentence into a piece of new corpus data; in response to an amount of new corpus data accumulated exceeding a threshold, uploading the accumulated new corpus data from the client device to a backend server… and in response to the classified language model being updated, downloading the updated classified language model from the backend server to the client device… Burger teaches recording sentences that are unrecognizable by the client device through speech recognition technologies (para. 0033 “The input data to the machine learning tool 170 can be spoken speech, images, time-series data such as temperatures from a temperature sensor, and so forth.”; para. 0048 “Examples of suitable applications for such neural network implementations include, but are not limited to: performing image recognition, performing speech recognition….”; para. 0080 “The quality analyzer 535 can use various techniques to determine the quality of the results from the machine learning tool 510 for a given set of input data. As one example, a misclassification by the machine learning tool 510 can indicate that the quality of the results is poor…The quality analyzer 535 can mark the input data as data that was misclassified and the misclassified input data can be uploaded (with or without a correct label) using the upload logic 540 and server interface 542 to the server computer 110. The uploaded input data can be used to incrementally train the machine learning tool so that the model parameters 534 can be adjusted based on the new training data.”), and converting each sentence into a piece of new corpus data (para. 0033 “The input data to the machine learning tool 170 can be spoken speech, images, time-series data such as temperatures from a temperature sensor, and so forth.”; para. 0080 “The uploaded input data can be used to incrementally train the machine learning tool so that the model parameters 534 can be adjusted based on the new training data.”), in response to an amount of new corpus data accumulated exceeding a threshold (para. 0083 “As another example, the upload logic 540 can upload the data from the collected training data set 550 after a given amount of data has been collected…”), uploading the accumulated new corpus data from the client device to a backend server (para. 0080 “The quality analyzer 535 can mark the input data as data that was misclassified and the misclassified input data can be uploaded (with or without a correct label) using the upload logic 540 and server interface 542 to the server computer 110…”), and in response to the classified language model being updated, downloading the updated classified language model from the backend server to the client device (para. 0080 “For example, the server computer can perform the incremental training using the uploaded input data to generate updated model parameters that can be redistributed to the client computing device 500.”; para. 0081 “After incremental training, the updated model parameters can be downloaded to the model parameters 534. The updated model parameters 534 can improve the accuracy of the machine learning tool 510.”; Rao discloses classified language model (see above claim mapping)); Rao and Burger are considered to be analogous to the claimed invention as they both are in the same field of updating machine learning models. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Rao to incorporate the teachings of Burger in order to have record sentences that are unrecognizable by the client device through speech recognition technologies, and convert the sentences into a piece of new corpus data, in response to an amount of new corpus data accumulated exceeding a threshold, uploading the accumulated new corpus data to a backend server for updating the classified language model, and responsive to the classified language model being updated, the client device downloads the updated classified language model from the backend server. Doing so would beneficial, as this would collect new data which can be used to further update the model to improve accuracy with respect to the new data while also reducing retraining costs and communication workloads between edge and server devices (Burger, para. 0023-0024). Rao in view of Burger does not specifically disclose wherein the model-building module updates the classified language model by only updating probability scores between words in the piece of new corpus data, and not updating probability scores of words that are not in the piece of new corpus data, wherein the probability scores are calculated based on occurrence frequency of word sequences in the piece of new corpus data. Itoh teaches wherein the model-building model (unsupervised training system 202) updates the classified language model (baseline model: para. 0047 “The corpus A302 is a baseline corpus used to build a part that forms the basis for the N-gram language model. As an example, the corpus A302 may be a corpus having a domain and a style consistent with a target application.”) by only updating probability scores between words in the piece of new corpus data, and not updating probability scores of words that are not in the piece of new corpus data (model updated with recognition results of speech data: para. 0048 “The corpus B304 is a corpus composed of recognition results of automatic speech recognition of speech data without manual intervention.”; para. 0054 “The language model training section 308 uses both the corpus A302 and the corpus B304 as training text to build an N-gram language model.”; may use only new corpus data B304 to add on to existing n-gram language model by only adding probabilities of new corpus data n-grams to the original model, instead of building from scratch (updating all probabilities): para. 0057 “Note that the use of the corpus A302 is optional, and N-gram entries may be selected without using the corpus A302.”; para. 0062 “The language model training section 308 may add, to the base N-gram language model, the selected one or more N-gram entries and their probabilities obtained as a result of the training, or build the N-gram language model from scratch.”), wherein the probability scores are calculated based on occurrence frequencies of word sequences in the piece of new corpus data (probabilities for N-grams based on occurrence frequencies of word sequences: para. 0061 “The probability calculation section 312 uses all the recognition results included in the corpus A302 and the corpus B304 to train the N-gram language model about one or more N-gram entries selected by the selection section 310, or when only the corpus B is used as training data, the probability calculation section 312 uses all the recognition results included in the corpus B to train the N-gram language model about the selected one or more N-gram entries.”; para. 0040 “The conditional probabiliites of the above mathematical expression (2) can be determined by using maximum likelihood estimation from the number of appearances of a string of N words and a string of (N-1) words appears in a corpus…”, see Eqs. (1)-(3)). Rao, Burger, and Itoh are considered to be analogous to the claimed invention as they are all in the same field of updating language models. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Rao in view of Burger to incorporate the teachings of Itoh in order to have the model building model update the classified language model by updating probability scores only between words that appear in the piece of new corpus data, and does not update probability scores between words that do not appear in the piece of new corpus data, wherein the probability scores are calculated based on occurrence frequencies of word sequences in the piece of new corpus data. Doing so would beneficial, as this would allow for necessary updates to language models to be obtained without the need to necessarily train from scratch (para. 0004, para. 0062). Regarding claim 13, claim 13 is rejected for analogous reasons to claim 2. Regarding claim 15, claim 15 is rejected for analogous reasons to claim 4. Regarding claim 16, claim 16 is rejected for analogous reasons to claim 5. Regarding claim 18, claim 18 is rejected for analogous reasons to claim 7. Regarding claim 19, claim 19 is rejected for analogous reasons to claim 8. Regarding claim 20, claim 20 is rejected for analogous reasons to claim 9. 4. Claims 6 and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Rao in view of Burger and Itoh, and further in view of Conneau et al. (NPL Very Deep Convolutional Networks for Text Classification, hereinafter Conneau). Regarding claim 6, Rao in view of Burger and Itoh does not specifically disclose wherein the classification model is a Fully-Connected Neural Network. Conneau teaches wherein the classification model is a Fully-Connected Neural Network (a text classification system is taught (pg. 4, Fig. 1), which is used to classify text in one or more categories (see pg. 6, Table 3, e.g. classifying DBPedia text dataset into one of 14 classes). The system uses a classification model which is a fully-connected neural network (pg. 4, Fig. 1, top 3 layers denoted “fc”; pg. 4, para 3 “The 512 x k resulting features are transformed into a single vector which is the input to a three layer fully connected classifier with ReLU hidden units and softmax outputs. The number of output neurons depends on the classification task…”)) Rao, Burger, Itoh, and Conneau are considered to be analogous to the claimed invention as they are in the same field of classifying text into corpus categories. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the teachings of Conneau in order to use a fully-connected neural network as the classification model. Doing so would beneficial, as fully-connected neural networks are structure agnostic, allowing for greater model input flexibility (NPL Ramsundar et al., TensorFlow for Deep Learning, Chapter 4, first paragraph). Regarding claim 17, claim 17 is rejected for analogous reasons to claim 6. 5. Claims 10 and 21 are rejected under 35 U.S.C. 103 as being unpatentable over Rao in view of Burger and Itoh, and further in view of Biadsy et al. (US 2018/0053502 A1, hereinafter Biadsy). Regarding claim 10, Rao in view of Burger and Itoh does not specifically disclose wherein the data-storage module further stores a category label of the category that corresponds to each piece of corpus data. Biadsy teaches wherein the data-storage module further stores a category label of the category that corresponds to each piece of corpus data (Biadsy, Fig. 8, for “Training Example”, there are accompanying labels “Application ID: Maps”, and “Location: New York City, NY”, both of which are labels indicating respective domains (see Fig. 7, domains 710c and 710a respectively)). Rao, Burger, Itoh, and Biadsy are considered to be analogous to the claimed invention as they are all in the same field of updating language models. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the teachings of Biadsy in order to store the category label that corresponds to each piece of corpus data. Doing so would beneficial, as the category label could be later used as a feature for indicating which classified language models to update for a particular training example. Regarding claim 21, claim 21 is rejected for analogous reasons to claim 10. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure: Breckenridge et al. (US 8,595,154 B2): uploading training data from a client device to a server in batches for retraining (Col. 4, Lines 36-49) Kwon et al. (US 2021/0287663 A1): obtain feedback for speech recognition model, determine whether or not to update model based on feedback data (Fig. 2) Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to CODY DOUGLAS HUTCHESON whose telephone number is (703)756-1601. The examiner can normally be reached M-F 8:00AM-5:00PM EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Pierre-Louis Desir can be reached at (571)-272-7799. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /CODY DOUGLAS HUTCHESON/Examiner, Art Unit 2659 /PIERRE LOUIS DESIR/Supervisory Patent Examiner, Art Unit 2659
Read full office action

Prosecution Timeline

Show 8 earlier events
Oct 27, 2025
Request for Continued Examination
Nov 05, 2025
Response after Non-Final Action
Feb 11, 2026
Non-Final Rejection mailed — §103
Apr 09, 2026
Interview Requested
Apr 28, 2026
Applicant Interview (Telephonic)
Apr 28, 2026
Examiner Interview Summary
May 10, 2026
Response Filed
Aug 05, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12743450
QUERY FORMATTING SYSTEM, QUERY FORMATTING METHOD, AND INFORMATION STORAGE MEDIUM
3y 6m to grant Granted Sep 22, 2026
Patent 12737563
REGIONAL SIGN LANGUAGE TRANSLATION
3y 5m to grant Granted Sep 15, 2026
Patent 12718828
AUDIO SOURCE CLASSIFICATION FOR HANDSFREE COMMUNICATIONS
3y 5m to grant Granted Aug 25, 2026
Patent 12664970
SPEECH TRANSLATION WITH PERFORMANCE CHARACTERISTICS
3y 2m to grant Granted Jun 23, 2026
Patent 12626715
ROLE SEPARATION METHOD, ELECTRONIC DEVICE, AND COMPUTER STORAGE MEDIUM
3y 4m to grant Granted May 12, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

5-6
Expected OA Rounds
62%
Grant Probability
99%
With Interview (+37.5%)
2y 9m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 32 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month