Prosecution Insights
Last updated: October 01, 2026
Application No. 18/835,869

Continuous Training of Machine Learning Models on Changing Data

Non-Final OA §101§102§103§112
Filed
Aug 05, 2024
Priority
Feb 03, 2022 — nonprovisional of PCTUS2022015035
Examiner
BOSTWICK, SIDNEY VINCENT
Art Unit
Tech Center
Assignee
Google LLC
OA Round
1 (Non-Final)
51%
Grant Probability
Moderate
1-2
OA Rounds
2y 3m
Est. Remaining
86%
With Interview

Examiner Intelligence

Grants 51% of resolved cases
51%
Career Allowance Rate
78 granted / 152 resolved
-8.7% vs TC avg
Strong +35% interview lift
Without
With
+35.1%
Interview Lift
resolved cases with interview
Typical timeline
4y 5m
Avg Prosecution
40 currently pending
Career history
216
Total Applications
across all art units

Statute-Specific Performance

§101
24.6%
-15.4% vs TC avg
§103
46.5%
+6.5% vs TC avg
§102
4.6%
-35.4% vs TC avg
§112
24.0%
-16.0% vs TC avg
Black line = Tech Center average estimate • Based on career data from 152 resolved cases

Office Action

§101 §102 §103 §112
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Detailed Action This action is in response to the claims filed 8/5/2024: Claims 1 – 20 are pending. Claims 1, 16, and 20 are independent. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 5, 6, 18, 12, 14, and 18 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Regarding claim 5, "subject to user-defined handling obligations" is indefinite. "Subject to user-defined handling obligations" fails to convey what type of handling constitutes an obligation, what entity is subject to the obligation, or what objective criterion distinguishes content subject to such an obligation from content that is not. For these reasons one of ordinary skill in the art could not reasonably determine the scope of the claim. In the interest of further examination any user input is interpreted as user-defined handling obligations. Regarding claims 6 and 18, "sampling [...] to generate the current set of training data comprises [...] deleting that training data" is logically indefinite. It would be unclear to one of ordinary skill in the art how deleting training data could generate the same training data. Regarding claim 12, "the model" lacks antecedent basis. Claim 9 recites both "the updated model and "one or more other machine learning models" such that it's unclear which model is "the model." In the interest of further examination the claim 12 is interpreted as "a degree to which a distribution fits to the performance of the one or more other machine learning models." Regarding claim 12, "a degree to which a distribution fits to the performance of the model" is indefinite. This is literally a term of degree without a relative basis for comparison. In the interest of further examination any test is interpreted as satisfying a degree to which a distribution fits the performance of the model. Regarding claim 14, "when the performance of the updated model relative to the current set of testing data deviates from the respective performance of the one or more other machine learning models on the current set of testing data" is indefinite. This is further complicated by the optional one to one or one to many relationship in the claim. There are multiple contradicting interpretations of performance deviation such that one of ordinary skill in the art could not reasonably determine the scope of the claim. In the interest of further examination. Claim Rejections - 35 USC § 101 101 Rejection 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-20 are rejected under 35 USC § 101 because the claimed invention is directed to non-statutory subject matter. Regarding Claim 1: Claim 1 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Step 1 Analysis: Claim 1 is directed to a method, which is directed to a process, one of the statutory categories. Step 2A Prong One Analysis: Claim 1 under its broadest reasonable interpretation is a series of mental processes. For example, but for the generic computer components language, the above limitations in the context of this claim encompass machine learning processing, including the following: for each of one or more update iterations: sampling, by a computing system comprising one or more computing devices, from a pool of data associated with one or more ancillary systems to generate a current set of training data (observation, evaluation, and judgement), evaluating, by the computing system, a performance of the updated model relative to a current set of testing data; (observation, evaluation, and judgement) performing, by the computing system, a comparison of the performance of the updated model relative to the current set of testing data with a respective performance of one or more other machine learning models on the current set of testing data or one or more past sets of testing data (observation, evaluation, and judgement) selecting, by the computing system, either the updated model or one of the one or more other machine learning models for deployment based on the comparison of the performance of the updated model relative to the current set of testing data with the respective performance of the one or more other machine learning models on the current set of testing data or the one or more past sets of testing data (observation, evaluation, and judgement) Therefore, claim 1 recites an abstract idea which is a judicial exception. Step 2A Prong Two Analysis: Claim 1 recites additional elements “training, by the computing system, a machine learning model on the current set of training data to generate an updated model” and “by the computing system”. However, these additional features are computer components recited at a high-level of generality, such that they amount to no more than mere instructions to apply the judicial exception using a generic computer component. An additional element that merely recites the words “apply it” (or an equivalent) with the judicial exception, or merely includes instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea, does not integrate the judicial exception into a practical application (See MPEP 2106.05(f)). Therefore, claim 1 is directed to a judicial exception. Step 2B Analysis: Claim 1 does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to the lack of integration of the abstract idea into a practical application, the additional elements recited in claim 1 amount to no more than mere instructions to apply the judicial exception using a generic computer component. For the reasons above, claim 1 is rejected as being directed to non-patentable subject matter under §101. This rejection applies equally to independent claims 16 and 20, which recite a system and a non-transitory computer readable medium, respectively, as well as to dependent claims 2-15 and 17-19. Independent claim 16 recites additional instructions to apply the judicial exception using generic computer components “A computing system for training machine learning models on changing data, the computing system comprising: “one or more processors; and one or more non-transitory computer-readable media that collectively store instructions that, when executed by the one or more processors, cause the computing system to perform operations” (MPEP 2106.05(f)). Independent claim 20 recites additional instructions to apply the judicial exception using generic computer components “One or more non-transitory computer-readable media that collectively store instructions that, when executed by a computing system, cause the computing system to perform operations for each of one or more update iterations, the operations comprising”. The additional limitations of the dependent claims are addressed briefly below: Dependent claim 2 recites additional observation, evaluation, and judgement “for each of the one or more update iterations, sampling, […], from the pool of data associated with the one or more ancillary systems to generate the current set of testing data” Dependent claim 3 recites additional observation, evaluation, and judgement “the current set of testing data comprises a fixed set of testing data” Dependent claim 4 recites additional insignificant extra-solution activity of gathering and outputting data (See MPEP 2106.05(g)) “the one or more other machine learning models comprise previous checkpoints of the machine learning model” which is well-understood, routine, and conventional in the art (See MPEP 2106.05(d)(II)(i)) (a previous checkpoint is just a stored data). Dependent claim 5 recites additional observation, evaluation, and judgement “the pool of data associated with the one or more ancillary systems comprises user-generated content that is subject to user-defined handling obligations” Dependent claim 6 recites additional insignificant extra-solution activity “sampling, […], from the pool of data associated with the one or more ancillary systems to generate the current set of training data comprises: associating, by the computing system, a wipeout-compliant flag with the current set of training data, wherein the wipeout-compliant flag causes deletion of the current set of training data upon occurrence of a condition;” as well as additional instructions to apply the judicial exception using generic computer components “deleting, by the computing system, the current set of training data upon occurrence of the condition” Dependent claim 7 recites additional observation, evaluation, and judgement “sampling, […], from the pool of data associated with the one or more ancillary systems to generate the current set of training data comprises: randomly sampling, by the computing system, from the pool of data associated with the one or more ancillary systems to generate the current set of training data.” Dependent claim 8 recites additional observation, evaluation, and judgement “sampling, […], from the pool of data associated with the one or more ancillary systems to generate the current set of training data comprises: sampling, […] from the pool of data associated with the one or more ancillary systems, only data examples that have been newly generated within a defined period of time” Dependent claim 9 recites additional observation, evaluation, and judgement “performing, […], the comparison of the performance of the updated model relative to the current set of testing data with the respective performance of the one or more other machine learning models on the current set of testing data or the one or more past sets of testing data comprises: determining, by the computing system, a first set of statistical tests for the performance of the updated model relative to the current set of testing data; determining, by the computing system, a second set of statistical tests for the respective performance of the one or more other machine learning models on the current set of testing data; and performing, […], a comparison of the first set of statistical tests and the second set of statistical tests” Dependent claim 10 recites additional observation, evaluation, and judgement “performing, […], the comparison of the performance of the updated model relative to the current set of testing data with the respective performance of the one or more other machine learning models on the current set of testing data or the one or more past sets of testing data comprises: determining, […], a first set of statistical tests for the performance of the updated model relative to the current set of testing data; determining, […], a second set of statistical tests for the respective performance of the one or more other machine learning models on the one or more past sets of testing data; and performing, […], a comparison of the first set of statistical tests and the second set of statistical tests” Dependent claim 11 recites additional observation, evaluation, and judgement “the first set of statistical tests and the second set of statistical tests each comprise a set of error bounds” Dependent claim 12 recites additional observation, evaluation, and judgement “wherein the first set of statistical tests and the second set of statistical tests each comprise one or more of: mean or standard deviation; min or max score; skew; quartile ranges; or a degree to which a distribution fits to the performance of the model” Dependent claim 13 recites additional observation, evaluation, and judgement “performing, […], the comparison of the first set of statistical tests and the second set of statistical tests comprises normalizing at least the first set of statistical tests based on feature values associated with the current set of testing data” Dependent claim 14 recites additional observation, evaluation, and judgement “providing, […], an automated alert when the performance of the updated model relative to the current set of testing data deviates from the respective performance of the one or more other machine learning models on the current set of testing data” Dependent claim 15 recites additional observation, evaluation, and judgement “sampling, […], from the pool of data associated with the one or more ancillary systems to generate the current set of training data comprises: accessing, by the computing system, the one or more ancillary systems using one or more application programming interfaces” Dependent claim 17 recites additional observation, evaluation, and judgement “selecting, […], either the updated model or one of the one or more other machine learning models for deployment based on the comparison of the performance of the updated model relative to the current set of testing data with the respective performance of the one or more other machine learning models on the current set of testing data” Dependent claim 18 recites additional observation, evaluation, and judgement “sampling, […], from the pool of data associated with the one or more ancillary systems to generate the current set of training data comprises: associating, […], a wipeout-compliant flag with the current set of training data, wherein the wipeout-compliant flag causes deletion of the current set of training data upon occurrence of a condition;” and additional instructions to apply the judicial exception using generic computer components “deleting, by the computing system, the current set of training data upon occurrence of the condition” Dependent claim 19 recites additional observation, evaluation, and judgement “performing, […], the comparison of the performance of the updated model relative to the current set of testing data with the respective performance of the one or more other machine learning models on the current set of testing data or the one or more past sets of testing data comprises: determining, […], a first set of statistical tests for the performance of the updated model relative to the current set of testing data; determining, […], a second set of statistical tests for the respective performance of the one or more other machine learning models on the current set of testing data; and performing, […] a comparison of the first set of statistical tests and the second set of statistical tests” Therefore, when considering the elements separately and in combination, they do not add significantly more to the inventive concept. Accordingly, claims 1-20 are rejected under 35 U.S.C. § 101. Claim Rejections - 35 USC § 102 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. Claims 1-3, 5, 9, 10, 12, 15-17, 19, and 20 are rejected under U.S.C. §102(a)(1) as being unpatentable over the combination of Bonawitz (US20180144265A1). PNG media_image1.png 680 576 media_image1.png Greyscale FIG. 2 of US20180144265A1 PNG media_image2.png 638 676 media_image2.png Greyscale FIG. 2 of US20180144265A1 Regarding claim 1, Bonawitz teaches A computer-implemented method to train machine learning models on changing data, the method comprising:([¶0054] "The user computing device 102 can also include a model trainer 122. The model trainer 122 can train or re-train one or more of the machine-learned models 120 stored at the user computing device 102 using various training or learning techniques, such as, for example, backwards propagation of errors (e.g., truncated backpropagation through time). In particular, the model trainer 122 can train or re-train one or more of the machine-learned models 120 using the locally logged data 119 as training data." [¶0024] "can provide a local, on-device database to which log entries are written") for each of one or more update iterations:([¶0104] "method 300 can be performed iteratively over a number of corresponding first and second portions of ground-truth data." [¶0088] "The various steps of the methods of FIGS. 2-5 can be omitted, rearranged, combined, and/or adapted in various ways without deviating from the scope of the present disclosure") sampling, by a computing system comprising one or more computing devices, from a pool of data associated with one or more ancillary systems to generate a current set of training data;([¶0076] "Each application 1 through N also includes local application data. In particular, each application can provide or otherwise communicate with a local, on-device database in memory to which log entries are written. One database can be used for all applications or different respective databases can be used for each application" [¶0054] " the model trainer 122 can train or re-train one or more of the machine-learned models 120 using the locally logged data 119 as training data" [¶0026] "The user computing device can be instructed to use the entries in the local database (or some subset thereof) as training data for the training algorithm" Bonawitz explicitly discloses using a subset (sample) of the local dataset (pool) for training) training, by the computing system, a machine learning model on the current set of training data to generate an updated model;([¶0033] "rather than receive and/or train a single machine-learned model, the user computing device can receive and/or train a plurality of machine-learned models" See also FIG. 2-4 where FIG. 2 describes train/update step, FIG. 3 elaborates and provides iterative evaluation and update, and FIG. 4 extends that evaluation to multiple models) evaluating, by the computing system, a performance of the updated model relative to a current set of testing data;([¶0026] "the trained model can be evaluated prior to activation." See also FIG. 2 which places activation after obtain/train and evaluate) performing, by the computing system, a comparison of the performance of the updated model relative to the current set of testing data with a respective performance of one or more other machine learning models on the current set of testing data or one or more past sets of testing data; and([Abstract] "the user computing device can evaluate a plurality of machine-learned models against locally stored data" [¶0021] "the user computing device can obtain a plurality of machine-learned models and can evaluate at least one performance metric for each of the plurality of machine-learned models relative to the data that is stored locally at the user computing device" [¶0025] "This process also yields a history of multiple versions of the model, and each of these could be updated on each new dataset to evaluate the best candidate.") selecting, by the computing system, either the updated model or one of the one or more other machine learning models for deployment based on the comparison of the performance of the updated model relative to the current set of testing data with the respective performance of the one or more other machine learning models on the current set of testing data or the one or more past sets of testing data.([¶0025] "This process also yields a history of multiple versions of the model, and each of these could be updated on each new dataset to evaluate the best candidate." [¶0035] "The user computing device and/or server computing device can select one of the machine-learned models based at least in part on the performance metrics. For example, the model with the best performance metric(s) can be selected for activation and use at the user computing device." [¶0120] "the server computing device can instruct the user computing device to activate and use the selected at least one machine-learned model."). Regarding claim 2, Bonawitz teaches The computer-implemented method of claim 1, further comprising, for each of the one or more update iterations, sampling, by the computing system, from the pool of data associated with the one or more ancillary systems to generate the current set of testing data.(Bonawitz [¶0104] "method 300 can be performed iteratively over a number of corresponding first and second portions of ground-truth data." [¶0088] "The various steps of the methods of FIGS. 2-5 can be omitted, rearranged, combined, and/or adapted in various ways without deviating from the scope of the present disclosure" [¶0026] "The user computing device can be instructed to use the entries in the local database (or some subset thereof) as training data for the training algorithm" See also FIG. 2 and 3). Regarding claim 3, Bonawitz teaches The computer-implemented method of claim 1, wherein the current set of testing data comprises a fixed set of testing data.(Bonawitz [¶0060] " the model manager 124 can use historical data 119 that was previously logged at the user computing device 102 to evaluate the at least one performance metric" [¶0034] "The user computing device can evaluate at least one performance metric for each of the plurality of machine-learned models. In particular, the user computing device can evaluate the at least one performance metric for each machine-learned model relative to data that is stored locally at the user computing device. For example, the locally stored data can be newly received data or can be previously logged data, as described above." the locally stored historical/previously logged data correspond to a fixed set of testing data). Regarding claim 5, Bonawitz teaches The computer-implemented method of claim 1, wherein the pool of data associated with the one or more ancillary systems comprises user-generated content that is subject to user-defined handling obligations.(Bonawitz [¶0023] " the data stays with the user, thereby increasing user privacy" [¶0024] "log entries may or may not be expungable by a user"). Regarding claim 9, Bonawitz teaches The computer-implemented method of claim 1, wherein performing, by the computing system, the comparison of the performance of the updated model relative to the current set of testing data with the respective performance of the one or more other machine learning models on the current set of testing data or the one or more past sets of testing data comprises:(Bonawitz [Abstract] "the user computing device can evaluate a plurality of machine-learned models against locally stored data" [¶0021] "the user computing device can obtain a plurality of machine-learned models and can evaluate at least one performance metric for each of the plurality of machine-learned models relative to the data that is stored locally at the user computing device" [¶0025] "This process also yields a history of multiple versions of the model, and each of these could be updated on each new dataset to evaluate the best candidate.") determining, by the computing system, a first set of statistical tests for the performance of the updated model relative to the current set of testing data;(Bonawitz [¶0103] " the performance metric can be an average of the errors over time") determining, by the computing system, a second set of statistical tests for the respective performance of the one or more other machine learning models on the current set of testing data; and(Bonawitz [¶0103] " the performance metric can be an average of the errors over time" [¶0008] "The method includes evaluating, by the user computing device, at least one performance metric for each of the plurality of machine-learned models") performing, by the computing system, a comparison of the first set of statistical tests and the second set of statistical tests.(Bonawitz [¶0025] "This process also yields a history of multiple versions of the model, and each of these could be updated on each new dataset to evaluate the best candidate."). Regarding claim 10, Bonawitz teaches The computer-implemented method of claim 1, wherein performing, by the computing system, the comparison of the performance of the updated model relative to the current set of testing data with the respective performance of the one or more other machine learning models on the current set of testing data or the one or more past sets of testing data comprises: determining, by the computing system, a first set of statistical tests for the performance of the updated model relative to the current set of testing data;(Bonawitz [¶0103] " the performance metric can be an average of the errors over time" [¶0008] "The method includes evaluating, by the user computing device, at least one performance metric for each of the plurality of machine-learned models") determining, by the computing system, a second set of statistical tests for the respective performance of the one or more other machine learning models on the one or more past sets of testing data; and(Bonawitz [¶0103] " the performance metric can be an average of the errors over time" [¶0008] "The method includes evaluating, by the user computing device, at least one performance metric for each of the plurality of machine-learned models") performing, by the computing system, a comparison of the first set of statistical tests and the second set of statistical tests.(Bonawitz [Abstract] "the user computing device can evaluate a plurality of machine-learned models against locally stored data" [¶0021] "the user computing device can obtain a plurality of machine-learned models and can evaluate at least one performance metric for each of the plurality of machine-learned models relative to the data that is stored locally at the user computing device" [¶0025] "This process also yields a history of multiple versions of the model, and each of these could be updated on each new dataset to evaluate the best candidate."). Regarding claim 12, Bonawitz teaches The computer-implemented method of claim 9, wherein the first set of statistical tests and the second set of statistical tests each comprise one or more of: mean or standard deviation; min or max score; skew; quartile ranges; or a degree to which a distribution fits to the performance of the model.(Bonawitz [¶0103] " the performance metric can be an average of the errors over time"). Regarding claim 15, Bonawitz teaches The computer-implemented method of claim 1, wherein sampling, by the computing system, from the pool of data associated with the one or more ancillary systems to generate the current set of training data comprises: accessing, by the computing system, the one or more ancillary systems using one or more application programming interfaces.(Bonawitz [¶0076] "Each application 1 through N also includes local application data. In particular, each application can provide or otherwise communicate with a local, on-device database in memory to which log entries are written"). Regarding claim 16, Bonawitz teaches A computing system for training machine learning models on changing data, the computing system comprising:([¶0054] "The user computing device 102 can also include a model trainer 122. The model trainer 122 can train or re-train one or more of the machine-learned models 120 stored at the user computing device 102 using various training or learning techniques, such as, for example, backwards propagation of errors (e.g., truncated backpropagation through time). In particular, the model trainer 122 can train or re-train one or more of the machine-learned models 120 using the locally logged data 119 as training data." [¶0024] "can provide a local, on-device database to which log entries are written") one or more processors; and one or more non-transitory computer-readable media that collectively store instructions that, when executed by the one or more processors, cause the computing system to perform operations for each of one or more update iterations, the operations comprising:([¶0048] "The user computing device 102 includes one or more processors 112 and a memory 114. The one or more processors 112 can be any suitable processing device") sampling, by the computing system, from a pool of data associated with one or more ancillary systems to generate a current set of training data;([¶0076] "Each application 1 through N also includes local application data. In particular, each application can provide or otherwise communicate with a local, on-device database in memory to which log entries are written. One database can be used for all applications or different respective databases can be used for each application" [¶0054] " the model trainer 122 can train or re-train one or more of the machine-learned models 120 using the locally logged data 119 as training data" [¶0026] "The user computing device can be instructed to use the entries in the local database (or some subset thereof) as training data for the training algorithm" Bonawitz explicitly discloses using a subset (sample) of the local dataset (pool) for training) training, by the computing system, a machine learning model on the current set of training data to generate an updated model;([¶0033] "rather than receive and/or train a single machine-learned model, the user computing device can receive and/or train a plurality of machine-learned models" See also FIG. 2-4 where FIG. 2 describes train/update step, FIG. 3 elaborates and provides iterative evaluation and update, and FIG. 4 extends that evaluation to multiple models) evaluating, by the computing system, a performance of the updated model relative to a current set of testing data; and([¶0026] "the trained model can be evaluated prior to activation." See also FIG. 2 which places activation after obtain/train and evaluate) performing, by the computing system, a comparison of the performance of the updated model relative to the current set of testing data with a respective performance of one or more other machine learning models on the current set of testing data or one or more past sets of testing data.([Abstract] "the user computing device can evaluate a plurality of machine-learned models against locally stored data" [¶0021] "the user computing device can obtain a plurality of machine-learned models and can evaluate at least one performance metric for each of the plurality of machine-learned models relative to the data that is stored locally at the user computing device" [¶0025] "This process also yields a history of multiple versions of the model, and each of these could be updated on each new dataset to evaluate the best candidate."). Regarding claim 17, Bonawitz teaches The computing system of claim 16, wherein the operations further comprise: selecting, by the computing system, either the updated model or one of the one or more other machine learning models for deployment based on the comparison of the performance of the updated model relative to the current set of testing data with the respective performance of the one or more other machine learning models on the current set of testing data.(Bonawitz [¶0025] "This process also yields a history of multiple versions of the model, and each of these could be updated on each new dataset to evaluate the best candidate." [¶0035] "The user computing device and/or server computing device can select one of the machine-learned models based at least in part on the performance metrics. For example, the model with the best performance metric(s) can be selected for activation and use at the user computing device." [¶0120] "the server computing device can instruct the user computing device to activate and use the selected at least one machine-learned model."). Regarding claim 19, Bonawitz teaches The computing system of claim 16, wherein performing, by the computing system, the comparison of the performance of the updated model relative to the current set of testing data with the respective performance of the one or more other machine learning models on the current set of testing data or the one or more past sets of testing data comprises: determining, by the computing system, a first set of statistical tests for the performance of the updated model relative to the current set of testing data;(Bonawitz [¶0103] " the performance metric can be an average of the errors over time" [¶0008] "The method includes evaluating, by the user computing device, at least one performance metric for each of the plurality of machine-learned models") determining, by the computing system, a second set of statistical tests for the respective performance of the one or more other machine learning models on the current set of testing data; and(Bonawitz [¶0103] " the performance metric can be an average of the errors over time" [¶0008] "The method includes evaluating, by the user computing device, at least one performance metric for each of the plurality of machine-learned models") performing, by the computing system, a comparison of the first set of statistical tests and the second set of statistical tests.(Bonawitz [Abstract] "the user computing device can evaluate a plurality of machine-learned models against locally stored data" [¶0021] "the user computing device can obtain a plurality of machine-learned models and can evaluate at least one performance metric for each of the plurality of machine-learned models relative to the data that is stored locally at the user computing device" [¶0025] "This process also yields a history of multiple versions of the model, and each of these could be updated on each new dataset to evaluate the best candidate."). Regarding claim 20, claim 20 is substantially similar to claim 16. Therefore, the rejection applied to claim 16 also applies to claim 20. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claim 4 is rejected under U.S.C. §103 as being unpatentable over the combination of Bonawitz and Bourtoule (“Machine Unlearning”, 2021). Regarding claim 4, Bonawitz teaches The computer-implemented method of claim 1. However, Bonawitz doesn't explicitly teach wherein the one or more other machine learning models comprise previous checkpoints of the machine learning model. Bourtoule, in the same field of endeavor, teaches the one or more other machine learning models comprise previous checkpoints of the machine learning model. ([p. 156] "it is already a common practice to regularly checkpoint models during training"). Bonawitz as well as Bourtoule are directed towards distributed learning. Therefore, Bonawitz as well as Bourtoule are analogous art in the same field of endeavor. It would have been obvious before the effective filing date of the claimed invention to combine the teachings of Bonawitz with the teachings of Bourtoule by using model checkpointing. Bourtoule provides as additional motivation for combination ([p. 156] "it is already a common practice to regularly checkpoint models during training"). This motivation for combination also applies to the remaining claims which depend on this combination. Claims 6 and 18 are rejected under U.S.C. §103 as being unpatentable over the combination of Bonawitz and Sarferaz (US20200380155A1). Regarding claim 6, Bonawitz teaches The computer-implemented method of claim 1. However, Bonawitz doesn't explicitly teach, wherein sampling, by the computing system, from the pool of data associated with the one or more ancillary systems to generate the current set of training data comprises: associating, by the computing system, a wipeout-compliant flag with the current set of training data, wherein the wipeout-compliant flag causes deletion of the current set of training data upon occurrence of a condition; and deleting, by the computing system, the current set of training data upon occurrence of the condition.. Sarferaz, in the same field of endeavor, teaches The computer-implemented method of claim 1, wherein sampling, by the computing system, from the pool of data associated with the one or more ancillary systems to generate the current set of training data comprises: associating, by the computing system, a wipeout-compliant flag with the current set of training data, wherein the wipeout-compliant flag causes deletion of the current set of training data upon occurrence of a condition; and deleting, by the computing system, the current set of training data upon occurrence of the condition. ([0045] "The machine learning algorithm 220 can access data, such as for use in generating the trained model 212" [0054] "The retention manager 236 can include a retention policy executor 274. The retention policy executor 274 can, at least for some data objects 264, periodically determine if an expiration date has passed, and, if so, and there are no status flags (e.g., legal holds) set, delete the data object from the archive 240"). Bonawitz as well as Sarferaz are directed towards distributed training. Therefore, Bonawitz as well as Sarferaz are analogous art in the same field of endeavor. It would have been obvious before the effective filing date of the claimed invention to combine the teachings of Bonawitz with the teachings of Sarferaz by checking flags to determine whether training data should be deleted. This is legally required in certain jurisdictions and Sarferaz provides as additional motivation for combination ([¶0004] “The data subject may be able to request, such as under applicable laws or regulations of a jurisdiction, that an organization delete their data or “forget” them”). This motivation for combination also applies to the remaining claims which depend on this combination. Regarding claim 18, Bonawitz teaches The computing system of claim 16. However, Bonawitz doesn't explicitly teach, wherein sampling, by the computing system, from the pool of data associated with the one or more ancillary systems to generate the current set of training data comprises: associating, by the computing system, a wipeout-compliant flag with the current set of training data, wherein the wipeout-compliant flag causes deletion of the current set of training data upon occurrence of a condition; and deleting, by the computing system, the current set of training data upon occurrence of the condition. Sarferaz, in the same field of endeavor, teaches sampling, by the computing system, from the pool of data associated with the one or more ancillary systems to generate the current set of training data comprises: associating, by the computing system, a wipeout-compliant flag with the current set of training data, wherein the wipeout-compliant flag causes deletion of the current set of training data upon occurrence of a condition; and deleting, by the computing system, the current set of training data upon occurrence of the condition. ([0045] "The machine learning algorithm 220 can access data, such as for use in generating the trained model 212" [0054] "The retention manager 236 can include a retention policy executor 274. The retention policy executor 274 can, at least for some data objects 264, periodically determine if an expiration date has passed, and, if so, and there are no status flags (e.g., legal holds) set, delete the data object from the archive 240"). Bonawitz as well as Sarferaz are directed towards distributed training. Therefore, Bonawitz as well as Sarferaz are analogous art in the same field of endeavor. It would have been obvious before the effective filing date of the claimed invention to combine the teachings of Bonawitz with the teachings of Sarferaz by checking flags to determine whether training data should be deleted. This is legally required in certain jurisdictions and Sarferaz provides as additional motivation for combination ([¶0004] “The data subject may be able to request, such as under applicable laws or regulations of a jurisdiction, that an organization delete their data or “forget” them”). This motivation for combination also applies to the remaining claims which depend on this combination. Claims 7 and 8 are rejected under U.S.C. §103 as being unpatentable over the combination of Bonawitz and Alizadeh (US20220383142A1). Regarding claim 7, Bonawitz teaches The computer-implemented method of claim 1. However, Bonawitz doesn't explicitly teach, wherein sampling, by the computing system, from the pool of data associated with the one or more ancillary systems to generate the current set of training data comprises: randomly sampling, by the computing system, from the pool of data associated with the one or more ancillary systems to generate the current set of training data. Alizadeh, in the same field of endeavor, teaches The computer-implemented method of claim 1, wherein sampling, by the computing system, from the pool of data associated with the one or more ancillary systems to generate the current set of training data comprises: randomly sampling, by the computing system, from the pool of data associated with the one or more ancillary systems to generate the current set of training data. ([¶0092] "a random sample of normal users (and optionally political users), hereafter “organic activity”, over a given period to form the training data"). Bonawitz as well as Alizadeh are directed towards machine learning training. Therefore, Bonawitz as well as Alizadeh are reasonably pertinent analogous art. It would have been obvious before the effective filing date of the claimed invention to combine the teachings of Bonawitz with the teachings of Alizadeh by randomly sampling the training dataset. Alizadeh provides as additional motivation for combination ([¶0010] “The improvement also includes retraining the classifier to distinguish between a post-URL pair produced from a coordinated influence effort and a post-URL pair produced from a random user using the extracted plurality of content-based features based on the additional resulting label.”). This motivation for combination also applies to the remaining claims which depend on this combination. Regarding claim 8, Bonawitz teaches The computer-implemented method of claim 1. However, Bonawitz doesn't explicitly teach wherein sampling, by the computing system, from the pool of data associated with the one or more ancillary systems to generate the current set of training data comprises: sampling, by the computing system and from the pool of data associated with the one or more ancillary systems, only data examples that have been newly generated within a defined period of time.. Alizadeh, in the same field of endeavor, teaches sampling, by the computing system, from the pool of data associated with the one or more ancillary systems to generate the current set of training data comprises: sampling, by the computing system and from the pool of data associated with the one or more ancillary systems, only data examples that have been newly generated within a defined period of time. ([¶0092] "a random sample of normal users (and optionally political users), hereafter “organic activity”, over a given period to form the training data"). Bonawitz as well as Alizadeh are directed towards machine learning training. Therefore, Bonawitz as well as Alizadeh are reasonably pertinent analogous art. It would have been obvious before the effective filing date of the claimed invention to combine the teachings of Bonawitz with the teachings of Alizadeh by randomly sampling the training dataset over a defined period of time. Alizadeh provides as additional motivation for combination ([¶0010] “The improvement also includes retraining the classifier to distinguish between a post-URL pair produced from a coordinated influence effort and a post-URL pair produced from a random user using the extracted plurality of content-based features based on the additional resulting label.”). This motivation for combination also applies to the remaining claims which depend on this combination. Claims 11, 13, and 14 are rejected under U.S.C. §103 as being unpatentable over the combination of Bonawitz and Wang (US20210150379A1). Regarding claim 11, Bonawitz teaches The computer-implemented method of claim 9. However, Bonawitz doesn't explicitly teach, wherein the first set of statistical tests and the second set of statistical tests each comprise a set of error bounds. Wang, in the same field of endeavor, teaches the first set of statistical tests and the second set of statistical tests each comprise a set of error bounds. ([¶0060] "Examples of metrics include sensitivity metrics, accuracy metrics, longevity metrics, specificity, etc. More particularly, examples of metrics include Logarithmic Loss, True Positive Rate (Sensitivity), False Positive Rate, True Negative Rate (Specificity), F1 Score, Precision, Recall, Mean Absolute Error, Mean Squared Error, etc. The distribution analysis module 310 is further configured to use the results of metric measurements across multiple different analytical models to create a distribution, such as the approximate normal distribution 400 illustrated in FIG. 4. The distribution 400 represents a range of metric results for a given metric"). Bonawitz as well as Wang are directed towards distributed training. Therefore, Bonawitz as well as Wang are analogous art in the same field of endeavor. It would have been obvious before the effective filing date of the claimed invention to combine the teachings of Bonawitz with the teachings of Wang by providing alerts that model training is deviating from a normalized range or distribution. Wang provides as additional motivation for combination ([¶0059] “The evaluation engine 340 is configured to use output of the distribution analysis module 310, the survival analysis module 320, and/or ensemble module 330 to improve the respective modules based on feedback from output results.”). This motivation for combination also applies to the remaining claims which depend on this combination. Regarding claim 13, Bonawitz teaches The computer-implemented method of claim 10. However, Bonawitz doesn't explicitly teach performing, by the computing system, the comparison of the first set of statistical tests and the second set of statistical tests comprises normalizing at least the first set of statistical tests based on feature values associated with the current set of testing data.. Wang, in the same field of endeavor, teaches performing, by the computing system, the comparison of the first set of statistical tests and the second set of statistical tests comprises normalizing at least the first set of statistical tests based on feature values associated with the current set of testing data.([¶0060] "The distribution analysis module 310 is further configured to use the results of metric measurements across multiple different analytical models to create a distribution, such as the approximate normal distribution 400 illustrated in FIG. 4"). Bonawitz as well as Wang are directed towards distributed training. Therefore, Bonawitz as well as Wang are analogous art in the same field of endeavor. It would have been obvious before the effective filing date of the claimed invention to combine the teachings of Bonawitz with the teachings of Wang by providing alerts that model training is deviating from a normalized range or distribution. Wang provides as additional motivation for combination ([¶0059] “The evaluation engine 340 is configured to use output of the distribution analysis module 310, the survival analysis module 320, and/or ensemble module 330 to improve the respective modules based on feedback from output results.”). This motivation for combination also applies to the remaining claims which depend on this combination. Regarding claim 14, Bonawitz teaches The computer-implemented method of claim 1. However, Bonawitz doesn't explicitly teach further comprising: providing, by the computing system, an automated alert when the performance of the updated model relative to the current set of testing data deviates from the respective performance of the one or more other machine learning models on the current set of testing data. Wang, in the same field of endeavor, teaches providing, by the computing system, an automated alert when the performance of the updated model relative to the current set of testing data deviates from the respective performance of the one or more other machine learning models on the current set of testing data. ([¶0003] "The method further includes comparing the model metric values to the normal distributions for model metric results for each of the received model metric values, and alerting to model degradation of the analytical model based on the comparison of the model metric values to the normal distributions for model metric results."). Bonawitz as well as Wang are directed towards distributed training. Therefore, Bonawitz as well as Wang are analogous art in the same field of endeavor. It would have been obvious before the effective filing date of the claimed invention to combine the teachings of Bonawitz with the teachings of Wang by providing alerts that model training is deviating from a normalized range or distribution. Wang provides as additional motivation for combination ([¶0059] “The evaluation engine 340 is configured to use output of the distribution analysis module 310, the survival analysis module 320, and/or ensemble module 330 to improve the respective modules based on feedback from output results.”). This motivation for combination also applies to the remaining claims which depend on this combination. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Kratzwald (“Learning from On-Line User Feedback in Neural Question Answering on the Web”, 2019) is directed towards continuous online learning and model selection. Any inquiry concerning this communication or earlier communications from the examiner should be directed to SIDNEY VINCENT BOSTWICK whose telephone number is (571)272-4720. The examiner can normally be reached M-F 7:30am-5:00pm EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Miranda Huang can be reached on (571)270-7092. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /SIDNEY VINCENT BOSTWICK/Examiner, Art Unit 2124
Read full office action

Prosecution Timeline

Aug 05, 2024
Application Filed
Sep 11, 2026
Non-Final Rejection mailed — §101, §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12699874
Leveraging Redundancy in Attention with Reuse Transformers
3y 10m to grant Granted Aug 04, 2026
Patent 12675673
NEURAL NETWORK PROCESSING DEVICE, METHOD, AND COMPUTER-READABLE RECORDING MEDIUM
3y 7m to grant Granted Jul 07, 2026
Patent 12645914
INSTRUCTION PRUNING FOR NEURAL NETWORKS
3y 6m to grant Granted Jun 02, 2026
Patent 12626139
SECRET SOFTMAX FUNCTION CALCULATION SYSTEM, SECRET SOFTMAX FUNCTION CALCULATION APPARATUS, SECRET SOFTMAX FUNCTION CALCULATION METHOD, SECRET NEURAL NETWORK CALCULATION SYSTEM, SECRET NEURAL NETWORK LEARNING SYSTEM, AND PROGRAM
4y 3m to grant Granted May 12, 2026
Patent 12619815
Magnitude Invariant Multimodal Agent for Efficient Image-Text Interface Automation
1y 6m to grant Granted May 05, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
51%
Grant Probability
86%
With Interview (+35.1%)
4y 5m (~2y 3m remaining)
Median Time to Grant
Low
PTA Risk
Based on 152 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month