DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statement (IDS) submitted on April 23, 2024 has been considered by the Examiner.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 1, 5-8, 12-15 and 18-20 are rejected under 35 U.S.C. 103 as being unpatentable over U.S. Patent Application Publication No. 2024/0046151 to Singh et al. (“Singh”), over the article entitled “Language Models are Few-Shot Learners” by Brown et al. (“Brown”), and also over U.S. Patent Application Publication No. 2025/0265827 to Kitayama et al. (“Kitayama”).
Regarding claims 1, 8 and 15, Singh describes a method and electronic device for automated Machine Learning (ML) model retraining (see e.g. paragraphs 0002 and 0008). Similar to the claims, Singh particularly teaches:
capturing a set of context data within a machine learning model (MLM) (see e.g. paragraph 0011: Singh teaches predicting an accuracy degradation of a first MLM by using a second MLM, whereupon if the predicted accuracy degradation meets a pre-defined threshold, the first MLM is automatically retrained. Singh particularly discloses that predicting the accuracy degradation comprises receiving data regarding the accuracy of the first MLM, such as a model type, parameters and hyper parameters, model training time, model prediction accuracies, resources used for model training, extraction times, a time window of data extraction, data generation patterns, model accuracy data, and an execution time for each training pipeline – see e.g. paragraph 0013. Such data, or particular portions thereof, is considered a set of context data captured within the first MLM.),
wherein the MLM generates outputs using an established context (see e.g. paragraph 0011: as noted above, Singh teaches predicting an accuracy degradation of a first MLM by using a second MLM. Singh teaches that the first MLM can be a neural network that generates outputs using an established context, e.g. using weights established through machine learning – see e.g. paragraph 0056.);
assessing the MLM through a second MLM (see e.g. paragraph 0011: as noted above, Singh teaches predicting an accuracy degradation of a first MLM by using a second MLM. The second MLM is thus considered to assess the first MLM to predict its accuracy degradation.),
wherein assessing the MLM through the second MLM determines if the captured set of context data indicates the established context within the MLM will be revised (see e.g. paragraph 0011: as noted above, Singh teaches predicting an accuracy degradation of a first MLM by using a second MLM, whereupon if the predicted accuracy degradation meets a pre-defined threshold, the first MLM is automatically retrained. Retraining the first MLM would understandably entail revising the established context, e.g. the one or more weights, of the first MLM. Like further noted above, Singh particularly discloses that predicting the accuracy degradation of the first MLM comprises capturing a set of context data within the first MLM, such as a model type, parameters and hyper parameters, etc. – see e.g. paragraph 0013. Singh further discloses that the second MLM is used to analyze this captured set of context data to predict the accuracy degradation of the first MLM – see e.g. paragraph 0013. Accordingly, assessing the first MLM through the second MLM would entail determining if the captured set of context data, e.g. the type of the first MLM, its parameters, etc. indicates the established context within the first MLM will be revised, i.e. that the first MLM needs to be retrained.); and
revising the established context within the MLM (see e.g. paragraph 0011: as noted above, Singh teaches predicting an accuracy degradation of a first MLM by using a second MLM, whereby the first MLM is automatically retrained if the predicted accuracy degradation meets a pre-defined threshold. The established context, e.g. weights, within the MLM are thus revised with updated weights through the retraining of the MLM.).
Singh thus teaches a method similar to that of claim 15. Singh further discloses that these teachings can be realized via computer executable instructions stored in the memory of a computing device that further comprises a processor coupled to the memory to execute the instructions (see e.g. paragraphs 0016, 0049, 0054 and 0129-0132). Such a computing device comprising a memory (i.e. a non-transitory storage device) and a processor (i.e. a processing device) to implement the above-described tasks taught by Singh is considered a system similar to that of claim 1. The memory storing the computer-executable instructions is considered a computer program product similar to that of claim 8. However, Singh does not explicitly disclose that the above-noted second MLM is a large language model (LLM), as is required by claims 1, 8 and 15. Singh also does not teach: (i) triggering parallel event streams within the MLM, wherein a first event stream within the MLM is introduced to the set of context data with the established context, and wherein the second event stream within the MLM is introduced to the set of context data with a revised context; (ii) validating outputs of the MLM through parallel event streams within the MLM generated through the set of context data; and (iii) wherein the established context within the MLM is revised based on validated outputs of the MLM and outputs generated through the revised context, as is also required by claims 1, 8 and 15.
Large Language Models and their uses and advantages are nevertheless well-known in the art. Brown in particular teaches that large language models are few-shot learners, whereby a large language model is able to learn a new task by being given a few examples of the task (see e.g. the Abstract, section 1 “Introduction” on pages 1-2, and section 2 “Approach” on page 3).
It would have been obvious to one of ordinary skill in the art, having the teachings of Singh and Brown before the effective filing date of the claimed invention, to modify the system, computer program product and method taught by Singh such that the second MLM is particularly a large language model like taught by Brown. It would have been advantageous to one of ordinary skill to utilize such a large language model because it can be trained (i.e. in a few-shot setting) with a smaller training set, as is suggested by Brown (see e.g. the Abstract, section 1 “Introduction” on pages 1-2, and section 2 “Approach” on page 3).
Kitayama generally teaches automatically verifying an updated MLM, such as a deep neural network (DNN), before deploying the updated MLM (see e.g. paragraphs 0002, 0004-0005 and 0012-0016). In particular, Kitayama teaches that such automatic verification entails:
triggering parallel event streams within the MLM, wherein a first event stream within the MLM is introduced to a set of context data with an established context and wherein a second event stream within the MLM is introduced to the set of context data with a revised context (see e.g. paragraphs 0013 and 0037-0039, and FIG. 1: Kitayama teaches receiving inference results from an existing MLM and an updated MLM when the same input stream is provided to both the existing and updated MLM, and determining whether there is a discrepancy between the two inference results. Accordingly, Kitayama teaches triggering parallel event streams within the MLM, wherein a first event stream, i.e. an input stream, within the MLM is introduced to a set of context data with an established context, i.e. to the MLM with established existing weights, and wherein a second event stream, i.e. also the input stream, within the MLM is introduced to the set of context data with a revised context, i.e. to the MLM with updated weights.);
validating outputs of the MLM through parallel event streams within the MLM generated through the set of context data (see e.g. paragraphs 0013 and 0037-0039, and FIG. 1: as noted above, Kitayama teaches receiving inference results from both an existing MLM and an updated MLM when both are provided the same input, and determining whether there is a discrepancy between the two inference results. The process further verifies the performance of the updated MLM based in part on a comparison of the two inference results – see e.g. paragraphs 0014-0015, 0039-0042, 0047, and 0056-0057. The outputs of the MLM, particularly with the updated weights, are thus validated through the parallel event streams within the MLM generated through the set of context data, i.e. through the inputs to both the MLM with existing weights and the MLM with updated weights.); and
revising the established context within the MLM based on validated outputs of the MLM and outputs generated through the revised context (see e.g. paragraph 0005: Kitayama suggests that the updated MLM is first validated before being executed in a computing environment, understandably in place of the existing MLM. The established context within the MLM is thus revised, i.e. with the updated MLM weights, based on the validated outputs of the MLM and the outputs generated through the revised context, i.e. through the updated MLM weights.).
It would have been obvious to one of ordinary skill in the art, having the teachings of Singh, Brown and Kitayama before the effective filing date of the claimed invention, to modify the system, computer program product and method taught by Singh and Brown so as to validate the retrained (i.e. updated) MLM like taught by Kitayama before deploying the retrained MLM in place of the existing MLM, wherein such validation comprises: (i) triggering parallel event streams within the MLM, wherein a first event stream within the MLM is introduced to the set of context data with the established context (i.e. to the existing MLM), and wherein a second event stream within the MLM is introduced to the set of context data with a revised context (i.e. to the retrained MLM); (ii) validating outputs of the MLM through parallel event streams within the MLM generated through the set of context data; and (iii) wherein the established context within the MLM is revised based on validated outputs of the MLM and outputs generated through the revised context. It would have been advantageous to one of ordinary skill to utilize such a combination because it can reduce the human cost for verification, as is taught by Kitayama (see e.g. paragraphs 0005 and 0016). Accordingly, Singh, Brown and Kitayama are considered to teach, to one of ordinary skill in the art, a system like that of claim 1, a computer program product like that of claim 8 and a method like that of claim 15.
As per claims 5, 12 and 18, Singh suggests that capturing the set of context data is initiated via triggered modules within the MLM (see e.g. paragraph 0011: like noted above, Singh teaches predicting an accuracy degradation of a first MLM, whereupon if the predicted accuracy degradation meets a pre-defined threshold, the first MLM is automatically retrained. Like further noted above, Singh particularly discloses that predicting the accuracy degradation comprises receiving a set of context data, such as a model type, parameters and hyper parameters, model training time, model prediction accuracies, resources used for model training, extraction times, a time window of data extraction, data generation patterns, model accuracy data, and an execution time for each training pipeline – see e.g. paragraph 0013. The computer program code necessary for capturing such data are considered modules within the first MLM.). Accordingly, the above-described combination of Singh, Brown and Kitayama further teaches a system like that of claim 5, a computer-program product like that of claim 12, and a method like that of claim 18.
As per claims 6, 13 and 19, Singh further teaches that revising context within the MLM at least partially replaces established context with the revised context (see e.g. paragraph 0011: as noted above, Singh teaches predicting an accuracy degradation of a first MLM, whereby the first MLM is automatically retrained if the predicted accuracy degradation meets a pre-defined threshold. The weights, i.e. established context, within the first MLM are thus revised by being replaced with updated weights, i.e. a revised context, through the retraining of the first MLM.). As described above, it would have been obvious to modify the system, computer program product and method taught by Singh and Brown so as to validate the MLM (i.e. the retrained/updated MLM) and its outputs like taught by Kitayama. Accordingly, the above-described combination of Singh, Brown and Kitayama further teaches a system like that of claim 6, a computer-program product like that of claim 13, and a method like that of claim 19.
As per claims 7, 14 and 20, Singh further suggests capturing the set of context data in predetermined periodic intervals (see e.g. paragraphs 0045 and 0070: Singh teaches periodically predicting an accuracy degradation of the first MLM. Singh further teaches that the accuracy degradation is predicted based on a set of context data captured from, e.g., a datastore – see e.g. paragraph 0013. Accordingly, it follows that the set of context data is captured, e.g. from the datastore, in predetermined periodic intervals to predict the accuracy degradation of the first MLM.). Accordingly, the above-described combination of Singh, Brown and Kitayama further teaches a system like that of claim 7, a computer-program product like that of claim 14, and a method like that of claim 20.
Claims 2-4, 9-11, 16 and 17 are rejected under 35 U.S.C. 103 as being unpatentable over the combination of Singh, Brown and Kitayama described above, and also over the article entitled, “Model Versioning for ML Models: A Comprehensive Guide” by Tonye Harry (“Harry”).
Regarding claims 2, 9 and 16, Singh, Brown and Kitayama teach a system like that of claim 1, a computer program product like that of claim 8 and a method like that of claim 15, as is described above, and which entail revising an established context within an MLM based on validated outputs of the MLM and outputs generated through a revised context. Singh, Brown and Kitayama, however, do not explicitly teach that the established context and the revised context comprise corresponding markers indicating a time in which each context was generated, as is required by claims 2, 9 and 16.
Harry generally describes machine learning model versioning, which “ensures that modifications to the machine learning model are tracked and managed”:
In a typical software development process, changes are made rapidly and continuously due to the experimental nature of technology development. As a result, versioning has become an essential component to keep track of modifications made to the source code and identify the team member responsible for them.
This is particularly relevant in Machine Learning (ML) systems, where teams need to track changes in data, code, and the model being developed to achieve optimal results. Specifically, three types of versioning are crucial in an ML system:
Data Versioning: This mostly involves tracking and managing changes to the data utilized in creating the model.
Code Versioning: This enables the tracking of modifications made to the source code that powers an ML system, ensuring transparency and facilitating collaboration.
Model Versioning: This ensures that modifications to the machine learning model are tracked and managed. It can also involve some aspects of data versioning when needed.
In this article, we will focus on model versioning and provide a comprehensive guide to help you understand what it is and the various technicalities involved.
(“Introduction”).
Regarding the claimed invention, Harry particularly teaches that such versioning can comprise a calendar versioning scheme that associates corresponding markers (i.e. a version number) with each version indicating a time (i.e. a date) in which the version was created:
Calendar Versioning:Calendar versioning is a versioning scheme that uses the date of release as the version number. For example, a model released on January 1st, 2022 would have a version number of 2022.01.01. This scheme is often used for data science projects where the focus is on tracking changes over time.
(“Versioning Schemes”).
It would have been obvious to one of ordinary skill in the art, having the teachings of Singh, Brown, Kitayama and Harry before the effective filing date of the claimed invention, to modify the system, computer program product and method taught by Singh, Brown and Kitayama so as to employ a calendar versioning scheme like taught by Harry to track and manage modifications to the MLM, whereby a corresponding marker is associated with each version of the MLM to indicate the time in which the version was generated. It thus follows that the original MLM (i.e. the established context/original weights) and the retrained MLM (i.e. the revised context/updated weights) would comprise corresponding markers indicating a time in which context was generated. It would have been advantageous to one of ordinary skill to utilize such a combination because it would enable modifications to the MLM to be tracked over time, as is taught by Harry (see e.g. the portions of the “Introduction” and “Versioning Schemes” excerpted above). Accordingly, Singh, Brown, Kitayama and Harry are considered to teach, to one of ordinary skill in the art, a system like that of claim 2, a computer program product like that of claim 9 and a method like that of claim 16.
As per claims 3, 10 and 17, it would have been obvious, as is described above, to modify the system, computer program product and method taught by Singh, Brown and Kitayama so as to employ a calendar versioning scheme like taught by Harry to track and manage modifications to the MLM, whereby a corresponding marker is associated with each version of the MLM to indicate the time in which the version was generated. Moreover, Singh generally suggests that later revised machine learning model versions can be implemented before earlier revised machine learning model versions (see e.g. paragraph 0011: like noted above, Singh teaches predicting an accuracy degradation of a first MLM, whereupon if the predicted accuracy degradation meets a pre-defined threshold, the first MLM is automatically retrained. After the first MLM is retrained, an end user first accessing the model would understandably implement the latest, retrained model prior to any earlier versions of the model.). It thus follows that, in such circumstances in which the later model versions are implemented before earlier model versions, the revised context with later time markers (i.e. the later model/weight versions) are implemented before revised context with earlier time markers (i.e. earlier model/weight versions). Accordingly, the above-described combination of Singh, Brown, Kitayama and Harry is further considered to teach a system like that of claim 3, a computer program product like that of claim 10 and a method like that of claim 17.
As per claims 4 and 11, Singh suggests that the established context is revised based on time markers associated with the established context (see e.g. paragraph 0011: like noted above, Singh teaches predicting an accuracy degradation of a first MLM, whereupon if the predicted accuracy degradation meets a pre-defined threshold, the first MLM is automatically retrained. Singh particularly discloses that the accuracy degradation is predicted based on received data regarding the accuracy of the first MLM, such as a model type, parameters and hyper parameters, model training time, model prediction accuracies, resources used for model training, extraction times, a time window of data extraction, data generation patterns, model accuracy data, and an execution time for each training pipeline – see e.g. paragraph 0013. The model training time, extraction times, time window of data extraction and/or execution time for each training pipeline are considered time markers associated with the established context.). Accordingly, the above-described combination of Singh, Brown, Kitayama and Harry is further considered to teach a system like that of claim 4 and a computer program product like that of claim 11.
Conclusion
The prior art made of record on form PTO-892 and not relied upon is considered pertinent to applicant’s disclosure. The applicant is required under 37 C.F.R. §1.111(C) to consider these references fully when responding to this action. In particular, the U.S. Patent Application Publication to Unfried cited therein describes a method and system for changing over from a first data processing version that uses at least one data model to a second data processing version that also uses at least one data model, whereby in a first phase, the second data processing version is used in parallel to the first data processing version to continuously adapt the at least one data model related to the first version as well as the data model related to the second version. The article by Kurtis Pykes cited therein (“A Guide to Monitoring Machine Learning Models in Production”) generally teaches monitoring machine learning models to identify whether the models require updating.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to BLAINE T BASOM whose telephone number is (571)272-4044. The examiner can normally be reached Monday-Friday, 9:00 am - 5:30 pm, EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kieu Vu can be reached at (571)272-4057. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/BTB/
8/22/2026
/MATTHEW ELL/Supervisory Patent Examiner, Art Unit 2141