DETAILED ACTION
Claims 1-6, 8-18, and 20 are presented for examination.
This office action is in response to submission of application on 22-JANUARY-2026.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 31-AUGUST-2022 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
The information disclosure statement (IDS) submitted on 13-JANUARY-2023 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
The information disclosure statement (IDS) submitted on 13-FEBRUARY-2025 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
The information disclosure statement (IDS) submitted on 25-SEPTEMBER-2025 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Response to Amendment
The amendment filed 22-JANUARY-2026 in response to the non-final office action mailed 11-DECEMBER-2025 has been entered. Claims 1-6, 8-18, and 20 remain pending in the application.
With regards to the non-final office action’s rejection under 101, the amendments to the claims have overcome the original rejection with regards to the claims being directed towards an abstract idea.
With regards to the non-final office action’s rejection under 103, the amendment to the claims have overcome the original rejection. However, upon a new search for the amended limitations, a new 103 rejection over Shah in view of Ghosh further in view of Tiku further in view of Carbune has been written.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-4, 8-12, 14-18 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Shah et al. (Pub. No. US 20230036289 A1, filed July 27th 2021, hereinafter Shah) in view of Ghosh (Pub. No. US 20220353172 A1, filed April 30th 2021, hereinafter Ghosh) further in view of Tiku et al. (Pub. No. US 20220156560 A1, filed November 18th 2020, hereinafter Tiku), further in view of Carbune et al. (Pub. No. US 20190220755 A1, filed January 18th 2018, hereinafter Carbune).
Regarding claim 1:
Claim 1 recites:
A system, comprising: one or more processors configured to: receive a request to train a machine learning model; receive a training schedule specifying a plurality of values for one or more hardware- level performance settings and one or more model-level performance settings to be set at different points in time during a training process for the machine learning model; train the machine learning model in accordance with the training schedule by setting the one or more hardware-level and model-level performance settings to first values of the plurality of values that select training speed over model quality at a first point in time during the training process; and adjusting the one or more hardware-level and model-level performance settings to second values of the plurality of values that select model quality over training speed at a second point in time after the first point in time during the training process, and send the trained machine learning model to one or more computing devices.
Shah discloses one or more processors configured to: receive a request to train a machine learning model; receive a training schedule specifying a plurality of values for one or more hardware-level performance settings and one or more model-level performance settings:
Shah teaches a user device sending a user input describing a desired model to the network device that generates the machine learning model (Paragraph 20), which would be analogous to a request to train a machine learning model. Furthermore, this user input more specifically contains instructions for generating and testing the model, along with a desired model type and hyperparameters (Paragraph 12), which would a training scheduling specifying a plurality of values for one or more performance settings, wherein the model type would be an example of model-level performance settings.
Shah does not teach hardware-level performance settings. This limitation is taught further below by Tiku.
Shah discloses send the trained machine learning model to one or more computing devices:
Shah teaches that information is sent back from the network device to the user device regarding a completed machine learning model, such that the user device may recreate the model (Paragraph 30), which would functionally be analogous to sending the trained machine learning model to one or more computing devices.
However, Shah does not disclose train the machine learning model in accordance with the training schedule, wherein the settings are set to different values of the plurality of values of the training schedule at different points in time during training. This is instead disclosed by Ghosh.
Ghosh discloses train the machine learning model in accordance with the training schedule, wherein the one or more hardware-level performance settings, and the one or more model-level performance settings are set to different values of the plurality of values of the training schedule at different points in time during training:
Ghosh in the same field of endeavor of machine learning teaches a machine learning model that during a training process maps received metric information (which would be the volume of message traffic to be transmitted (Paragraph 14) for different levels of message activity (Paragraph 20). The different levels of activity would indicate that the metric information is set to different values of the plurality of values of the training schedule, wherein as the model is trained over time the metric information is at different points.
Model-level performance settings have previously been taught by Shah.
Ghosh does not teach hardware-level performance settings as in settings that adjust the training process for the machine learning model, as described in the arguments. This limitation is taught further below by Tiku.
Ghosh, Shah, and the present application are all analogous art because they are all in the same field of endeavor of machine learning.
Shah and Ghosh do not teach hardware-level performance settings. Instead, this limitation is taught by Tiku:
Tiku in the same field of endeavor of machine learning teaches that based on a selected resource criterion, the path of training for an artificial neural network may be altered (Paragraph 47). This determination of the training path would be hardware-level performance settings since this is a setting that adjusts the training process for the machine learning model on a hardware level, as the resource criterion may, for example, be the energy consumption of a memory device (Paragraph 44) which would be a criterion resulting from the hardware.
Tiku and the present application are analogous art because they are all in the same field of endeavor of machine learning
However, none of Shah, Ghosh, or Tiku disclose by setting the [one or more hardware-level and model-level performance settings] to first values of the plurality of values that select training speed over model quality at a first point in time during the training process; and adjusting the [one or more hardware-level and model-level performance settings] to second values of the plurality of values that select model quality over training speed at a second point in time after the first point in time during the training process. Instead, this limitation is taught by Carbune:
Carbune teaches a step size hyperparameter that is iteratively reduced until a realism score exceeds a threshold score (Paragraph 87). The iterative reduction would be a first value at a first point in time during the training process that becomes a second value at a second point in time after the first point in time during the training process. Furthermore, Carbune teaches that the threshold may also be used to determine an appropriate balance between speed and realism i.e. model quality (Paragraph 120), and therefore initial values in an embodiment of Carbune may select for training speed over model quality at a first point in time, and model quality over training speed at a second point in time.
Finally, Carbune teaches that the adversarial training system which tunes the hyperparameters is what trains the machine learning models (Paragraph 18-19). Therefore, the training process as seen in Carbune, and for that reason Carbune’s hyperparameter tuning is part of its training process.
Carbune does not specifically teach the hardware-level and model-level settings of Shah in view of Ghosh in view of Tiku, nor that there are a plurality of values as seen in the multiple settings. However, those settings may be integrated with the teachings of Carbune for reasons of the advantages described below.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to implement a system utilizing the teachings of Shah, the teachings of Ghosh, and the teachings of Tiku. This would have provided the advantage of improvements to hardware functionality such as dynamic routing and load balancing of network traffic (Ghosh, Paragraph 3) as well as improving hardware utilizations (Tiku, Paragraph 16) as well as improving the efficacy of hyperparameters while also saving time and resources (Carbune, Paragraph 5).
Regarding claim 2:
Claim 2 recites:
The system of claim 1, wherein the one or more model-level performance settings comprise one or more of: an input data size for input data to the machine learning model, one or more model hyperparameters specifying a size or shape of the machine learning model; or one or more training process hyperparameters modifying the training process implemented by the one or more processors for training the machine learning model.
Shah in view of Ghosh further in view of Tiku further in view of Carbune teach the system of claim 1 upon which claim 2 depends. Furthermore, regarding the limitation wherein the one or more model-level performance settings comprise one or more of: an input data size for input data to the machine learning model, one or more model hyperparameters specifying the size or shape of the machine learning model, and one or more training process hyperparameters modifying the training process implemented by the one or more processors for training the machine learning model:
Shah teaches that hyperparameters or setting may include a number of layers, input, and outputs (Paragraph 12) which would describe the size and shape of a machine learning model, as well as an epoch value (Paragraph 12), which would be a training process hyperparameter modifying the training process implement by the one or more processors for training the model as epochs are related to the process of training. These two examples of hyperparameters would be one or more model-level performance settings of the above list.
Regarding claim 3:
Claim 3 recites:
The system of claim 1, wherein the one or more hardware-level performance settings comprise settings for adjusting intra- or inter-data communication between the one or more processors.
Shah in view of Ghosh further in view of Tiku further in view of Carbune teach the system of claim 1 upon which claim 3 depends. Furthermore, regarding the limitation of claim 3:
Ghosh teaches receiving bandwidth information and a traffic volume classification to determine routing recommendation (Paragraph 21). The bandwidth information and traffic volume classification would be hardware-level performance settings for the network that comprise settings for intra or inter-data communication between one or more processors, wherein the routing recommendation would be the adjustments. Therefore Ghosh teaches the above hardware settings, which could be used in combination with the system of Shah to implement them as hyperparameters.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to implement a system utilizing the teachings of Shah, the teachings of Ghosh, and the teachings of Tiku. This would have provided the advantage of improvements to hardware functionality such as dynamic routing and load balancing of network traffic (Ghosh, Paragraph 3) as well as improving hardware utilizations (Tiku, Paragraph 16) as well as improving the efficacy of hyperparameters while also saving time and resources (Carbune, Paragraph 5).
Regarding claim 4:
Claim 4 recites:
The system of claim 3, wherein the one or more processors comprise a plurality of processors logically or physically grouped into a plurality of groups, and the one or more hardware-level performance settings comprise settings for a rate of inter-data communication between processors in different groups
Shah in view of Ghosh further in view of Tiku further in view of Carbune teach the system of claim 3 upon which claim 4 depends. Furthermore, regarding the limitation of claim 4:
Ghosh teaches receiving bandwidth information and a traffic volume classification to determine routing recommendation (Paragraph 21). The bandwidth information and traffic volume classification would be hardware-level performance settings for the network that comprise settings for intra or inter-data communication between one or more processors, wherein the routing recommendation would be the adjustments. Furthermore, Ghosh teaches that one or more processors are operably couples with the memory (Paragraph 35) which would be a plurality of processors grouped.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to implement a system utilizing the teachings of Shah, the teachings of Ghosh, and the teachings of Tiku. This would have provided the advantage of improvements to hardware functionality such as dynamic routing and load balancing of network traffic (Ghosh, Paragraph 3) as well as improving hardware utilizations (Tiku, Paragraph 16) as well as improving the efficacy of hyperparameters while also saving time and resources (Carbune, Paragraph 5).
Regarding claim 7:
Claim 7 recites:
The system of claim 1, wherein in training the machine learning model, the one or more processors are further configured to: set the one or more hardware-level and model-level performance settings to first values of the plurality of values of the training schedule; and at a first point in time setting the one or more hardware-level and model-level performance settings to first values but before the completion of the training process, adjust the one or more hardware-level and one or more model-level performance settings to second values of the plurality of values different from the first values.
Shah in view of Ghosh further in view of Tiku teach the system of claim 1 upon which claim 7 depends. Furthermore, regarding the limitation wherein in training the machine learning model, the one or more processors are further configured to: set the one or more hardware-level and model-level performance settings to first values of the plurality of values of the training schedule:
Shah teaches using hyperparameters settings to configure a plurality of machine learning models (Paragraph 21) wherein the first machine learning model would be set to the first values of the plurality of values of the training schedule.
Carbune discloses at a first point in time setting the one or more hardware-level and model-level performance settings to first values but before the completion of the training process, adjust the one or more hardware-level and one or more model-level performance settings to second values of the plurality of values different from the first values:
Carbune in the same field of endeavor of machine learning teaches that hyperparameters can be configurable parameters of the generation process (Paragraph 26) wherein the hyperparameters would include the one or more hardware-level and model-level performance settings. Configuring a parameter in a generation process would be giving it a value during an initialization phase, or setting it to a first value a first point in time before the completion of the training process.
Furthermore, Carbune teaches optimizing the hyperparameters that guide the process of generating adversarial examples (Paragraph 18), wherein optimizing would be analogous to adjusting the performance setting (or the hyperparameters) after initiation of the training of the machine learning model to a second value, wherein the second value is the next step in the optimization of the hyperparameters.
Carbune and the present application are analogous art because they are both in the same field of machine learning
It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to implement a system utilizing the teachings of Shah in view of Ghosh further in view of Tiku and the teachings of Carbune. This would have provided the advantage of improving the efficacy of hyperparameters while also saving time and resources (Carbune, Paragraph 5).
Regarding claim 8:
Claim 8 recites:
The system of claim 1, wherein the one or more processors are further configured to receive the training schedule from a training schedule machine learning model.
Shah in view of Ghosh further in view of Tiku further in view of Carbune teach the system of claim 1 upon which claim 8 depends. Furthermore, Carbune discloses the limitation of claim 8
Carbune teaches the use of federated learning (Paragraph 23) which describes a central model is generated through a system of client models collaborating with the central model. The system of client models would be analogous to the training schedule machine learning model as they generate the training schedule, or the settings of the central model to be used when it runs. Therefore the training schedule is received as it is generated.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to implement a system utilizing the teachings of Shah in view of Ghosh further in view of Tiku and the teachings of Carbune. This would have provided the advantage of improving the efficacy of hyperparameters while also saving time and resources (Carbune, Paragraph 5).
Regarding claim 9:
Claim 9 recites:
The system of claim 1, wherein the machine learning model is a neural network.
Shah in view of Ghosh further in view of Tiku further in view of Carbune teach the system of claim 1 upon which claim 9 depends. Furthermore, Shah discloses the limitation of claim 9:
Shah teaches a machine learning model which is a neural network (Paragraph 12).
Regarding claim 10:
Claim 10 recites:
The system of claim 1, wherein in receiving the training schedule, the one or more processors are further configured to: send a query to one or more memory devices storing a plurality of candidate training schedules, the query comprising data at least partially describing the machine learning model, and computing resources available for training the machine learning model; and receive the training schedule from the plurality of candidate training schedules in response to the query
Shah in view of Ghosh further in view of Tiku further in view of Carbune teach the system of claim 1 upon which claim 10 depends. Furthermore, regarding the limitation wherein in receiving the training schedule, the one or more processors are further configured to: send a query to one or more memory devices storing a plurality of candidate training schedules, the query comprising data at least partially describing the machine learning model, and computing resources available for training the machine learning model:
Shah teaches that the user input may describe the hyperparameters and source data to be used in the model, which leads the network device to select the model type (Paragraph 21). The model type would be analogous to the candidate training schedules as it affects the training of the model, while the user input would include the query in the form of the description of the model and the source data for the machine learning task that leads to the network device making its selection. The network device also receives data regarding the hardware configuration of the device, as the hardware configuration (or computing resources available for training the model) are part of the model selection process (Paragraph 3).
Regarding the limitation receive the training schedule from the plurality of candidate training schedules in response to the query:
Shah teaches the selection of the model type (Paragraph 21), or training schedule, which would consist of receiving the training schedule for the plurality of candidate training schedules as a selection for use in the model would require obtaining the training schedule itself.
Claims 11-14, 16-18 recite a method that parallels the system of claims 1-4 and 8-10 respectively. Therefore, the analysis discussed above with respect to claims 1-4 and 8-10 also applies to claims 11-14 and 16-18 respectively. Accordingly, claims 11-14 and 16-18 are rejected based on substantially the same rationale as set forth above with respect to claims 1-4 and 8-10 respectively.
Claim 20 recites a non-transitory computer readable storage medium that parallels the system of claim 1. Therefore, the analysis discussed above with respect to claim 1 also applies to claim 20. Accordingly, claim 20 is rejected based on substantially the same rationale as set forth above with respect to claim 1.
Claims 5-6 and 13 are rejected under 35 U.S.C. 103 as being unpatentable over Shah in view of Ghosh, further in view of Tiku, further in view of Carbune, further in view of Brady et al. (Pub. No. US 20190392296 A1, filed June 28th 2019, hereinafter Brady).
Regarding claim 5:
Claim 5 recites:
The system of claim 3, wherein the one or more hardware-level performance settings comprise settings for adjusting numerical precision of operations performed by the one or more processors
Shah in view of Ghosh further in view of Tiku further in view of Carbune teach the system of claim 3 upon which claim 5 depends. However, neither Shah nor Ghosh nor Tiku nor Carbune teaches the limitation of claim 5:
Brady in the same field of endeavor of machine learning teaches a target descriptor file that may include settings for data precision of a target which would be the numerical precision of operation performed by the one or more processors (Paragraph 78). This file is additionally used in combination with a generated neural network in order to optimize said neural network for specific targets (Paragraph 46), which would be adjusting the settings.
Brady and the present application are analogous art because they are both in the same field of machine learning
It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to implement a system utilizing the teachings of Shah in view of Ghosh further in view of Tiku further in view of Carbune and the teachings of Brady. This would have provided the advantage of efficient and effective implementation of machine learning models (Brady, Paragraph 2).
Regarding claim 6:
Claim 6 recites:
The system of claim 3, wherein the one or more hardware-level performance settings comprise settings for enabling or disabling hardware parallelism among the one or more processors.
Shah in view of Ghosh further in view of Tiku further in view of Carbune teach the system of claim 3 upon which claim 6 depends. However, neither Shah nor Ghosh nor Tiku nor Carbune teaches the limitation of claim 6:
Brady teaches a control model that may enable parallel execution for parts of a neural network (Paragraph 55). In combination with the hardware-level performance settings used, of Shah in view of Ghosh, the enablement of parallel execution can be used as a setting.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to implement a system utilizing the teachings of Shah in view of Ghosh further in view of Tiku further in view of Carbune and the teachings of Brady. This would have provided the advantage of efficient and effective implementation of machine learning models (Brady, Paragraph 2).
Claim 15 recites a method that parallels the system of claim 5. Therefore, the analysis discussed above with respect to claim 5 also applies to claim 15. Accordingly, claim 15 is rejected based on substantially the same rationale as set forth above with respect to claim 5.
Response to Arguments
Applicant’s arguments filed 22-JANUARY-2026 have been fully considered, but the examiner believes that not all are fully persuasive.
Regarding the applicant’s remarks on the non-final office action’s 103 rejection of the claims, the applicant argues that Shah in view of Ghosh, further in view of Tiku does not teach the amended limitations of these claims. As such, the applicant argues that all claims dependent on the above would additionally not be obvious under 103. However, the examiner believes that Shah in view of Ghosh, further in view of Tiku, further in view of Carbune does teach the amended limitations and respectfully requests applicant’s consideration of the following:
The applicant argues that Carbune does not teach the amended limitations because (a) Carbune does not tune the hyperparameters during the training step as they are tuned beforehand, and (b) Carbune does not initially select training speed and then adjust the values to select for model quality. The examiner addresses both points below.
With regards to (a), Carbune teaches that the adversarial training system which tunes the hyperparameters is what trains the machine learning models (Paragraph 18-19). Therefore, the examiner believes that the adversarial training system is the training process as seen in Carbune, and for that reason Carbune’s hyperparameter tuning is part of its training process.
With regards to (b), Carbune teaches a step size hyperparameter that is iteratively reduced until a realism score exceeds a threshold score (Paragraph 87). The iterative reduction would be a first value at a first point in time during the training process that becomes a second value at a second point in time after the first point in time during the training process. Furthermore, Carbune teaches that the threshold may also be used to determine an appropriate balance between speed and realism i.e. model quality (Paragraph 120), and therefore initial values in an embodiment of Carbune may select for training speed over model quality at a first point in time, and model quality over training speed at a second point in time.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ALEXANDRIA JOSEPHINE MILLER whose telephone number is (703)756-5684. The examiner can normally be reached Monday-Thursday: 7:30 - 5:00 pm, every other Friday 7:30 - 4:00.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Mariela Reyes can be reached at (571) 270-1006. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/A.J.M./Examiner, Art Unit 2142
/Mariela Reyes/Supervisory Patent Examiner, Art Unit 2142