Prosecution Insights
Last updated: October 01, 2026
Application No. 18/362,123

METHOD AND SYSTEM FOR SWITCHING BETWEEN HARDWARE ACCELERATORS FOR DATA MODEL TRAINING

Final Rejection §103§112
Filed
Jul 31, 2023
Priority
Aug 30, 2022 — IN 202221049467
Examiner
HADDAD, MAJD MAHER
Art Unit
2125
Tech Center
2100 — Computer Architecture & Software
Assignee
Tata Group
OA Round
2 (Final)
100%
Grant Probability
Favorable
3-4
OA Rounds
2m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 100% — above average
100%
Career Allowance Rate
5 granted / 5 resolved
+45.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
3y 4m
Avg Prosecution
21 currently pending
Career history
30
Total Applications
across all art units

Statute-Specific Performance

§101
29.0%
-11.0% vs TC avg
§103
51.2%
+11.2% vs TC avg
§102
3.1%
-36.9% vs TC avg
§112
14.2%
-25.8% vs TC avg
Black line = Tech Center average estimate • Based on career data from 5 resolved cases

Office Action

§103 §112
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . This action is in response to the amendment and remarks filed July 1st, 2025. In the amendment, claims 1, 3, and 5 were amended, no claims were cancelled or added. As such, claims 1-6 are presented for examination. Response to Arguments Applicant’s arguments, see Page 7, filed March June 30th, 2026, with respect to the claim objections have been considered and are persuasive. Amendments to the claims obviate the objections of record. The objections of claims have been withdrawn. See updated objections below. Applicant’s arguments with respect to the rejection under 35 U.S.C. 103 are not persuasive for the following reasons: 35 U.S.C 103: Applicant argues that Ma fails to teach determining whether the difference in accuracy between two consecutive epochs exceeds a threshold, because Ma teaches only loss-based convergence used for termination (Pages 8 to 9 of Remarks). The Examiner respectfully disagrees. Ma discloses that SGD is stopped when there is no significant drop in the loss across iterations, which is a determination of whether the change between two consecutive epochs is above or below a significance threshold. Because lower loss indicates higher prediction accuracy in Ma, evaluating whether the loss change is significant across consecutive epochs corresponds to determining whether the accuracy difference between two consecutive epochs exceeds a threshold. Ma therefore teaches the claimed determining limitation. Applicant argues that Ma does not teach a mechanism where an accuracy-based comparison governs switching from one hardware accelerator to another during ongoing training (Page 9 of Remarks). The Examiner respectfully disagrees. The rejection does not rely on Ma alone for the switching limitation. Ma is relied upon for the epoch-based accuracy determination, and Wheatley is relied upon for switching the training to a next hardware accelerator. The combination of Ma and Wheatley therefore teaches this limitation. Applicant argues that Wheatley teaches only configuration-based, deployment level switching and not accuracy-driven switching within an iterative epoch-based training loop (Page 10 of Remarks). The Examiner respectfully disagrees. Wheatley is relied upon only for the act of switching the training to a next hardware accelerator, and the accuracy-based per-epoch evaluation that triggers the switch is supplied by Ma. It is the combination that teaches accuracy-driven switching within an epoch-based loop and switching between hardware accelerators. The argument attacks Wheatley alone and does not address the combined teachings relied upon in the rejection. Applicant argues that even in combination Ma and Wheatley fail to teach threshold-based switching triggered by accuracy differences between epochs (Page 10 of Remarks). The Examiner respectfully disagrees. Ma teaches evaluating whether the loss changes significantly across consecutive epochs against a threshold, and Wheatley teaches switching the model training from one processor to a next processor. A person of ordinary skill combining these teachings would arrive at switching the training to the next accelerator based on the epoch-based accuracy determination taught by Ma in order to enhance training efficiency and resource utilization as suggested by Wheatley at Paragraph 156. The combination therefore teaches the claimed threshold-based switching. Applicant argues that the cited art fails to teach defining maximum accuracy based on stability over a predefined number of epochs, terminating training even when unutilized accelerators remain, and continuing training on the final accelerator even when accuracy is below the threshold (Pages 9 to 12 of Remarks). The Examiner respectfully disagrees. Applicant’s arguments have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument. The rejection now relies on Ananthanarayanan to teach terminating training even when there are unutilized hardware accelerators in the sequence and continuing training on the final hardware accelerator when the measured accuracy is less than the threshold. Ma additionally teaches the maximum accuracy limitation by stopping training when there is no significant drop in the loss across iterations, which corresponds to the accuracy remaining constant or changing minimally over a number of epochs. Applicant argues that a person of ordinary skill would lack motivation to arrive at the claimed invention and that the combination is conclusory under MPEP 2141 III (Page 13 of Remarks). The Examiner respectfully disagrees. The rejection provides an articulated reason drawn from the references themselves for combining Ma and Wheatley to enhance training efficiency and resource utilization, and further combines Ananthanarayanan to allocate computing resources efficiently and achieve higher accuracy for a given amount of accelerator resource as taught at Paragraph 40. This reasoning is based on an express teaching, suggestion, or motivation found in the prior art under MPEP 2143(I)(G). Each reference is in the same field of endeavor of training machine learning models on hardware accelerators. Applicant argues that Kang in claims 2, 4, and 6 fails to teach arranging the plurality of hardware accelerators in a predefined sequence of increasing order of specification for use during model training, and that Kang merely discloses heterogeneous processors and the assignment of tasks across such processors (Pages 13 to 14 of Remarks). The Examiner respectfully disagrees. The claim does not define the term specification, and under its broadest reasonable interpretation the term reads on any ordering of the processing elements by a processor characteristic such as clock frequency, throughput, profiled execution time, or precision. Kang discloses heterogeneous processors together with their specifications including a Mali-G72 MP18 GPU and big.LITTLE CPUs with a quad-core M3 CPU at 2.7 GHz and a quad-core Cortex-A55 CPU at 1.79 GHz, and defines a set of logical processing elements PE={PE1, PE2, ..., PEm} that the scheduler proceeds through based on each element's profiled performance, from a lower performing element to a higher performing one. Ordering the processing elements by these profiled specifications and proceeding through them in that order corresponds to arranging the plurality of hardware accelerators in a sequence of increasing order of specification used during model training. Specification The disclosure is objected to because of the following informalities: Paragraph 4: "Some of the tasks maybe more process intensive and some maybe less process intensive" should read "Some of the tasks may be more process intensive and some may be less process intensive". Paragraph 29: “…more number of hardware accelerators maybe used as per implementation… The hardware accelerators used maybe of different 20 specifications” should read “…more number of hardware accelerators may be used as per implementation… The hardware accelerators used may be of different specifications.” Paragraph 30: “Subsequently the remaining hardware accelerators maybe termed as second hardware accelerator…” should read “Subsequently the remaining hardware accelerators may be termed as second hardware accelerator.” Paragraph 42: "(Recall@20, MRR of 53.49, 18.67) 6" contains a stray number 6 that should be deleted. Appropriate correction is required. Claim Objections Claims 1-6 are objected to because of the following informalities: The limitation "iteratively performing until … b) a final hardware accelerator in a sequence of the plurality of heterogeneous hardware accelerators has reached" should read "iteratively performing until... b) a final hardware accelerator in a sequence of the plurality of heterogeneous hardware accelerators has been reached." The limitation "determining if difference between the accuracy…" should read "determining if a difference between the accuracy..." Appropriate correction is required. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 1-6 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. The term “minimal” in claims 1, 3, and 5 are relative terms which render the claims indefinite. The term “minimal” is not defined by the claim and the specification does not provide a standard for ascertaining the requisite degree. At most, paragraph 30 of the instant specification discloses that “the measured accuracy of the data model is minimal (which may be decided in terms of a threshold of change) over the pre-defined number of epochs”. One of ordinary skill in the art would not be reasonably apprised of the scope of the invention. For purposes of examination, the Examiner interprets "minimal" under its broadest reasonable interpretation as any change in the measured accuracy of the data model that falls at or below a threshold and reaches a plateau. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention. Claims 1, 3, and 5 are rejected under 35 U.S.C. 103 as being unpatentable over Ma (“Heterogeneous CPU+GPU Stochastic Gradient Descent Algorithms”, 2020) in view of Wheatley (US 20240352839 A1) in further view of Ananthanarayanan (US 20220188569 A1). Regarding claim 1, Ma teaches [a] processor implemented method, comprising: initiating a data model training using a first hardware accelerator from among a plurality of heterogeneous hardware accelerators (Page 7 Section 5, “CPUs and GPUs perform concurrent asynchronous SGD algorithms– specialized for their specific architecture– on data assigned dynamically and adaptively at runtime based on the current execution state.”, Page 7 Section 5.1 of Ma, “The architecture of the heterogeneous CPU+GPU framework for deep learning is depicted in Figure 3. It consists of a series of asynchronous worker threads corresponding to each of the CPUs and GPUs in Figure 2, and a central coordinator. The worker assigned to a hardware component is in charge of managing the resources, e.g., cores, memory, threads, and operation of that component.” Ma teaches training a deep learning model using a first hardware accelerator (CPU or GPU worker) from a set of heterogeneous accelerators.) wherein the data model training is spread across a plurality of epochs (Page 5 Section 3, “SGD can be stopped either after a fixed number of iterations, i.e., epochs, or when there is no significant drop in the loss across iterations. In practice, due to the large data set size and number of iterations it takes to converge, each SGD iteration is performed only over a randomly selected batch of B training examples…”, Page 10 Section 5.2, “As such, the framework performs loss computation after each complete pass– or a given number of batches– over the training data.” Ma teaches training and performing loss computation after each complete pass across multiple epochs.); and iteratively performing until one of a) a maximum accuracy is achieved for the data model, and b) a final hardware accelerator in a sequence of the plurality of heterogeneous hardware accelerators has reached (Page 5 Section 3, “SGD can be stopped either after a fixed number of iterations, i.e., epochs, or when there is no significant drop in the loss across iterations." Ma teaches that the model is iteratively performed until the loss is minimized, meaning that the accuracy is at its maximum level.): measuring accuracy of the data model after each of the plurality of epochs (Page 5 Section 3, “SGD can be stopped either after a fixed number of iterations, i.e., epochs, or when there is no significant drop in the loss across iterations.", Page 10 Section 5.2, “As such, the framework performs loss computation after each complete pass– or a given number of batches– over the training data. The loss is computed with a DNN forward pass over the training– or test– data… The size of the batch is proportional to the worker speed… Each worker computes a partial loss on its data batch and then sends it back to the coordinator, which aggregates it into the overall loss. This strategy is optimized for execution time by prioritizing the fast workers and minimizing the coordinator overhead.” The loss is measured after each epoch and the SGD is stopped once there is minimal change in the loss across epochs, meaning that the accuracy has reached its highest point.); determining if difference between the accuracy of the data model measured in two consecutive epochs of the plurality of epochs exceeds a threshold of accuracy (Page 5 Section 3, “SGD can be stopped either after a fixed number of iterations, i.e., epochs, or when there is no significant drop in the loss across iterations." Ma stops training when there is no significant drop in loss between successive epochs, and a drop in loss corresponds to a gain in accuracy. The significance criterion is the threshold such that a significant loss drop corresponds to an accuracy difference exceeding the threshold and an insignificant drop corresponds to an accuracy difference below it.); if the difference between the accuracy of the data model measured in two consecutive epochs of the plurality of epochs exceeds the threshold of accuracy (Page 5 Section 3, “SGD can be stopped either after a fixed number of iterations, i.e., epochs, or when there is no significant drop in the loss across iterations." Ma stops training when there is no significant drop in loss between successive epochs, and a drop in loss corresponds to a gain in accuracy. The significance criterion is the threshold such that a significant loss drop corresponds to an accuracy difference exceeding the threshold and an insignificant drop corresponds to an accuracy difference below it.) wherein when the measured accuracy of the data model remains constant over a pre-defined number of epochs or change in the measured accuracy of the data model is minimal over the pre-defined number of epochs, then the data model achieved the maximum accuracy… the model training is terminated (Page 5 Section 3, “SGD can be stopped either after a fixed number of iterations, i.e., epochs, or when there is no significant drop in the loss across iterations." Ma teaches stopping SGD once there is no significant drop in the loss across iterations, meaning that the accuracy changes minimally across epochs and that the model achieved a maximum accuracy.) Ma does not teach and switching the data model training to next hardware accelerator in the sequence of the plurality of heterogeneous hardware accelerators, if the difference between the accuracy of the data model measured in two consecutive epochs of the plurality of epochs is exceeding the threshold of accuracy. Wheatley, in the same field of endeavor, teaches and switching the data model training to a next hardware accelerator … (Paragraph 184, “As an example, a convolution neural network system (CNNS) can be implemented using one or more platforms… The GAFFE framework includes models and optimization defined by configuration where switching between CPU and GPU setting can be achieved via a single flag to train on a GPU machine then deploy to clusters.” Wheatley teaches switching a neural-network computation from a CPU to a GPU setting via a single configuration during training.) Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date to combine Ma’s teaching of learning data model training across heterogeneous CPU and GPU accelerators with Wheatley’s teaching of switching processors during training in a neural network in order to enhance training efficiency and resource utilization (Paragraph 156 of Wheatley). Although Ma in view of Wheatley substantially teaches the claimed invention, Ma in view of Wheatley is not relied on to teach and even when there are unutilized hardware accelerators in the sequence, the model training is terminated. However, Ananthanarayanan, in the same field of endeavor, teaches and even when there are unutilized hardware accelerators in the sequence, the model training is terminated (Paragraph 86, “the micro-profiler tests each configuration of a superset of configurations for a small number (e.g., 5) of training epochs, and then terminates the testing early, well before convergence… After early termination on the sampled training data, the (validation) accuracy of each configuration is obtained at each epoch it was trained.”, Paragraph 92, “To avoid expensive inter-GPU communication, the allocations may be quantized to inverse powers of two (e.g., ½, ¼, ⅛). This may make the jobs amenable to packing. The jobs may then be allocated to GPUs in descending order of demands to reduce fragmentation.”, Paragraph 38, “In each retraining window, the resource scheduler makes the decisions described above to (1) decide which of the edge models to retrain; (2) allocate the edge server's GPU resources among the retraining and inference jobs, and (3) select the configurations of the retraining and inference jobs… The scheduler decides against retraining the models which do not improve a target metric… First, it can simplify the spatial complexity by considering GPU allocations in coarse fractions (e.g., 10%) that are accurate enough for the scheduling decisions, while also being mindful of the granularity achievable in modern GPUs. Second, it can avoid changing allocations to jobs during the re-training, which helps to avoid temporal complexity.”, Paragraph 143, “Processors of the logic machine may be single-core or multi-core, and the instructions executed thereon may be configured for sequential, parallel, and/or distributed processing.” Ananthanarayanan performs the deep-learning training on GPUs and FPGAs, which are the specialized hardware components that accelerate the model training, which corresponds to the hardware accelerators. Ananthanarayanan's scheduler allocates these GPU/FPGA resources to jobs in fractions and decides against continuing to retrain models that do not improve a target metric, terminating training early before convergence. Because the scheduler apportions the accelerators fractionally, a given accelerator may receive no tasks, i.e. remain unutilized where the termination of training can happen.) and wherein when the final hardware accelerator in the sequence of the plurality of heterogeneous hardware accelerators is reached, then the model training is continued with the hardware accelerator even when the measured accuracy is less than the threshold of accuracy in the subsequent epochs (Paragraph 92, “To avoid expensive inter-GPU communication, the allocations may be quantized to inverse powers of two (e.g., ½, ¼, ⅛). This may make the jobs amenable to packing. The jobs may then be allocated to GPUs in descending order of demands to reduce fragmentation.”, Paragraph 94, “Every few epochs (e.g., every 5 epochs), the current accuracy of the model being retrained is used to estimate its eventual accuracy when all the epochs are complete.”, Paragraph 38, “In each retraining window, the resource scheduler makes the decisions described above to (1) decide which of the edge models to retrain; (2) allocate the edge server's GPU resources among the retraining and inference jobs, and (3) select the configurations of the retraining and inference jobs… it can simplify the spatial complexity by considering GPU allocations in coarse fractions (e.g., 10%) that are accurate enough for the scheduling decisions…”, Paragraph 105, “The capacity of the thief scheduling approach (e.g., the maximum number of concurrent video streams subject to an accuracy threshold is compared with that of the uniform baseline, as more GPUs are available. An accuracy threshold may be set, since some applications may not be usable when accuracy is below the threshold in some instances.”, Paragraph 35, “Implementing continuous retraining may involve making the following decisions: (1) in each retraining window, decide which edge models of a plurality of models to retrain; (2) allocate the edge server's GPU resources among the retraining and inference jobs, and (3) select the configurations of the retraining and inference jobs. Decisions may also be constrained such that the inference accuracy at any point in time does not drop below a minimum value (so that the outputs continue to remain useful to the application).” Ananthanarayanan allocates jobs across the GPUs in descending order such that the last allocated accelerator in that ordering is reached while training continues and the current accuracy is monitored every few epochs against a set accuracy threshold. Reaching the last accelerator in the ordering and continuing to train it while accuracy is under the threshold corresponds to continuing training on the final accelerator even when the measured accuracy is less than the threshold in subsequent epochs.) Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date to combine Ma in view of Wheatley's data model training and switching across heterogeneous accelerators with Ananthanarayanan's per-epoch accuracy driven scheduling that allocates accelerators in fractional ordered amounts and terminates models which do not improve a target metric in order to allocate computing resources efficiently and achieve higher accuracy for a given amount of accelerator resource (Paragraph 40 of Ananthanarayanan). Regarding claim 3, Ma teaches [a] system, comprising: one or more hardware processors; a communication interface; and a memory storing a plurality of instructions, wherein the plurality of instructions when executed, cause the one or more hardware processors to (Page 7 Section 5.1 of Ma, “The architecture of the heterogeneous CPU+GPU framework for deep learning is depicted in Figure 3. It consists of a series of asynchronous worker threads corresponding to each of the CPUs and GPUs in Figure2, and a central coordinator… The coordinator and workers are implemented as stand-alone system threads that exist over the entire duration of the program. The worker assigned to a hardware component is in charge of managing the resources, e.g., cores, memory, threads, and operation of that component. The coordinator assigns data and tasks to workers, and schedules their interaction. The communication between the coordinator and workers– workers do not communicate directly– is realized through control messages, while data are passed through references in the shared memory space.”): initiate a data model training using a first hardware accelerator from among a plurality of heterogeneous hardware accelerators (Page 7 Section 5, “CPUs and GPUs perform concurrent asynchronous SGD algorithms– specialized for their specific architecture– on data assigned dynamically and adaptively at runtime based on the current execution state.”, Page 7 Section 5.1 of Ma, “The architecture of the heterogeneous CPU+GPU framework for deep learning is depicted in Figure 3. It consists of a series of asynchronous worker threads corresponding to each of the CPUs and GPUs in Figure 2, and a central coordinator. The worker assigned to a hardware component is in charge of managing the resources, e.g., cores, memory, threads, and operation of that component.” Ma teaches training a deep learning model using a first hardware accelerator (CPU or GPU worker) from a set of heterogeneous accelerators.) wherein the data model training is spread across a plurality of epochs (Page 5 Section 3, “SGD can be stopped either after a fixed number of iterations, i.e., epochs, or when there is no significant drop in the loss across iterations. In practice, due to the large data set size and number of iterations it takes to converge, each SGD iteration is performed only over a randomly selected batch of B training examples…”, Page 10 Section 5.2, “As such, the framework performs loss computation after each complete pass– or a given number of batches– over the training data.” Ma teaches training and performing loss computation after each complete pass across multiple epochs.); and iteratively performing until one of a) a maximum accuracy is achieved for the data model, and b) a final hardware accelerator in a sequence of the plurality of heterogeneous hardware accelerators has reached (Page 5 Section 3, “SGD can be stopped either after a fixed number of iterations, i.e., epochs, or when there is no significant drop in the loss across iterations." Ma teaches that the model is iteratively performed until the loss is minimized, meaning that the accuracy is at its maximum level.): measuring accuracy of the data model after each of the plurality of epochs (Page 5 Section 3, “SGD can be stopped either after a fixed number of iterations, i.e., epochs, or when there is no significant drop in the loss across iterations.", Page 10 Section 5.2, “As such, the framework performs loss computation after each complete pass– or a given number of batches– over the training data. The loss is computed with a DNN forward pass over the training– or test– data… The size of the batch is proportional to the worker speed… Each worker computes a partial loss on its data batch and then sends it back to the coordinator, which aggregates it into the overall loss. This strategy is optimized for execution time by prioritizing the fast workers and minimizing the coordinator overhead.” The loss is measured after each epoch and the SGD is stopped once there is minimal change in the loss across epochs, meaning that the accuracy has reached its highest point.); determining if difference between the accuracy of the data model measured in two consecutive epochs of the plurality of epochs exceeds a threshold of accuracy (Page 5 Section 3, “SGD can be stopped either after a fixed number of iterations, i.e., epochs, or when there is no significant drop in the loss across iterations." Ma stops training when there is no significant drop in loss between successive epochs, and a drop in loss corresponds to a gain in accuracy. The significance criterion is the threshold such that a significant loss drop corresponds to an accuracy difference exceeding the threshold and an insignificant drop corresponds to an accuracy difference below it.); if the difference between the accuracy of the data model measured in two consecutive epochs of the plurality of epochs exceeds the threshold of accuracy (Page 5 Section 3, “SGD can be stopped either after a fixed number of iterations, i.e., epochs, or when there is no significant drop in the loss across iterations." Ma stops training when there is no significant drop in loss between successive epochs, and a drop in loss corresponds to a gain in accuracy. The significance criterion is the threshold such that a significant loss drop corresponds to an accuracy difference exceeding the threshold and an insignificant drop corresponds to an accuracy difference below it.) wherein when the measured accuracy of the data model remains constant over a pre-defined number of epochs or change in the measured accuracy of the data model is minimal over the pre-defined number of epochs, then the data model achieved the maximum accuracy… the model training is terminated (Page 5 Section 3, “SGD can be stopped either after a fixed number of iterations, i.e., epochs, or when there is no significant drop in the loss across iterations." Ma teaches stopping SGD once there is no significant drop in the loss across iterations, meaning that the accuracy changes minimally across epochs and that the model achieved a maximum accuracy.) Ma does not teach and switching the data model training to next hardware accelerator in the sequence of the plurality of heterogeneous hardware accelerators, if the difference between the accuracy of the data model measured in two consecutive epochs of the plurality of epochs is exceeding the threshold of accuracy. Wheatley, in the same field of endeavor, teaches and switching the data model training to a next hardware accelerator … (Paragraph 184, “As an example, a convolution neural network system (CNNS) can be implemented using one or more platforms… The GAFFE framework includes models and optimization defined by configuration where switching between CPU and GPU setting can be achieved via a single flag to train on a GPU machine then deploy to clusters.” Wheatley teaches switching a neural-network computation from a CPU to a GPU setting via a single configuration during training.) Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date to combine Ma’s teaching of learning data model training across heterogeneous CPU and GPU accelerators with Wheatley’s teaching of switching processors during training in a neural network in order to enhance training efficiency and resource utilization (Paragraph 156 of Wheatley). Although Ma in view of Wheatley substantially teaches the claimed invention, Ma in view of Wheatley is not relied on to teach and even when there are unutilized hardware accelerators in the sequence, the model training is terminated. However, Ananthanarayanan, in the same field of endeavor, teaches and even when there are unutilized hardware accelerators in the sequence, the model training is terminated (Paragraph 86, “the micro-profiler tests each configuration of a superset of configurations for a small number (e.g., 5) of training epochs, and then terminates the testing early, well before convergence… After early termination on the sampled training data, the (validation) accuracy of each configuration is obtained at each epoch it was trained.”, Paragraph 92, “To avoid expensive inter-GPU communication, the allocations may be quantized to inverse powers of two (e.g., ½, ¼, ⅛). This may make the jobs amenable to packing. The jobs may then be allocated to GPUs in descending order of demands to reduce fragmentation.”, Paragraph 38, “In each retraining window, the resource scheduler makes the decisions described above to (1) decide which of the edge models to retrain; (2) allocate the edge server's GPU resources among the retraining and inference jobs, and (3) select the configurations of the retraining and inference jobs… The scheduler decides against retraining the models which do not improve a target metric… First, it can simplify the spatial complexity by considering GPU allocations in coarse fractions (e.g., 10%) that are accurate enough for the scheduling decisions, while also being mindful of the granularity achievable in modern GPUs. Second, it can avoid changing allocations to jobs during the re-training, which helps to avoid temporal complexity.” Ananthanarayanan performs the deep-learning training on GPUs and FPGAs, which are the specialized hardware components that accelerate the model training, which corresponds to the hardware accelerators. Ananthanarayanan's scheduler allocates these GPU/FPGA resources to jobs in fractions and decides against continuing to retrain models that do not improve a target metric, terminating training early before convergence. Because the scheduler apportions the accelerators fractionally, a given accelerator may receive no tasks, i.e. remain unutilized where the termination of training can happen.) and wherein when the final hardware accelerator in the sequence of the plurality of heterogeneous hardware accelerators is reached, then the model training is continued with the hardware accelerator even when the measured accuracy is less than the threshold of accuracy in the subsequent epochs (Paragraph 92, “To avoid expensive inter-GPU communication, the allocations may be quantized to inverse powers of two (e.g., ½, ¼, ⅛). This may make the jobs amenable to packing. The jobs may then be allocated to GPUs in descending order of demands to reduce fragmentation.”, Paragraph 94, “Every few epochs (e.g., every 5 epochs), the current accuracy of the model being retrained is used to estimate its eventual accuracy when all the epochs are complete.”, Paragraph 38, “In each retraining window, the resource scheduler makes the decisions described above to (1) decide which of the edge models to retrain; (2) allocate the edge server's GPU resources among the retraining and inference jobs, and (3) select the configurations of the retraining and inference jobs… it can simplify the spatial complexity by considering GPU allocations in coarse fractions (e.g., 10%) that are accurate enough for the scheduling decisions…”, Paragraph 105, “The capacity of the thief scheduling approach (e.g., the maximum number of concurrent video streams subject to an accuracy threshold is compared with that of the uniform baseline, as more GPUs are available. An accuracy threshold may be set, since some applications may not be usable when accuracy is below the threshold in some instances.”, Paragraph 35, “Implementing continuous retraining may involve making the following decisions: (1) in each retraining window, decide which edge models of a plurality of models to retrain; (2) allocate the edge server's GPU resources among the retraining and inference jobs, and (3) select the configurations of the retraining and inference jobs. Decisions may also be constrained such that the inference accuracy at any point in time does not drop below a minimum value (so that the outputs continue to remain useful to the application).” Ananthanarayanan allocates jobs across the GPUs in descending order such that the last allocated accelerator in that ordering is reached while training continues and the current accuracy is monitored every few epochs against a set accuracy threshold. Reaching the last accelerator in the ordering and continuing to train it while accuracy is under the threshold corresponds to continuing training on the final accelerator even when the measured accuracy is less than the threshold in subsequent epochs.) Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date to combine Ma in view of Wheatley's data model training and switching across heterogeneous accelerators with Ananthanarayanan's per-epoch accuracy driven scheduling that allocates accelerators in fractional ordered amounts and terminates models which do not improve a target metric in order to allocate computing resources efficiently and achieve higher accuracy for a given amount of accelerator resource (Paragraph 40 of Ananthanarayanan). Regarding claim 5, Ma teaches [o]ne or more non-transitory machine-readable information storage mediums comprising one or more instructions which when executed by one or more hardware processors cause (Page 13 Section 7.1, “We implement the heterogeneous CPU+GPU framework for deep learning in C/C++ using the pthreads library. The coordinator and each worker is managed by a stand-alone thread. The threads communicate using our custom asynchronous message queue.”): initiating a data model training using a first hardware accelerator from among a plurality of heterogeneous hardware accelerators (Page 7 Section 5, “CPUs and GPUs perform concurrent asynchronous SGD algorithms– specialized for their specific architecture– on data assigned dynamically and adaptively at runtime based on the current execution state.”, Page 7 Section 5.1 of Ma, “The architecture of the heterogeneous CPU+GPU framework for deep learning is depicted in Figure 3. It consists of a series of asynchronous worker threads corresponding to each of the CPUs and GPUs in Figure 2, and a central coordinator. The worker assigned to a hardware component is in charge of managing the resources, e.g., cores, memory, threads, and operation of that component.” Ma teaches training a deep learning model using a first hardware accelerator (CPU or GPU worker) from a set of heterogeneous accelerators.) wherein the data model training is spread across a plurality of epochs (Page 5 Section 3, “SGD can be stopped either after a fixed number of iterations, i.e., epochs, or when there is no significant drop in the loss across iterations. In practice, due to the large data set size and number of iterations it takes to converge, each SGD iteration is performed only over a randomly selected batch of B training examples…”, Page 10 Section 5.2, “As such, the framework performs loss computation after each complete pass– or a given number of batches– over the training data.” Ma teaches training and performing loss computation after each complete pass across multiple epochs.); and iteratively performing until one of a) a maximum accuracy is achieved for the data model, and b) a final hardware accelerator in a sequence of the plurality of heterogeneous hardware accelerators has reached (Page 5 Section 3, “SGD can be stopped either after a fixed number of iterations, i.e., epochs, or when there is no significant drop in the loss across iterations." Ma teaches that the model is iteratively performed until the loss is minimized, meaning that the accuracy is at its maximum level.): measuring accuracy of the data model after each of the plurality of epochs (Page 5 Section 3, “SGD can be stopped either after a fixed number of iterations, i.e., epochs, or when there is no significant drop in the loss across iterations.", Page 10 Section 5.2, “As such, the framework performs loss computation after each complete pass– or a given number of batches– over the training data. The loss is computed with a DNN forward pass over the training– or test– data… The size of the batch is proportional to the worker speed… Each worker computes a partial loss on its data batch and then sends it back to the coordinator, which aggregates it into the overall loss. This strategy is optimized for execution time by prioritizing the fast workers and minimizing the coordinator overhead.” The loss is measured after each epoch and the SGD is stopped once there is minimal change in the loss across epochs, meaning that the accuracy has reached its highest point.); determining if difference between the accuracy of the data model measured in two consecutive epochs of the plurality of epochs exceeds a threshold of accuracy (Page 5 Section 3, “SGD can be stopped either after a fixed number of iterations, i.e., epochs, or when there is no significant drop in the loss across iterations." Ma stops training when there is no significant drop in loss between successive epochs, and a drop in loss corresponds to a gain in accuracy. The significance criterion is the threshold such that a significant loss drop corresponds to an accuracy difference exceeding the threshold and an insignificant drop corresponds to an accuracy difference below it.); if the difference between the accuracy of the data model measured in two consecutive epochs of the plurality of epochs exceeds the threshold of accuracy (Page 5 Section 3, “SGD can be stopped either after a fixed number of iterations, i.e., epochs, or when there is no significant drop in the loss across iterations." Ma stops training when there is no significant drop in loss between successive epochs, and a drop in loss corresponds to a gain in accuracy. The significance criterion is the threshold such that a significant loss drop corresponds to an accuracy difference exceeding the threshold and an insignificant drop corresponds to an accuracy difference below it.) wherein when the measured accuracy of the data model remains constant over a pre-defined number of epochs or change in the measured accuracy of the data model is minimal over the pre-defined number of epochs, then the data model achieved the maximum accuracy… the model training is terminated (Page 5 Section 3, “SGD can be stopped either after a fixed number of iterations, i.e., epochs, or when there is no significant drop in the loss across iterations." Ma teaches stopping SGD once there is no significant drop in the loss across iterations, meaning that the accuracy changes minimally across epochs and that the model achieved a maximum accuracy.) Ma does not teach and switching the data model training to next hardware accelerator in the sequence of the plurality of heterogeneous hardware accelerators, if the difference between the accuracy of the data model measured in two consecutive epochs of the plurality of epochs is exceeding the threshold of accuracy. Wheatley, in the same field of endeavor, teaches and switching the data model training to a next hardware accelerator … (Paragraph 184, “As an example, a convolution neural network system (CNNS) can be implemented using one or more platforms… The GAFFE framework includes models and optimization defined by configuration where switching between CPU and GPU setting can be achieved via a single flag to train on a GPU machine then deploy to clusters.” Wheatley teaches switching a neural-network computation from a CPU to a GPU setting via a single configuration during training.) Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date to combine Ma’s teaching of learning data model training across heterogeneous CPU and GPU accelerators with Wheatley’s teaching of switching processors during training in a neural network in order to enhance training efficiency and resource utilization (Paragraph 156 of Wheatley). Although Ma in view of Wheatley substantially teaches the claimed invention, Ma in view of Wheatley is not relied on to teach and even when there are unutilized hardware accelerators in the sequence, the model training is terminated. However, Ananthanarayanan, in the same field of endeavor, teaches and even when there are unutilized hardware accelerators in the sequence, the model training is terminated (Paragraph 86, “the micro-profiler tests each configuration of a superset of configurations for a small number (e.g., 5) of training epochs, and then terminates the testing early, well before convergence… After early termination on the sampled training data, the (validation) accuracy of each configuration is obtained at each epoch it was trained.”, Paragraph 92, “To avoid expensive inter-GPU communication, the allocations may be quantized to inverse powers of two (e.g., ½, ¼, ⅛). This may make the jobs amenable to packing. The jobs may then be allocated to GPUs in descending order of demands to reduce fragmentation.”, Paragraph 38, “In each retraining window, the resource scheduler makes the decisions described above to (1) decide which of the edge models to retrain; (2) allocate the edge server's GPU resources among the retraining and inference jobs, and (3) select the configurations of the retraining and inference jobs… The scheduler decides against retraining the models which do not improve a target metric… First, it can simplify the spatial complexity by considering GPU allocations in coarse fractions (e.g., 10%) that are accurate enough for the scheduling decisions, while also being mindful of the granularity achievable in modern GPUs. Second, it can avoid changing allocations to jobs during the re-training, which helps to avoid temporal complexity.” Ananthanarayanan performs the deep-learning training on GPUs and FPGAs, which are the specialized hardware components that accelerate the model training, which corresponds to the hardware accelerators. Ananthanarayanan's scheduler allocates these GPU/FPGA resources to jobs in fractions and decides against continuing to retrain models that do not improve a target metric, terminating training early before convergence. Because the scheduler apportions the accelerators fractionally, a given accelerator may receive no tasks, i.e. remain unutilized where the termination of training can happen.) and wherein when the final hardware accelerator in the sequence of the plurality of heterogeneous hardware accelerators is reached, then the model training is continued with the hardware accelerator even when the measured accuracy is less than the threshold of accuracy in the subsequent epochs (Paragraph 92, “To avoid expensive inter-GPU communication, the allocations may be quantized to inverse powers of two (e.g., ½, ¼, ⅛). This may make the jobs amenable to packing. The jobs may then be allocated to GPUs in descending order of demands to reduce fragmentation.”, Paragraph 94, “Every few epochs (e.g., every 5 epochs), the current accuracy of the model being retrained is used to estimate its eventual accuracy when all the epochs are complete.”, Paragraph 38, “In each retraining window, the resource scheduler makes the decisions described above to (1) decide which of the edge models to retrain; (2) allocate the edge server's GPU resources among the retraining and inference jobs, and (3) select the configurations of the retraining and inference jobs… it can simplify the spatial complexity by considering GPU allocations in coarse fractions (e.g., 10%) that are accurate enough for the scheduling decisions…”, Paragraph 105, “The capacity of the thief scheduling approach (e.g., the maximum number of concurrent video streams subject to an accuracy threshold is compared with that of the uniform baseline, as more GPUs are available. An accuracy threshold may be set, since some applications may not be usable when accuracy is below the threshold in some instances.”, Paragraph 35, “Implementing continuous retraining may involve making the following decisions: (1) in each retraining window, decide which edge models of a plurality of models to retrain; (2) allocate the edge server's GPU resources among the retraining and inference jobs, and (3) select the configurations of the retraining and inference jobs. Decisions may also be constrained such that the inference accuracy at any point in time does not drop below a minimum value (so that the outputs continue to remain useful to the application).” Ananthanarayanan allocates jobs across the GPUs in descending order such that the last allocated accelerator in that ordering is reached while training continues and the current accuracy is monitored every few epochs against a set accuracy threshold. Reaching the last accelerator in the ordering and continuing to train it while accuracy is under the threshold corresponds to continuing training on the final accelerator even when the measured accuracy is less than the threshold in subsequent epochs.) Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date to combine Ma in view of Wheatley's data model training and switching across heterogeneous accelerators with Ananthanarayanan's per-epoch accuracy driven scheduling that allocates accelerators in fractional ordered amounts and terminates models which do not improve a target metric in order to allocate computing resources efficiently and achieve higher accuracy for a given amount of accelerator resource (Paragraph 40 of Ananthanarayanan). Claims 2, 4, 6 are rejected under 35 U.S.C. 103 as being unpatentable over Ma (“Heterogeneous CPU+GPU Stochastic Gradient Descent Algorithms”, 2020) in view of Wheatley (US 20240352839 A1), in view of Ananthanarayanan (US 20220188569 A1), and in further view of Kang (“Scheduling of Deep Learning Applications Onto Heterogeneous Processors in an Embedded Device”, 2020). Regarding claim 2, Ma does not teach the plurality of hardware accelerators in the sequence of the plurality of heterogeneous hardware accelerators are arranged in increasing order of specification. Kang, in the same field of endeavor, teaches the plurality of hardware accelerators in the sequence of the plurality of heterogeneous hardware accelerators are arranged in increasing order of specification (Page 4 Section 3, “Galaxy S9 is a heterogeneous system that consists of a Mali-G72 MP18 GPU and big.LITTLE CPUs with a quad-core M3 CPU running at 2.7GHz and a quad-core Cortex-A55 CPU at 1.79GHz.”, Page 7 Section 5, “Let L={L1,L2, ...,Ln} be a set of layers, or tasks, in a DL application sorted in the topology order, and PE={PE1, PE2,..., PEm} be a set of logical PEs in the device…” See Figure 1, PNG media_image1.png 161 420 media_image1.png Greyscale Kang teaches heterogeneous processors (CPU, GPU, NPU) with their specifications and the CPUS are arranged by increasing specification. P1, P2, P3, P4 represents a sequence of hardware accelerators. The scheduler determines which PE to proceed first which can start off with the worst accelerator to the best performing one according to the specification.). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date to combine Ma’s teaching of learning data model training across heterogeneous CPU and GPU accelerators with Kang’s teaching of switching neural network computation across heterogeneous processors in further view of Ananthanarayanan's per-epoch accuracy driven scheduling in order to enhance training efficiency and resource utilization (Introduction of Kang). Claim 4 recites similar limitations to claim 2. Therefore, claim 4 is rejected using the same rationale as claim 2. Claim 6 recites similar limitations to claim 2. Therefore, claim 6 is rejected using the same rationale as claim 2. Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to MAJD MAHER HADDAD whose telephone number is (571)272-2265. The examiner can normally be reached Mon-Friday 8-5 pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kamran Afshar, can be reached at (571) 272-7796. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /M.M.H./Examiner, Art Unit 2125 /KAMRAN AFSHAR/Supervisory Patent Examiner, Art Unit 2125
Read full office action

Prosecution Timeline

Jul 31, 2023
Application Filed
Apr 01, 2026
Non-Final Rejection mailed — §103, §112
Jun 30, 2026
Response Filed
Aug 20, 2026
Final Rejection mailed — §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12737630
FIRST NETWORK NODE AND METHOD PERFORMED THEREIN FOR HANDLING DATA IN A COMMUNICATION NETWORK
3y 11m to grant Granted Sep 15, 2026
Patent 12705535
Systems and Methods for Grouping Records Associated with Like Media Items
3y 6m to grant Granted Aug 11, 2026
Study what changed to get past this examiner. Based on 2 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
100%
Grant Probability
99%
With Interview (+0.0%)
3y 4m (~2m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 5 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month