DETAILED ACTION
This office action is in response to the Application No. 16418799 filed on
09/15/2025. Claims 1-25 are presented for examination and are currently pending. Applicant’s arguments have been carefully and respectfully considered.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Reopening of Prosecution After Appeal Brief
3. In view of the appeal brief filed on 03/26/2026, PROSECUTION IS HEREBY REOPENED. A new ground of rejection is set forth below.
To avoid abandonment of the application, appellant must exercise one of the following two options:
(1) file a reply under 37 CFR 1.111 (if this Office action is non-final) or a reply under 37 CFR 1.113 (if this Office action is final); or,
(2) initiate a new appeal by filing a notice of appeal under 37 CFR 41.31 followed by an appeal brief under 37 CFR 41.37. The previously paid notice of appeal fee and appeal brief fee can be applied to the new appeal. If, however, the appeal fees set forth in 37 CFR 41.20 have been increased since they were previously paid, then appellant must pay the difference between the increased fees and the amount previously paid.
A Supervisory Patent Examiner (SPE) has approved of reopening prosecution by signing below:
Response to Arguments
4. Upon further review, the final rejection has been withdrawn. As a result,
a new non-final rejection has been issued because the Applicant’s argument on page 3 of the appeal that “Kulkarni does not describe comparing training performance values during training, switching between data parallelism and model parallelism, or switching "based on comparing training performance values achieved by the one or more iterations of the machine learning model parallelism technique with training performance values achieved by the one or more iterations of the data parallelism technique," as recited in claim 1” is persuasive.
In addition, a new reference has been applied to the independent claims 1, 7, 14 and 20.
The Examiner notes that since a new non-final rejection has been issued, i.e., new grounds of rejection have been presented. Claims 2-6, 8-13, 15-19 and 21-25 are not patentable in light of the new grounds of rejection.
On page 5 of the appeal, the Applicant argued that dependent claim 12 was not rejected.
The Examiner notes, while there was a typo in the header and 12 was not included in the header, claim 12 was rejected on the bottom of page 11 to the beginning of page 12.
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
5. Claims 1, 5-9, 14 and 19-22 are rejected under 35 U.S.C 102(a)(1) as being anticipated by over Lewis et al. (US20180314935)
Regarding claim 1, Lewis teaches one or more processors, comprising: circuitry (Any processes relating to framework 860 may be performed by processing logic that may comprise hardware (e.g., circuitry, dedicated logic, programmable logic, etc.) [0181]) to train a machine learning model (FIG. 7 illustrates training mechanism 610 of FIG. 6 according to one embodiment. ... training mechanism 610 may include ... adaptive runtime logic 709 [0153]; FIG. 8B illustrates a novel framework 840 for adaptive runtime for efficient deep learning according to one embodiment [0179]) by, during training (the adaptive runtime component is further to use the input data sets along with a compiled code for data model and hybrid parallelism and a novel profiling table to communicate with and select one or more of data parallelism, model parallelism, and hybrid parallelism for efficiently training one or more neural networks [0185]),
iteratively switching between applying one or more iterations of a machine learning model parallelism technique and one or more iterations of a data parallelism technique (The process begins by adaptive runtime layer logic 709 executing a network model on a subset of input data using data parallelism (such as dividing the mini-batch size across a number of cores, nodes, etc.). It then executes the same model on another subset of input data using model parallelism (such as dividing the amount of parallelism in each layer into multiple cores, nodes, etc.) [0164]; The Examiner notes the adaptive runtime layer logic 709 executing a network model on a subset of input data using data parallelism is done at one iteration and then executes the same model on another subset of input data using model parallelism which is the next iteration. The changing from data parallelism to model parallelism is switching) based on comparing training performance values achieved by the one or more iterations of the machine learning model parallelism technique with training performance values achieved by the one or more iterations of the data parallelism technique (adaptive runtime layer logic 709 computes a relative throughput on each model on the underlying hardware. If one model seems clearly better than the other, then adaptive runtime layer logic 709 may then use the best model for the remaining input data [0165]; In one embodiment, adaptive runtime layer logic 709 may be used to seamlessly determine the best parallelism for a given DNN application execution on a hardware platform. [0164]; The Examiner notes adaptive runtime layer logic computes a throughput of the model using data parallelism and the model using model parallelism and based on which performance model is better the adaptive layer logic switches to that model (i.e. switches to data parallelism or model parallelism). The throughput is a training performance value that is compared and each model switches based on the comparison) until a level of training efficiency is achieved (At block 911, through efficient training, a neural network, such as a DNN architecture, is trained and maintained [0185]).
Regarding claim 5, Lewis teaches the one or more processors of claim 1, Lewis teaches wherein the machine learning model is trained by using a combination of the one or more iterations of the data parallelism technique and the one or more iterations of the machine learning model parallelism technique within one or more layers of the machine learning model (In one embodiment, adaptive runtime layer logic 709 may be used to seamlessly determine the best parallelism for a given DNN application execution on a hardware platform. The process begins by adaptive runtime layer logic 709 executing a network model on a subset of input data using data parallelism (such as dividing the mini-batch size across a number of cores, nodes, etc.). It then executes the same model on another subset of input data using model parallelism (such as dividing the amount of parallelism in each layer into multiple cores, nodes, etc. [0164]) to achieve the level of training efficiency (Further, in one embodiment, adaptive runtime component 823 may communicate with one or more of data parallelism 851, model parallelism 853, and hybrid parallelism 855 to provide an efficient neural network, such as efficient DNN architecture 857 [0180]) .
Regarding claim 6, Lewis teaches the one or more processors of claim 1, Lewis teaches wherein the circuitry is to use the machine learning model to infer information (FIG. 17 illustrates an exemplary inferencing system on a chip (SOC) suitable for performing inferencing using a trained model [0025]),
wherein the machine learning model is trained by using the one or more iterations of the data parallelism technique followed by the one or more iterations of the machine learning model parallelism technique (The process begins by adaptive runtime layer logic 709 executing a network model on a subset of input data using data parallelism (such as dividing the mini-batch size across a number of cores, nodes, etc.). It then executes the same model on another subset of input data using model parallelism (such as dividing the amount of parallelism in each layer into multiple cores, nodes, etc.) [0164]; The Examiner notes the adaptive runtime layer logic 709 executing a network model on a subset of input data using data parallelism is done at one iteration and then executes the same model on another subset of input data using model parallelism which is the next iteration. The change from data parallelism to model parallelism is switching).
Regarding claim 7, claim 7 is similar to claim 1. It is rejected in the same manner and reasoning applying. Further Lewis teaches a system, comprising: one or more computers having one or more processors (FIG. 18 is a block diagram of an embodiment of a computer system with a processor having one or more processor cores and graphics processors [0026])
Regarding claim 8, Lewis teaches the system of claim 7, Lewis teaches wherein a number of training data threads is split into subsets to be distributed among the one or more processors to train the machine learning model (data parallelism is often used to efficiently map mini-batch sizes to the underlying hardware architecture that is also data parallel. In the context of DNN, data parallelism refers to using the same network model for every thread and/or process, but feeding it with different parts of the input data [0163]; The process begins by adaptive runtime layer logic 709 executing a network model on a subset of input data using data parallelism (such as dividing the mini-batch size across a number of cores, nodes, etc.). It then executes the same model on another subset of input data using model parallelism (such as dividing the amount of parallelism in each layer into multiple cores, nodes, etc.[0164]).
Regarding claim 9, Lewis teaches the system of claim 7, Lewis teaches wherein a number of portions of the machine learning model trained in parallel is split into components to be distributed among the one or more processors to train the machine learning model (There also exists another variant, called model parallelism, which refers to using the same data for every thread, but splitting the network model among threads [0163]; It then executes the same model on another subset of input data using model parallelism (such as dividing the amount of parallelism in each layer into multiple cores, nodes, etc.)[0164])
Regarding claim 14, claim 14 similar to claim 1. It is rejected in the same manner and reasoning applying. Further, Lewis teaches a non-transitory machine-readable medium having stored thereon a set of instructions, which if performed by one or more processors, cause the one or more processors to at least (one non-transitory or tangible machine-readable medium comprising a plurality of instructions, when executed on a computing device, to implement or perform a method or realize an apparatus as claimed [0370]):
Regarding claim 19, Lewis teaches the machine-readable medium of claim 14, Lewis teaches wherein the set of instructions further cause the one or more processors to at least train the neural network using a number of parallel training data threads prior to using a number of portions of the neural network in parallel using the number of parallel training data threads. (The process begins by adaptive runtime layer logic 709 executing a network model on a subset of input data using data parallelism (such as dividing the mini-batch size across a number of cores, nodes, etc.). It then executes the same model on another subset of input data using model parallelism (such as dividing the amount of parallelism in each layer into multiple cores, nodes, etc.) [0164]. The Examiner notes that data parallelism occurs because model parallelism)
Regarding claim 20, claim 20 similar to claim 1. It is rejected in the same manner and reasoning applying.
Regarding claim 21, claim 21 similar to claim 8. It is rejected in the same manner and reasoning applying.
Regarding claim 22, claim 22 similar to claim 9. It is rejected in the same manner and reasoning applying.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
6. Claims 3 and 12 are rejected under 35 U.S.C. 103 as being unpatentable over Lewis et al. (US20180314935) further in view of Koch et al. (US20180240041)
Regarding claim 3, Lewis teaches the one or more processors of claim 1, Lewis does not explicitly teach wherein the training performance values are measured based at least in part on training times associated with using the one or more iterations of the data parallelism technique and the level of training efficiency is measured based at least in part on training times associated with using the combination of the one or more iterations of the data parallelism technique and the one or more iterations of the machine learning model parallelism technique.
Koch teaches wherein the training performance values are one or more performance metrics is measured based at least in part on training times associated with using the one or more iterations of the data parallelism technique (Thus, the training and scoring process is repeated F−1 times with different training datasets [0213]; Each session of the plurality of sessions executes training and scoring of the model type using the input dataset in parallel with other sessions of the plurality of sessions [0005]; computation times used to perform the hyperparameter tuning are computed for example, using the training time, … included in the data structure associated with each session manager device 400 that contains times for the model train and score executions for that session [0193]) and the level of training efficiency is measured based at least in part on training times associated with using the combination of the one or more iterations of the data parallelism technique and the one or more iterations of the machine learning model parallelism technique (model train/score worker application 432 provide efficient distributed and parallel computing device implementations for training and tuning models. The presented results demonstrate the improved model accuracies and the improved execution times [0262]. The Examiner notes the improved model accuracies and the improved execution times means improved training times).
It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of Lewis to incorporate the method of Koch for the benefit of providing efficient distributed and parallel computing device implementations for training and tuning models which demonstrates the improved model accuracies and the improved execution times (Koch [0262]).
Regarding claim 12, Lewis teaches the system of claim 7, Lewis does not explicitly teach wherein a number of portions of the machine learning model are to be trained in parallel using a number of training data threads.
Koch teaches wherein a number of portions of the machine learning model are to be trained in parallel using a number of training data threads (Each session of a plurality of sessions executes training … of a model type using an input dataset in parallel, abstract; when a train or score action is executed, each computing device of that session also may use multiple threads [0071])
It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of Lewis to incorporate the method of Koch for the benefit of providing efficient distributed and parallel computing device implementations for training and tuning models which demonstrates the improved model accuracies and the improved execution times (Koch [0262])
7. Claims 4, 15, 16 and 25 are rejected under 35 U.S.C. 103 as being unpatentable over Lewis et al. (US20180314935) in view of Harlap et al. ("Pipedream: Fast and efficient pipeline parallel dnn training." arXiv preprint arXiv:1806.03377 (2018))
Regarding claim 4, Lewis teaches the one or more processors of claim 1, Lewis does not explicitly teach wherein the machine learning model is trained by: comparing training times associated with using the one or more iterations of the data parallelism technique and training times associated with using a combination of the one or more iterations of the data parallelism technique and the one or more iterations of the machine learning model parallelism technique; and using the combination of the one or more iterations of the data parallelism technique and the one or more iterations of the machine learning model parallelism technique based on the comparing.
Harlap teaches wherein the machine learning model is trained by: comparing training times associated with using the one or more iterations of the data parallelism technique and training times associated with using a combination of the one or more iterations of the data parallelism technique and the one or more iterations of the machine learning model parallelism technique (Table 1 summarizes results comparing PipeDream with data-parallel training (BSP). For the three models, the table shows PipeDream’s auto-generated configuration and the corresponding speedup in training time over single machine and data-parallel training (BSP). The Examiner notes this represents training times associated with machine learning model parallelism techniques; It also shows the communication reduction achieved by PipeDream compared to data-parallel training, pg. 10, right col, last para., 5.2 PipeDream vs. Data Parallelism. The Examiner notes this represents training times associated with data parallelism techniques); and
using the combination of the one or more iterations of the data parallelism technique and the one or more iterations of the machine learning model parallelism technique based on the comparing (Figure 11 shows accuracy vs. training time for VGG16 and Inception-v3, using 8 machines in Cluster-B, for both BSP and PipeDream. Compared to Cluster-A, ClusterB employs faster V100 GPUs with 10Gbps interconnect between the machines (as granted by the cloud provider). Thus, models running on Cluster-B have lower computation-to-communication ratios. We note that the faster GPUs result in faster end-to-end training time — e.g., training time for VGG16 reduces from 220 hours on Cluster-A to little less than 100 hours on Cluster-B (Figures 10 and 11). We also observe that the higher communication overhead causes both BSP and PipeDream to scale less effectively to 8 machines.), pg. 11, right col, second to the last para.).
It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of Lewis to incorporate the method of Harlap for the benefit of using PipeDream which reduces communication by up to 95% for large DNNs relative to data-parallel training, and allows perfect overlap of communication and computation (Harlap, abstract).
Regarding claim 15, Lewis teaches the non-transitory machine-readable medium of claim 14, Lewis does not explicitly teach wherein a first level of training efficiency is based at least in part on training times associated with training the machine learning model using the number of parallel training data threads.
Harlap teaches wherein a first level of training efficiency is based at least in part on training times associated with training the machine learning model using the number of parallel training data threads (achieving a first level of training efficiency of about 3 speedup over 1 machine (y-axis) with 4 PipeDream (x-axis) pg. 12, Fig. 13; … PipeDream eliminates 95% of this communication overhead thereby improving performance by 7.04x, pg. 11, left col, last para.; a machine may have multiple GPUs each running a worker thread, Footnote 1, pg. 3. The Examiner notes that since Harlap's system trains DNNs (abstract) the worker threads can be considered training threads).
It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of Lewis to incorporate the method of Harlap for the benefit of using PipeDream which reduces communication by up to 95% for large DNNs relative to data-parallel training, and allows perfect overlap of communication and computation (Harlap, abstract).
Regarding claim 16, Lewis teaches the non-transitory machine-readable medium of claim 14, Lewis does not explicitly teach wherein a second level of training efficiency indicates training times associated with training the machine learning model using a number of portions of the machine learning model in parallel in combination with using a number of parallel training data threads.
Harlap teaches wherein a second level of training efficiency indicates training times associated with training the machine learning model using a number of portions of the machine learning model in parallel in combination with using a number of parallel training data threads (achieving a second level of efficiency of about 7 speedup with 8 PipeDream (x-axis), pg. 12, Fig. 13; … PipeDream eliminates 95% of this communication overhead thereby improving performance by 7.04x, pg. 11, left col, last para. pg. 10, right col, Models and training Methodology; a machine may have multiple GPUs each running a worker thread, Footnote 1, pg. 3. Using these estimates, PipeDream’s optimizer partitions layers across available machines, Fig. 7).
It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of Lewis to incorporate the method of Harlap for the benefit of using PipeDream which reduces communication by up to 95% for large DNNs relative to data-parallel training, and allows perfect overlap of communication and computation (Harlap, abstract).
Regarding claim 25, Lewis teaches the method of claim 20, Lewis does not explicitly teach wherein a number of portions of the machine learning model that are to be trained in parallel using a number of parallel training data threads is based at least in part on attaining speedup associated with using the number of portions to train the machine learning model.
Harlap teaches wherein a number of portions of the machine learning model that are to be trained in parallel (Using these estimates, PipeDream’s optimizer partitions layers across available machines, Fig. 7. The Examiner notes that the DNN has four layers or components in Fig. 7 which implies the layers are partitioned and distributed across different machines that are trained in parallel)
using a number of parallel training data threads is based at least in part on attaining speedup associated with using the number of portions to train the machine learning model (PipeDream provides the ML worker thread (Caffe) pointers to GPU memory containing layer input data, pg. 9, Fig. 9; a machine may have multiple GPUs each running a worker thread, Footnote 1, pg. 3. The Examiner notes that the worker threads can be considered as a training thread. We use the ILSVRC12 dataset to train VGG16 and Inception-v3, pg. 10, right col, Models and Training Methodology Section.).
It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of Lewis to incorporate the method of Harlap for the benefit of using PipeDream which reduces communication by up to 95% for large DNNs relative to data-parallel training, and allows perfect overlap of communication and computation (Harlap, abstract).
8. Claims 10, 17 and 23 are rejected under 35 U.S.C. 103 as being unpatentable over Lewis et al. (US20180314935) in view of Ioannou et al. (US20200356815 filed 05/07/2019)
Regarding claim 10, Lewis teaches the system of claim 7, Lewis does not explicitly teach wherein the one or more computers having the one or more processors further train the machine learning model by increasing a number of training data threads until the level of training efficiency is achieved.
Ioannou teaches wherein the one or more computers having the one or more processors further train the machine learning model by increasing a number of training data threads until the level of training efficiency is achieved (we use a prior method in a multi-threaded setting to train a logistic regression model … Results depicted in FIG. 2 show that as we increase the number of worker threads the number of epochs to converge increases significantly, achieving a mere 2.7× speedup with 32 cores [0044]. The Examiner notes 2.7× speedup as the target level of training efficiency achieved).
It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of Lewis to incorporate the method of Ioannou for the benefit of increasing data parallelism of the algorithm to improve scalability (Ioannou [0045]).
Regarding claim 17, Lewis teaches the non-transitory machine-readable medium of claim 14, Lewis does not explicitly teach wherein the set of instructions further cause the one or more processors to at least train the machine learning model by adjusting the number of parallel training data threads until the level of training efficiency is achieved.
Ioannou teaches wherein the set of instructions further cause the one or more processors to at least train the machine learning model by adjusting the number of parallel training data threads until the level of training efficiency is achieved (we use a prior method in a multi-threaded setting to train a logistic regression model … Results depicted in FIG. 2 show that as we increase the number of worker threads the number of epochs to converge increases significantly, achieving a mere 2.7× speedup with 32 cores [0044]. The Examiner notes 2.7× speedup as the target level of training efficiency achieved).
It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of Lewis to incorporate the method of Ioannou for the benefit of increasing data parallelism of the algorithm to improve scalability (Ioannou [0045]).
Regarding claim 23, Lewis teaches the method of claim 20, Lewis does not explicitly teach wherein a number of parallel training data threads to train the machine learning model is increased until the machine learning model is trained at or above a first level of training efficiency, wherein the first level of training efficiency is based at least in part on training speedup associated with using the number of parallel training data threads to train the machine learning model.
Ioannou teaches wherein a number of parallel training data threads to train the machine learning model is increased until the machine learning model is trained at or above a first level of training efficiency (we use a prior method in a multi-threaded setting to train a logistic regression model … Results depicted in FIG. 2 show that as we increase the number of worker threads the number of epochs to converge increases significantly, achieving a mere 2.7× speedup with 32 cores [0044]. The Examiner notes 2.7× speedup as the target level of training efficiency achieved),
wherein the first level of training efficiency is based at least in part on training speedup associated with using the number of parallel training data threads to train the machine learning model (Results depicted in FIG. 2 show that as we increase the number of worker threads the number of epochs to converge increases significantly, achieving a mere 2.7× speedup [0044]).
It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of Lewis to incorporate the method of Ioannou for the benefit of increasing data parallelism of the algorithm to improve scalability (Ioannou [0045]).
9. Claims 2 and 24 are rejected under 35 U.S.C. 103 as being unpatentable over Lewis et al. (US20180314935) in view of Ioannou et al. (US20200356815 filed 05/07/2019) further in view of Harlap et al. ("Pipedream: Fast and efficient pipeline parallel dnn training." arXiv preprint arXiv:1806.03377 (2018))
Regarding claim 2, Lewis teaches the one or more processors of claim 1, Lewis does not explicitly teach wherein the machine learning model is trained by increasing one or more iterations of the data parallelism technique of data until the level of training efficiency is achieved and increasing the one or more iterations of the machine learning model parallelism technique until the level of training efficiency is achieved.
Ioannou teaches wherein the machine learning model is trained by increasing one or more iterations of the data parallelism technique of data until the level of training efficiency is achieved (we use a prior method in a multi-threaded setting to train a logistic regression model … Results depicted in FIG. 2 show that as we increase the number of worker threads the number of epochs to converge increases significantly, achieving a mere 2.7× speedup with 32 cores [0044]. The Examiner notes 2.7× speedup as the target level of training efficiency achieved) and
It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of Lewis to incorporate the method of Ioannou for the benefit of increasing data parallelism of the algorithm to improve scalability (Ioannou [0045]).
Modified Lewis does not explicitly teach increasing the one or more iterations of the machine learning model parallelism technique until the level of training efficiency is achieved.
Harlap teaches increasing the one or more iterations of the machine learning model parallelism technique until the level of training efficiency is achieved (Using these estimates, PipeDream’s optimizer partitions layers across available machines, Fig. 7.; The figure 13 shows these results for training VGG16 using 4 and 8 machines on Cluster-A, pg. 12, right col, section 5.3; PipeDream speedup over 1 machine of 7.04x, pg. 11, Table 1, first row .The Examiner notes that 4 machines implies partitioning neural network into 4 layers and 8 machines implies partitioning neural network into 8 layers which means that the number of portions are increased from 4 layers to 8 layers and 7.04x is the level of training efficiency.)
It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of Lewis in view of Ioannou to incorporate the teachings of Harlap for the benefit of using PipeDream which reduces communication by up to 95% for large DNNs relative to data-parallel training, and allows perfect overlap of communication and computation (Harlap, abstract).
Regarding claim 24, Lewis in view of Ioannou teaches the method of claim 23, Harlap teaches wherein a number of portions of the machine learning model to train in parallel using the number of parallel training data threads is determined in response to the machine learning model being trained at or above the first level of training efficiency (Given a model and a set of machines, PipeDream’s first challenge is to automatically partition layers of the model across available machines so as to minimize overall training time. Figure 7 shows the workflow adopted by PipeDream to partition the layers of the DNN among the available machines (pg. 5, right col., section 3.2 to pg. 6 left col., first para.); The partitioning algorithm tries to minimize the overall training time of the model, pg. 6, right col., section PipeDream’s Partitioning Algorithm. The Examiner notes that the partitioning algorithm determines the number of portion of the machine learning model by minimizing training time which is the training efficiency.).
It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of Lewis in view of Ioannou to incorporate the method of Harlap for the benefit of using PipeDream which reduces communication by up to 95% for large DNNs relative to data-parallel training, and allows perfect overlap of communication and computation (Harlap, abstract).
10. Claims 11 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Lewis et al. (US20180314935) in view of Seide et al. ("1-bit stochastic gradient descent and its application to data-parallel distributed training of speech dnns.", hereinafter “seide2” Fifteenth annual conference of the international speech communication association. 2014.)
Regarding claim 11, Lewis teaches the system of claim 7, Lewis teaches wherein the one or more computers having the one or more processors further train the machine learning model by (FIG. 6 illustrates a computing device 600 hosting an efficient training mechanism (“training mechanism”) 610 according to one embodiment [0140]):
Lewis does not explicitly teach comparing training times associated with using a number of training data threads and training times associated with using a number of portions of the machine learning model trained in parallel using the number of training data threads; and using the number of portions of the machine learning model trained in parallel using the number of training data threads based on the comparison.
Seide2 teaches comparing training times associated with using a number of training data threads and training times associated with using a number of portions of the machine learning model trained in parallel (Table 5: combining data and model parallelism, pg. 1061, left-right col, 5.5 Combination with Model Parallelism.; Row “partial gradients” adds data parallelism over K = 4 compute nodes, combined with 2-GPU model parallelism in each compute node, pg. 1061, left col, first para., Table 4.The examiner notes that the table 5 shows the comparison of the training speed of data parallelism and model parallelism)
using the number of training data threads; and using the number of portions of the machine learning model trained in parallel using the number of training data threads based on the comparing (Table 3 analyzes where to apply AdaGrad in the process. First, we can see that applying AdaGrad to the raw gradients before momentum, rather than the momentum-smoothed gradient, leads to higher training frame accuracy and 0.3 points better WER. We believe that this is because momentum smoothing reduces the standard deviation and thus the effect of AdaGrad. Row “partial gradients” adds data parallelism over K = 4 compute nodes, combined with 2-GPU model parallelism in each compute node (more on that in Section 5.5). AdaGrad is applied locally before quantization. while it leads to a small WER gain, the training frame accuracy drops a little, pg. 1061, left col, first para.; Comparing the same number of GPUs, MP only helps in one configuration, the communication-bound minibatch size 2880 with 16 GPUs, pg. 1061, left-right col, 5.4. Impact of MB-Size Selection and Double Buffering).
It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of Lewis to incorporate the teachings of Seide2 for the benefit of reducing end-to-end training time from 35h to 8.1h (Seide2, pg. 1061, fourth Para., Table 3)
Regarding claim 18, Lewis teaches the non-transitory machine-readable medium of claim 14, Lewis teaches wherein the set of instructions further cause the one or more processors to at least train the machine learning model by (FIG. 6 illustrates a computing device 600 hosting an efficient training mechanism (“training mechanism”) 610 according to one embodiment [0140]):
Lewis does not explicitly teach comparing training times associated with using a number of parallel training data threads and training times associated with using a number of portions of the machine learning model in parallel using the number of parallel training data threads; and using the number of portions of the machine learning model in parallel using the number of parallel training data threads based on the comparing.
Seide2 teaches comparing training times associated with using a number of parallel training data threads and training times associated with using a number of portions of the machine learning model in parallel (Table 5: combining data and model parallelism, pg. 1061, left-right col, 5.5 Combination with Model Parallelism.; Row “partial gradients” adds data parallelism over K = 4 compute nodes, combined with 2-GPU model parallelism in each compute node, pg. 1061, left col, first para., Table 4.The examiner notes that the table 5 shows the comparison of the training speed of data parallelism and model parallelism)
using the number of parallel training data threads; and using the number of portions of the machine learning model in parallel using the number of parallel training data threads based on the comparing (Table 3 analyzes where to apply AdaGrad in the process. First, we can see that applying AdaGrad to the raw gradients before momentum, rather than the momentum-smoothed gradient, leads to higher training frame accuracy and 0.3 points better WER. We believe that this is because momentum smoothing reduces the standard deviation and thus the effect of AdaGrad. Row “partial gradients” adds data parallelism over K = 4 compute nodes, combined with 2-GPU model parallelism in each compute node (more on that in Section 5.5). AdaGrad is applied locally before quantization. while it leads to a small WER gain, the training frame accuracy drops a little, pg. 1061, left col, first para.; Comparing the same number of GPUs, MP only helps in one configuration, the communication-bound minibatch size 2880 with 16 GPUs, pg. 1061, left-right col, 5.4. Impact of MB-Size Selection and Double Buffering).
It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of Lewis to incorporate the method of Seide2 for the benefit of reducing end-to-end training time from 35h to 8.1h (Seide2, pg. 1061, fourth Para., Table 3)
11. Claim 13 is rejected under 35 U.S.C. 103 as being unpatentable over Lewis et al. (US20180314935) in view of Li et al. ("Strategies for energy-efficient resource management of hybrid programming models." IEEE Transactions on parallel and distributed Systems 24.1 (2012): 144-157)
Regarding claim 13, Lewis teaches the system of claim 7, Lewis does not explicitly teach wherein the level training efficiency is based at least in part on information directed to power consumption of the system.
Li teaches wherein the level training efficiency is based at least in part on information directed to power consumption of the system (Dynamic Concurrently Throttling (DCT) can reduce dynamic power consumption by putting cores that do not improve application execution time in a low-power state. DCT can also save execution time by alleviating contention for shared resources, pg. 145, right col, Applying DCT).
It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of Lewis to incorporate the teachings of Li for the benefit of achieving substantial energy savings (8.74 percent on average and up to 13.8 percent) with some performance gain (up to 7.5 percent) or negligible performance loss (Li, abstract)
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to MORIAM MOSUNMOLA GODO whose telephone number is (571)272-8670. The examiner can normally be reached Monday-Friday 8:00am-5:00pm EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Michelle T. Bechtold can be reached on (571) 431-0762. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/M.G./Examiner, Art Unit 2148
/MICHELLE T BECHTOLD/Supervisory Patent Examiner, Art Unit 2148