Prosecution Insights
Last updated: August 17, 2026
Application No. 18/443,052

MODEL DISTILLATION METHOD AND RELATED DEVICE

Non-Final OA §101§103
Filed
Feb 15, 2024
Priority
Aug 20, 2021 — CN 202110962700.9 +1 more
Examiner
LE, HUNG VAN
Art Unit
Tech Center
Assignee
Huawei Technologies Co., Ltd.
OA Round
1 (Non-Final)
Grant Probability
Favorable
1-2
OA Rounds

Examiner Intelligence

Grants only 0% of cases
0%
Career Allowance Rate
0 granted / 0 resolved
-60.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
Avg Prosecution
19 currently pending
Career history
3
Total Applications
across all art units

Statute-Specific Performance

§101
35.9%
-4.1% vs TC avg
§103
64.1%
+24.1% vs TC avg
Black line = Tech Center average estimate • Based on career data from 0 resolved cases

Office Action

§101 §103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Information Disclosure Statement The information disclosure statement (IDS) submitted on 2025/05/22; 2025/01/29; 2024/11/21 & 2024/07/30. The submission is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1–20 are not rejected under 35 U.S.C. 101. The claims are directed to eligible subject matter for the reasons set forth below. Step 1 -- whether the claims fall within a statutory category. See MPEP 2106.03. Claims 1–8 are drawn to a method (process); claims 9–16 are drawn to a first model distillation apparatus (machine); and claims 17–20 are drawn to a non-transitory computer-readable storage medium (manufacture). Each claim therefore falls within one of the four statutory categories of invention. (Step 1: YES.) Step 2A Prong One -- whether the claims recite a judicial exception. See MPEP 2106.04, subsection II. Independent claims 1, 9, and 17 recite, in pertinent part, "the first intermediate output and the second intermediate output are used to determine a first gradient," and "distilling … the first sub-model based on the first gradient." Determining a gradient encompasses a mathematical concept, specifically a mathematical calculation. See MPEP § 2106.04(a)(2), subsection I. The claims are therefore identified as reciting a mathematical concept for purposes of Step 2A Prong One. (Step 2A Prong One: YES.) Step 2A Prong Two -- whether the claims as a whole integrate the recited judicial exception into a practical application. See MPEP 2106.04(d). Regarding independent claims 1, 9, and 17, the claims recite additional elements beyond the identified mathematical concept, including: a first computing node and a second computing node; a first sub-model and a second sub-model deployed on the first computing node and a third sub-model and a fourth sub-model deployed on the second computing node, the first and second sub-models being partial models of a student model and a teacher model respectively; obtaining, by the first computing node, the first input data and the second input data from the second computing node, the first input data being output data of the third sub-model and the second input data being output data processed by the fourth sub-model; processing the first and second input data by the respective sub-models to obtain a first and second intermediate output; and distilling the first sub-model based on the first gradient to obtain an updated first sub-model. These additional elements, considered as a whole, integrate the recited mathematical concept into a practical application because they reflect an improvement to the functioning of a computer and to the technical field of distributed neural-network training. See MPEP 2106.04(d)(1) and 2106.05(a). The specification sets forth the technical problem and the asserted improvement: in existing gradient back-propagation for knowledge distillation, "update is gradually performed from an output layer to an input layer, and update of a previous level of network layer depends on completion of update of a current level of network layer," such that "when update of a network layer of one or more current levels of computing nodes is not completed, a large quantity of computing nodes are in a resource idle state" (specification, ¶0024). The claimed architecture addresses this problem by partitioning the student and teacher models across computing nodes and providing that "a gradient calculated by each computing node is not back-propagated to a previous level of computing node," whereby each node performs its gradient back-propagation internally without dependency on a next-level node (specification, ¶0024), thereby achieving "higher utilization of computing resources" and "accelerat[ing] a distillation process" (specification, ¶0028). The claims reflect this improvement through the recited limitations of deploying the partitioned sub-models across the first and second computing nodes and distilling the first sub-model locally based on the first gradient. The improvement is therefore both described in the specification and reflected in the claims, consistent with the eligibility analysis of the 2024 Guidance Update on Patent Subject Matter Eligibility, Including on Artificial Intelligence, and the eligible claims of the accompanying examples. Because the claims as a whole integrate the recited judicial exception into a practical application (Step 2A Prong Two: YES), the claims are not directed to the judicial exception (Step 2A: NO), and the claims are eligible. The analysis need not proceed to Step 2B. Dependent claims 2–8, 10–16, and 18–20 incorporate the limitations of their respective independent claims and further narrow the architecture (for example, the transformer-model limitations of claims 2/10/18, the first-in-first-out queue limitations of claims 5/13, and the third-computing-node and queue limitations of claims 8/16). These claims are eligible at least for the same reasons as their respective independent claims. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The text of those sections of Title 35, U.S. Code not included in this action can be found in a prior Office action. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claims 1, 3, 9, 11, 17, and 19 are rejected under 35 U.S.C. 103 as being unpatentable over (Shao), Shao et al., U.S. Patent Application Publication No. US 2023/0092147 A1, published March 23, 2023 (effectively filed June 23, 2020 as PCT/US2020/039098, in view of (Amirguliyev), Amirguliyev et al., U.S. Patent Application Publication No. US 2021/0182660 A1, and further in view of (Blakeney), Blakeney et al., Non-Patent Literature, "Parallel Blockwise Knowledge Distillation for Deep Neural Network Compression," arXiv:2012.03096, published December 2020. Regarding claim 1, (Shao) teaches a model distillation method ((Shao), Title, "NEIGHBORHOOD DISTILLATION OF DEEP NEURAL NETWORKS"; ¶0019), comprising: the first sub-model is a partial model of a student model that comprises the third sub-model, the second sub-model is a partial model of a teacher model that comprises the fourth sub-model; processing, by the first computing node, the first input data by using the first sub-model, to obtain a first intermediate output; processing, by the first computing node, the second input data by using the second sub-model, to obtain a second intermediate output, wherein the first intermediate output and the second intermediate output are used to determine a first gradient; and distilling, by the first computing node, the first sub-model based on the first gradient, to obtain an updated first sub-model. (Shao) teaches the foregoing limitations as follows. With respect to "the first sub-model is a partial model of a student model that comprises the third sub-model, the second sub-model is a partial model of a teacher model that comprises the fourth sub-model," (Shao) teaches partitioning a teacher network into a plurality of non-overlapping pieces and training, for each such piece, a partial student model intended to replace that piece, such that the resulting student network is composed of plural partial student models ((Shao), ¶0021, "teacher network 202 is broken down into four teacher network neighborhoods 204, 206, 208, and 210… Neighborhoods will each include a separate, non-overlapping piece of the teacher network"; ¶0019, "breaking a teacher network down into smaller sub-networks, referred to herein as neighborhoods, which may each be distilled independently and then reassembled into a student network"; ¶0022, "each teacher network neighborhood 204, 206, 208, and 210 (or a copy thereof) is used by a separate trainer 212, 214, 216, and 218, respectively, to train a set of individual candidate student models"). Each teacher neighborhood reads on a partial model of a teacher model, and each candidate student model reads on a partial model of a student model. With respect to "processing, by the first computing node, the first input data by using the first sub-model, to obtain a first intermediate output; processing, by the first computing node, the second input data by using the second sub-model, to obtain a second intermediate output, wherein the first intermediate output and the second intermediate output are used to determine a first gradient," (Shao) teaches that the candidate student model and the teacher network neighborhood each generate an output from the input received, and that the two outputs are compared to compute an error that is used to generate a training gradient ((Shao), ¶0031, "The teacher network neighborhood 306 and the candidate student model 308 will then both generate outputs based on the inputs they received from the teacher root 304. The outputs from the teacher network neighborhood 306 and the candidate student model 308 are then compared to create a mean square error 310… [which] may be used by the trainer to recursively generate training gradients and tune the parameters of the candidate student model"; ¶0029, "minimizing the mean square error between its output and the output of the teacher network neighborhood"). The candidate student model's output reads on the first intermediate output, the teacher neighborhood's output reads on the second intermediate output, and the mean square error computed therefrom and used to generate training gradients reads on the first gradient. With respect to "distilling, by the first computing node, the first sub-model based on the first gradient, to obtain an updated first sub-model," (Shao) teaches using the training gradient to tune the parameters of the candidate student model until the loss between the student output and the teacher-neighborhood output is minimized, thereby producing the trained (updated) partial student model ((Shao), ¶0031, "used by the trainer to recursively generate training gradients and tune the parameters of the candidate student model until the objective loss between the output of the candidate student model 308 and the output of the teacher network neighborhood 306 has been minimized"; ¶0022, "tuning the parameters of each candidate student model (using training gradients and back-propagation) until the output of each candidate student model begins to approximate… the teacher network neighborhood"). (Shao) teaches subject matter related to distributing the distillation workload, in that the candidate student models for different neighborhoods may be trained in parallel and the processor(s), memory, and other elements of a computing device may be distributed between two or more housings ((Shao), ¶0019, "each candidate student model for a given neighborhood can be trained in parallel"; ¶0016, "The processor(s), memory, and other elements of a single computing device… may be distributed between two or more housings… references to a processor or computing device will be understood to include references to a collection of processors or computing devices… as well as one or more servers of a load-balanced server farm or cloud-based system"). However, (Shao) does not teach obtaining, by a first computing node, first input data and second input data from a second computing node, wherein a first sub-model and a second sub-model are deployed on the first computing node, a third sub-model and a fourth sub-model are deployed on the second computing node. In the same field of endeavor, (Amirguliyev) teaches obtaining, by a first computing node, first input data and second input data from a second computing node, wherein a first sub-model and a second sub-model are deployed on the first computing node, a third sub-model and a fourth sub-model are deployed on the second computing node ((Amirguliyev), ¶0050, "The system 100 includes a master device 110 and a slave device 120. The master device 110 and the slave device 120 include computing devices, i.e. devices with at least one processor and a memory"; ¶0052, "the master device 110 and the slave device 120 are communicatively coupled by at least one network 150"; ¶0006/Abstract, "the neural network model to be split between the master device and the slave device"; ¶0055, "the master device 110 generates first configuration data (CD₁) 180 that is sent to the slave device 120," and ¶0057, "The second configuration data 190 is transmitted from the slave device 120 to the master device 110 over the network 150… The master device 110 is configured to use the second configuration data 190 to update parameters for the first version of the neural network model 160"). (Amirguliyev) thereby teaches two distinct, communicatively-coupled computing nodes (a master device and a slave device), with a neural network model split and deployed across the two nodes, and with one node obtaining data (configuration/gradient data) from the other node over the network. The master device reads on the first computing node and the slave device reads on the second computing node, with the model split between them reading on the sub-models being deployed across the first and second computing nodes. (Shao) and (Amirguliyev) are analogous to the claimed invention as both are from the same field of endeavor of distillation and distributed training of neural-network models. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the neighborhood-distillation method of (Shao) with the master/slave distributed-computing-node deployment of (Amirguliyev) by deploying the partitioned teacher and student sub-models of (Shao) across the two communicatively-coupled computing nodes of (Amirguliyev), such that one node obtains the data needed for distillation from the other node. The motivation to combine (Shao) and (Amirguliyev) is as recited by (Amirguliyev) ((Amirguliyev), ¶0012, "the master device may incorporate information from a plurality of independent heterogenous devices that do not need to be part of a common system or under common control. This provides a flexible and scalable system for implementation in real world production environments"), such that one would distribute the partitioned distillation of (Shao) across separate computing nodes in order to provide a flexible and scalable training system, consistent with (Shao)'s own recognition that the per-neighborhood training may be performed in parallel across a collection of computing devices ((Shao), ¶0016, ¶0019). The combination of (Shao) and (Amirguliyev), however, does not teach wherein the first input data is output data of the third sub-model, and the second input data is output data processed by the fourth sub-model. In the same field of endeavor, (Blakeney) teaches wherein the first input data is output data of the third sub-model, and the second input data is output data processed by the fourth sub-model ((Blakeney), § 3.1, p. 4, "The activations from the proceeding preceding block are used as inputs for the student block as described in equation 3, and the activations from the original block being replaced become our labels"; Eq. (3), p. 4, "L_k^local = (1/N) Σ ‖f_k(x_i) − g_k ∘ f_{k−1}(x_i)‖²"). (Blakeney) teaches a blockwise distillation in which a model is represented as a composite of subnetwork blocks and the activations output by the preceding block, f_{k−1}(x_i), are used as the input to both the student replacement block g_k and the teacher block f_k. The preceding block that produces the activations supplied to the student block reads on the third sub-model, and the activations it outputs that are used as input read on the first input data being output data of the third sub-model; correspondingly, the activations supplied along the teacher path read on the second input data being output data processed by the fourth sub-model. (Shao), (Amirguliyev), and (Blakeney) are analogous to the claimed invention as all three are from the same field of endeavor of distillation and distributed/parallel training of neural-network models. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the partitioned, distributed distillation of (Shao) and (Amirguliyev) with the blockwise dataflow of (Blakeney), such that the input processed by the student and teacher sub-models on the first computing node is the output of the preceding sub-models on the second computing node. The motivation to combine (Shao), (Amirguliyev), and (Blakeney) is as recited by (Blakeney) ((Blakeney), § 3.1, p. 4, "all candidate blocks are trained independently and simultaneously. This not only eliminates most unwanted dependencies but also allows us to reduce the accumulated error from the model"; Abstract, "leverages local information to conduct independent blockwise distillation… [achieving] excellent speedup and scalability without compromising accuracy"), such that feeding each sub-model the output of the preceding sub-model and training each block on local information would eliminate inter-block dependencies and provide speedup and scalability when the partitioned distillation of (Shao) is distributed across the computing nodes of (Amirguliyev). Regarding claim 3, the combination of (Shao), (Amirguliyev), and (Blakeney) teaches the method according to claim 1, and claim 3 is rejected under the same rationale set forth for claim 1 above with respect to those limitations carried forward from claim 1. (Shao) further teaches "wherein an output layer in the student model is absent from the first sub-model." (Shao) teaches that each candidate student model replaces an intermediate, non-overlapping piece of the network and passes its output to a downstream teacher head that comprises the subsequent layers, such that the candidate student model itself does not include the output layer of the network ((Shao), ¶0032, "the teacher head 312a is a portion of the teacher network that immediately follows the teacher network neighborhood 306… [and] must include whatever layer, block, unit, element, node, etc. that receives output from the last layer, block, unit, element, node, etc. of the teacher network neighborhood 306"; ¶0033, "the candidate student model 308 will pass its output to teacher head 312b, which is an identical copy of teacher head 312a"; and ¶0050, "the neighborhood that comprises the final layer, block, unit, element, node, etc. of the teacher network" is treated as a separate neighborhood). The candidate student model, which replaces an intermediate neighborhood and passes its output to a downstream teacher head containing the subsequent layers, reads on the first sub-model, and the output layer being located in the downstream portion rather than in the candidate student model reads on an output layer in the student model being absent from the first sub-model. Regarding claim 9, claim 9 recites a first model distillation apparatus that performs operations corresponding to the method of claim 1, wherein the first model distillation apparatus corresponds to the first computing node of claim 1 and the second model distillation apparatus corresponds to the second computing node of claim 1. The limitations of claim 9 corresponding to the limitations of claim 1 are rejected under the same rationale set forth for claim 1 above. Claim 9 further recites the following additional limitations, which (Shao) teaches: "the first model distillation apparatus comprises a memory and at least one processor" ((Shao), FIG. 1, depicting processing system 102 having processor(s) 104 and memory 106; ¶0016, "The one or more processors included in each computing device may be any conventional processors, such as commercially available central processing units ('CPUs'), graphics processing units ('GPUs'), tensor processing units ('TPUs'), etc."). (Shao) teaches a processing system that includes one or more processors and a memory; the processing system having one or more processors and a memory reads on the first model distillation apparatus comprising a memory and at least one processor. "the memory stores programming instructions for execution by the at least one processor to perform operations comprising" ((Shao), FIG. 1, depicting memory 106 storing instructions 108; ¶0017, "The computing devices described herein may store instructions capable of being executed directly (such as machine code) or indirectly (such as scripts) by the processor(s) … Instructions may be stored as computing device code on a computing device-readable medium"). (Shao) teaches that the memory stores instructions that are executed by the processor(s) to perform the disclosed operations; the memory storing instructions executed by the processor to perform the operations reads on the memory storing programming instructions for execution by the at least one processor to perform operations. However, (Shao) does not expressly teach that the first model distillation apparatus is a discrete apparatus distinct from a second model distillation apparatus, each apparatus being a separate computing device having its own at least one processor and a memory. In the same field of endeavor, (Amirguliyev) teaches that the first and second apparatuses are separate, communicatively-coupled computing devices, each comprising at least one processor and a memory ((Amirguliyev), ¶0050, "The master device 110 and the slave device 120 include computing devices, i.e. devices with at least one processor and a memory wherein the at least one processor is configured to execute computer program code loaded into the memory to perform one or more functions"; ¶0052, "the master device 110 and the slave device 120 are communicatively coupled by at least one network 150"). (Amirguliyev) teaches a master device and a slave device, each being a discrete computing device having at least one processor and a memory storing executable program code, the two devices being communicatively coupled; the master device, being a discrete computing device having a processor and memory and communicatively coupled to the slave device, reads on the first model distillation apparatus being a discrete apparatus comprising a memory and at least one processor and distinct from the second model distillation apparatus. (Shao) and (Amirguliyev) are analogous to the claimed invention as both are from the same field of endeavor of distillation and distributed training of neural-network models. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the processing system of (Shao) with the discrete master/slave computing devices of (Amirguliyev), such that the operations of (Shao) are performed by a first model distillation apparatus having a memory and at least one processor that is a discrete apparatus distinct from a second model distillation apparatus. The motivation to combine (Shao) and (Amirguliyev) is as recited by (Amirguliyev) ((Amirguliyev), ¶0012, "the master device may incorporate information from a plurality of independent heterogenous devices that do not need to be part of a common system or under common control. This provides a flexible and scalable system for implementation in real world production environments"), such that one would implement the distillation operations of (Shao) on discrete, communicatively-coupled computing apparatuses in order to provide a flexible and scalable training system. Regarding claim 11, claim 11 recites, in apparatus form, the same additional limitation recited in claim 3. Claim 11 is rejected under the same rationale set forth for claims 3 and 9 above. Regarding claim 17, claim 17 recites a non-transitory computer-readable storage medium storing programming instructions for execution by at least one processor of a first model distillation apparatus to perform operations corresponding to the method of claim 1. The limitations of claim 17 corresponding to the limitations of claim 1 are rejected under the same rationale set forth for claim 1 above. Regarding claim 19, claim 19 recites, in non-transitory computer-readable storage medium form, the same additional limitation recited in claim 3. Claim 19 is rejected under the same rationale set forth for claims 3 and 17 above. Claims 2, 10, and 18 are rejected under 35 U.S.C. 103 as being unpatentable over (Shao), in view of (Amirguliyev), (Blakeney), and further in view of (Jiao), Jiao et al., Non-Patent Literature, "TinyBERT: Distilling BERT for Natural Language Understanding," arXiv:1909.10351, published 16 October 2020. Regarding claim 2, the combination of (Shao), (Amirguliyev), and (Blakeney) teaches the method according to claim 1, and claim 2 is rejected under the same rationale set forth for claim 1 above with respect to those limitations carried forward from claim 1. The combination of (Shao), (Amirguliyev), and (Blakeney) teaches subject matter related to the architecture of the student and teacher models, in that (Shao) teaches that the teacher network may be any kind of neural network having multiple layers, blocks, units, elements, or nodes ((Shao), ¶0021, "Teacher network 202 may be any kind of neural network having multiple layers, blocks, units, elements, nodes, etc."). However, the combination of (Shao), (Amirguliyev), and (Blakeney) does not teach the further limitations of claim 2. In the same field of endeavor, (Jiao) teaches these limitations as follows. With respect to "wherein the student model and the teacher model are transformer models," (Jiao) teaches that the teacher network is a transformer-based BERT model and the student network is a transformer-based TinyBERT model, each built from Transformer layers ((Jiao), page 3, right column, Section 3, "Problem Formulation," paragraph 1, "In this work, both the student and teacher networks are built with Transformer layers"; (Jiao), page 1, left column, Abstract, paragraph 1, "the plenty of knowledge encoded in a large 'teacher' BERT can be effectively transferred to a small 'student' TinyBERT"). The teacher BERT and the student TinyBERT, each constructed of Transformer layers, read on the student model and the teacher model being transformer models. With respect to "the first sub-model and the second sub-model each comprises one or more transformer layers," (Jiao) teaches that the student model has M Transformer layers and the teacher model has N Transformer layers, and that distillation is performed on a per-Transformer-layer basis by mapping a subset of the student's Transformer layers to corresponding teacher Transformer layers and matching their layer outputs ((Jiao), page 3, right column, Section 3, "Problem Formulation," paragraph 1, "Assuming that the student model has M Transformer layers and teacher model has N Transformer layers, we start with choosing M out of N layers from the teacher model for the Transformer-layer distillation"; (Jiao), page 4, left column, Section 3.1, "Transformer-layer Distillation," Equation (8), "L_hidn = MSE(H^S W_h, H^T)," where H^S and H^T refer to the hidden states of the student and teacher networks output by the Transformer layers being distilled). Each mapped portion of the student model, and each corresponding mapped portion of the teacher model, comprising one or more Transformer layers on which the per-layer distillation is performed, reads on the first sub-model and the second sub-model each comprising one or more transformer layers. (Shao), (Amirguliyev), (Blakeney), and (Jiao) are analogous to the claimed invention as all are from the same field of endeavor of distillation and distributed/parallel training of neural-network models. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the partitioned, distributed, blockwise distillation of (Shao), (Amirguliyev), and (Blakeney) with the transformer-model architecture and Transformer-layer distillation of (Jiao), such that the student model and the teacher model are transformer models and the first and second sub-models distilled at the first computing node each comprise one or more transformer layers. The motivation to combine (Shao), (Amirguliyev), (Blakeney), and (Jiao) is as recited by (Jiao) ((Jiao), page 1, left column, Abstract, paragraph 1, "pre-trained language models are usually computationally expensive, so it is difficult to efficiently execute them on resource-restricted devices. To accelerate inference and reduce model size while maintaining accuracy, we first propose a novel Transformer distillation method that is specially designed for knowledge distillation (KD) of the Transformer-based models"), such that one would apply the partitioned distillation of (Shao), (Amirguliyev), and (Blakeney) to transformer models composed of Transformer layers in order to accelerate inference and reduce model size while maintaining accuracy when compressing computationally expensive transformer-based models for resource-restricted devices. Regarding claim 10, claim 10 recites, in apparatus form, the same additional limitations recited in claim 2. Claim 10 is rejected under the same rationale set forth for claims 2 and 9 above. Regarding claim 18, claim 18 recites, in non-transitory computer-readable storage medium form, the same additional limitations recited in claim 2. Claim 18 is rejected under the same rationale set forth for claims 2 and 17 above. Claims 4-8, 12-16, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over (Shao), in view of (Amirguliyev), (Blakeney), and further in view of (Harlap), Harlap et al., Non-Patent Literature, "PipeDream: Generalized Pipeline Parallelism for DNN Training," arXiv:1806.03377, published 8 June 2018. Regarding claim 4, the combination of (Shao), (Amirguliyev), and (Blakeney) teaches the method according to claim 1, and claim 4 is rejected under the same rationale set forth for claim 1 above with respect to those limitations carried forward from claim 1, including the obtaining first input data from the second computing node. Claim 4 further recites that the obtaining first input data from the second computing node comprises the limitations addressed below. The combination of (Shao), (Amirguliyev), and (Blakeney) teaches subject matter related to the transfer of data between the second computing node and the first computing node, in that (Amirguliyev) teaches transmitting configuration data from the slave device to the master device over a network ((Amirguliyev), ¶0057, "The second configuration data 190 is transmitted from the slave device 120 to the master device 110 over the network 150"). However, the combination of (Shao), (Amirguliyev), and (Blakeney) does not expressly teach “obtaining, by the first computing node, the first input data from a first queue, wherein the first queue is used to store at least one piece of first data from the second computing node, each of the at least one piece of first data is output data obtained by the second computing node by based on processing input data by using the third sub-model, and the at least one piece of first data comprises the first input data”. In the same field of endeavor, (Harlap) teaches the further limitations of claim 4 as follows. With respect to "obtaining, by the first computing node, the first input data from a first queue," (Harlap) teaches that, upon receiving intermediate data from the prior pipeline stage, the runtime copies the intermediate data to memory and places a pointer to the associated buffer in a work queue, from which the stage retrieves the data for processing ((Harlap), page 7, right column, Section 4, "Implementation," "Intermediate State," "Upon receiving intermediate data from the prior stage (or from disk in the case of the input stage), PipeDream copies the intermediate data to GPU memory and places a pointer to the associated buffer in a work queue"). The work queue from which a stage retrieves the intermediate data received from the prior stage reads on the first queue from which the first input data is obtained. With respect to "wherein the first queue is used to store at least one piece of first data from the second computing node," (Harlap) teaches that the queued intermediate data received from the prior stage is stored and retained until the associated minibatch is processed ((Harlap), page 7, right column, Section 4, "Intermediate State," "Intermediate data from the forward pass is not discarded until the associated minibatch completes that stage's backward pass"; (Harlap), page 7, left column, Section 4, "Intermediate input activations and stashed weights for each active minibatch"). The intermediate data received from the prior stage and retained in the work queue reads on at least one piece of first data, from the second computing node, that the first queue is used to store. With respect to "each of the at least one piece of first data is output data obtained by the second computing node based on processing input data by using the third sub-model," the combination relies on the teachings already established for claim 1, in that (Blakeney) teaches that the activations output by the preceding block are used as inputs for the next block ((Blakeney), Section 3.1, page 4, "The activations from the proceeding [preceding] block are used as inputs for the student block"), and (Amirguliyev) teaches that the preceding (second/slave) node processes data using its deployed portion of the model ((Amirguliyev), ¶0056, "train the second version of the neural network model 170 using data from the first data source 130"). The activations output by the preceding block, produced by the second computing node using the third sub-model deployed thereon, read on each piece of first data being output data obtained by the second computing node based on processing input data by using the third sub-model. With respect to "and the at least one piece of first data comprises the first input data," (Harlap) teaches that the intermediate data placed in the work queue is the data subsequently retrieved and processed by the stage ((Harlap), page 7, right column, Section 4, "Intermediate State," as cited above), such that the stored intermediate data includes the particular piece of data obtained and processed as the first input data. The intermediate data stored in the work queue, which includes the piece retrieved and processed by the first computing node, reads on the at least one piece of first data comprising the first input data. (Shao), (Amirguliyev), (Blakeney), and (Harlap) are analogous to the claimed invention as all are from the same field of endeavor of distillation and distributed/parallel training of neural-network models. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the partitioned, distributed, blockwise distillation of (Shao), (Amirguliyev), and (Blakeney) with the inter-stage work queue of (Harlap), such that the first input data transferred from the second computing node is stored in, and obtained by the first computing node from, a first queue. The motivation to combine (Shao), (Amirguliyev), (Blakeney), and (Harlap) is as recited by (Harlap) ((Harlap), page 1, Abstract, "PipeDream… pipelines the execution of forward passes and intermludes them with backward passes in an attempt to minimize… time-to-target-accuracy"; and page 7, right column, Section 4, "PipeDream allocates all required GPU memory at the beginning of training, and reuses the allocated memory as appropriate. This significantly reduces the overhead of dynamically managing GPU memory"), such that one would buffer the data passed from the second computing node to the first computing node in a queue in order to decouple the producing and consuming nodes, allow the consuming node to retrieve its next input without waiting, and thereby reduce overhead and improve pipeline throughput in the distributed distillation of (Shao), (Amirguliyev), and (Blakeney). Regarding claim 5, the combination of (Shao), (Amirguliyev), (Blakeney), and (Harlap) teaches the method according to claim 4, and claim 5 is rejected under the same rationale set forth for claims 1 and 4 above with respect to those limitations carried forward from claims 1 and 4. The combination of (Shao), (Amirguliyev), and (Blakeney) teaches subject matter related to the first queue from which the first input data is obtained, in that (Amirguliyev) teaches transmitting the data from the slave device to the master device over a network ((Amirguliyev), ¶0057, "The second configuration data 190 is transmitted from the slave device 120 to the master device 110 over the network 150"). However, the combination of (Shao), (Amirguliyev), and (Blakeney) does not expressly teach wherein the first queue is a first-in-first-out queue. In the same field of endeavor, (Harlap) teaches wherein the first queue is a first-in-first-out queue. (Harlap) teaches that each stage processes minibatches in the order in which they enter the pipeline, admitting minibatches at the input stage and propagating them through to the output stage in arrival order under a one-forward-one-backward schedule, and that the intermediate data placed in the work queue is retained and consumed in that same order, not being discarded until the associated minibatch completes that stage's pass ((Harlap), page 5, right column, Section 3.2, "in steady state, each stage alternates between performing the forward and backward pass for a minibatch. We call this mechanism one-forward-one-backward (1F1B)"; (Harlap), page 6, left column, Section 3.2, "In the startup phase, the input stage admits exactly four minibatches that propagate their way to the output stage"; and (Harlap), page 7, right column, Section 4, "Intermediate State," "Intermediate data from the forward pass is not discarded until the associated minibatch completes that stage's backward pass"). The work queue, into which the intermediate data received from the prior stage is placed in arrival order and from which the data is consumed in that same arrival order, reads on the first queue being a first-in-first-out queue. (Shao), (Amirguliyev), (Blakeney), and (Harlap) are analogous to the claimed invention as all are from the same field of endeavor of distillation and distributed/parallel training of neural-network models. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the partitioned, distributed, blockwise distillation of (Shao), (Amirguliyev), and (Blakeney) with the first-in-first-out work queue of (Harlap), such that the first queue from which the first input data is obtained is a first-in-first-out queue. The motivation to combine (Shao), (Amirguliyev), (Blakeney), and (Harlap) is as recited by (Harlap) ((Harlap), page 5, right column, Section 3.2, "In a balanced pipeline, 1F1B ensures that no GPU is idle in steady state and that we make forward progress in learning from each minibatch"), such that one would process the data passed from the second computing node to the first computing node in first-in-first-out order in order to keep the pipeline full, avoid idle computing resources, and make forward progress in learning from each input in the distributed distillation of (Shao), (Amirguliyev), and (Blakeney). Regarding claim 6, the combination of (Shao), (Amirguliyev), (Blakeney), and (Harlap) teaches the method according to claim 4, and claim 6 is rejected under the same rationale set forth for claims 1 and 4 above with respect to those limitations carried forward from claims 1 and 4. (Shao) teaches: "after the distilling the first sub-model based on the first gradient, obtaining, by the first computing node, the third input data … in response to obtaining the updated first sub-model" ((Shao), ¶0050, "the processing system 102 will additionally modify one or more parameters of each of student models 205a-z based on the … training gradient … and then may repeat steps 406-412 of FIG. 4 and steps 504-510 of FIG. 5 using a new input and these modified parameters for student models 205a-z"). (Shao) teaches that, after the parameters of the candidate student model have been modified based on the training gradient to obtain the updated candidate student model, a new input is obtained for processing by that updated candidate student model; obtaining a new input after, and in response to, the modification of the candidate student model's parameters reads on obtaining, after the distilling the first sub-model based on the first gradient and in response to obtaining the updated first sub-model, the third input data. "wherein the third input data is used as input data in a feedforward process existing when model distillation is performed on the updated first sub-model" ((Shao), ¶0051, "steps 406-408 of FIG. 4 and steps 504-510 of FIG. 5 may then be repeated as already described so as to recursively feed new data to the given neighborhood and each candidate student model, and thus tune the parameters of each candidate student model … until the output of each candidate student model begins to approximate (or match) the output of the given neighborhood for each input"; ¶0022, "a recursive process of feeding a set of data to the teacher network neighborhood and each candidate student model"). (Shao) teaches that the new input is fed forward through the candidate student model having the modified parameters as part of the recursive distillation process; feeding the new input forward through the candidate student model having the modified parameters reads on the third input data being used as input data in a feedforward process existing when model distillation is performed on the updated first sub-model. However, (Shao) does not expressly teach that the third input data is output data of the third sub-model. In the same field of endeavor, (Blakeney) teaches "the third input data is output data of the third sub-model" ((Blakeney), Section 3.1, page 4, "The activations from the proceeding [preceding] block are used as inputs for the student block"). (Blakeney) teaches that the activations output by the preceding block are used as the input for the next block; the activations output by the preceding (third) sub-model, used as a subsequent input, read on the third input data being output data of the third sub-model. (Shao) and (Blakeney) are analogous to the claimed invention as both are from the same field of endeavor of distillation of neural-network models. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the recursive feedforward distillation of (Shao) with the blockwise dataflow of (Blakeney), such that the subsequent input fed forward through the updated first sub-model is the output data of the third sub-model. The motivation to combine (Shao) and (Blakeney) is as recited by (Blakeney) ((Blakeney), Section 3.1, page 4, "all candidate blocks are trained independently and simultaneously. This not only eliminates most unwanted dependencies but also allows us to reduce the accumulated error from the model"), such that using the output of the preceding sub-model as the input would eliminate unwanted dependencies and reduce accumulated error in the distillation of (Shao). The combination of (Shao) and (Blakeney) does not expressly teach that the at least one piece of first data comprises the third input data, and that the third input data is obtained from the first queue. In the same field of endeavor, (Harlap) teaches: "the at least one piece of first data comprises third input data" ((Harlap), page 7, left column, Section 4, "Intermediate input activations and stashed weights for each active minibatch"; page 7, left column, Section 4, "the input stage needs to do so for NOAM minibatches"). (Harlap) teaches that the queue holds plural pieces of intermediate input data for multiple in-flight minibatches; the queue holding a further piece of intermediate input data in addition to the first input data reads on the at least one piece of first data comprising third input data. "obtaining, by the first computing node, the third input data from the first queue" ((Harlap), page 7, right column, Section 4, "Once the minibatch is processed, the ML worker indicates the completion of work to PipeDream, and pulls its next work item"; page 7, right column, Section 4, "Intermediate State," "Upon receiving intermediate data from the prior stage … PipeDream copies the intermediate data to GPU memory and places a pointer to the associated buffer in a work queue"). (Harlap) teaches that, upon completing the processing of a minibatch, the worker pulls its next work item -- a further piece of intermediate data held in the work queue -- for processing; pulling the next queued piece of intermediate data upon completion of the current work reads on obtaining, by the first computing node, the third input data from the first queue. (Shao), (Blakeney), and (Harlap) are analogous to the claimed invention as all are from the same field of endeavor of distillation and distributed/parallel training of neural-network models. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the recursive feedforward distillation of (Shao) and the blockwise dataflow of (Blakeney) with the queue-based retrieval of a next input of (Harlap), such that, after the first sub-model is distilled and the updated first sub-model is obtained, the first computing node obtains the third input data from the first queue and feeds it forward through the updated first sub-model. The motivation to combine (Shao), (Blakeney), and (Harlap) is as recited by (Harlap) ((Harlap), page 7, right column, Section 4, "Once the minibatch is processed, the ML worker indicates the completion of work to PipeDream, and pulls its next work item"; and page 5, right column, Section 3.2, "1F1B ensures that no GPU is idle in steady state and that we make forward progress in learning from each minibatch"), such that the first computing node would retrieve its next input from the queue immediately upon completing distillation of the current sub-model, thereby keeping the computing resources fully utilized and making continuous forward progress in the distillation of (Shao) and (Blakeney). Regarding claim 7, the combination of (Shao), (Amirguliyev), (Blakeney), and (Harlap) teaches the method according to claim 1, and claim 7 is rejected under the same rationale set forth for claim 1 above with respect to those limitations carried forward from claim 1. (Shao) teaches: "each piece of second data is an output obtained … based on processing input data by using the fourth sub-model" ((Shao), ¶0030, "The trainer will then conduct training by passing an input 302 into the teacher root 304, which will pass output to both the teacher network neighborhood 306 and the candidate student model 308"). (Shao) teaches that the teacher portion preceding the neighborhood processes the input and produces an output that is passed onward for use in the distillation; the output produced by the teacher portion from the input it processes reads on each piece of second data being an output obtained based on processing input data by using the fourth sub-model. However, (Shao) does not expressly teach that the processing is performed by the second computing node. In the same field of endeavor, (Amirguliyev) teaches "each piece of second data is an output obtained by the second computing node" ((Amirguliyev), ¶0056, "Once instantiated, the slave device 120 is configured to train the second version of the neural network model 170 using data from the first data source 130. The training includes retrieving a training set from the first data source 130"; ¶0052, "the master device 110 and the slave device 120 are communicatively coupled by at least one network 150"). (Amirguliyev) teaches that the second (slave) computing node processes data using the portion of the model deployed thereon and communicates the result to the first (master) computing node; the processing performed by the slave device using its deployed model portion reads on the output being obtained by the second computing node. (Shao) and (Amirguliyev) are analogous to the claimed invention as both are from the same field of endeavor of distillation and distributed training of neural-network models. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the teacher-path processing of (Shao) with the master/slave distributed-computing-node deployment of (Amirguliyev), such that the output processed by the fourth sub-model is obtained by the second computing node. The motivation to combine (Shao) and (Amirguliyev) is as recited by (Amirguliyev) ((Amirguliyev), ¶0012, "the master device may incorporate information from a plurality of independent heterogenous devices that do not need to be part of a common system or under common control. This provides a flexible and scalable system for implementation in real world production environments"), such that one would distribute the teacher-path processing of (Shao) onto a separate computing node in order to provide a flexible and scalable training system. The combination of (Shao) and (Amirguliyev) does not expressly teach that the output processed by the fourth sub-model is used as second data in a distillation in which the activations of the teacher-path block serve as the distillation target. In the same field of endeavor, (Blakeney) teaches: "each piece of second data is an output obtained … based on processing input data by using the fourth sub-model" in the context of the distillation target ((Blakeney), Section 3.1, page 4, "the activations from the original block being replaced become our labels"). (Blakeney) teaches that the activations output by the block being replaced are used as the labels supplied to the distillation; the activations output by the teacher-path block, supplied as the distillation target, read on each piece of second data being the output obtained based on processing input data by using the fourth sub-model. (Shao), (Amirguliyev), and (Blakeney) are analogous to the claimed invention as all are from the same field of endeavor of distillation and distributed/parallel training of neural-network models. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the distributed teacher-path processing of (Shao) and (Amirguliyev) with the blockwise distillation target of (Blakeney), such that the output processed by the fourth sub-model and obtained by the second computing node is used as the second data serving as the distillation target. The motivation to combine (Shao), (Amirguliyev), and (Blakeney) is as recited by (Blakeney) ((Blakeney), Section 3.1, page 4, "all candidate blocks are trained independently and simultaneously. This not only eliminates most unwanted dependencies but also allows us to reduce the accumulated error from the model"), such that using the teacher-path block activations as the distillation target would eliminate unwanted dependencies and reduce accumulated error in the distributed distillation of (Shao) and (Amirguliyev). The combination of (Shao), (Amirguliyev), and (Blakeney) does not expressly teach a second queue. In the same field of endeavor, (Harlap) teaches: "obtaining, by the first computing node, the second input data from a second queue for storing at least one piece of second data from the second computing node" ((Harlap), page 7, right column, Section 4, "Intermediate State," "Upon receiving intermediate data from the prior stage … PipeDream copies the intermediate data to GPU memory and places a pointer to the associated buffer in a work queue"; page 7, left column, Section 4, "Intermediate input activations and stashed weights for each active minibatch"). (Harlap) teaches that intermediate data received from the prior stage is stored in, and retrieved by the stage from, a work queue; the work queue storing and supplying the intermediate data received from the prior stage reads on the second queue from which the second input data is obtained and which stores at least one piece of second data from the second computing node. "and the at least one piece of second data comprises the second input data" ((Harlap), page 7, right column, Section 4, "Once the minibatch is processed, the ML worker indicates the completion of work to PipeDream, and pulls its next work item"). (Harlap) teaches that the intermediate data placed in the queue is the data subsequently retrieved and processed by the stage; the stored intermediate data, which includes the particular piece retrieved and processed, reads on the at least one piece of second data comprising the second input data. (Shao), (Amirguliyev), (Blakeney), and (Harlap) are analogous to the claimed invention as all are from the same field of endeavor of distillation and distributed/parallel training of neural-network models. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the partitioned, distributed, blockwise distillation of (Shao), (Amirguliyev), and (Blakeney) with the inter-stage work queue of (Harlap), such that the second input data transferred from the second computing node is stored in, and obtained by the first computing node from, a second queue. The motivation to combine (Shao), (Amirguliyev), (Blakeney), and (Harlap) is as recited by (Harlap) ((Harlap), page 7, right column, Section 4, "PipeDream allocates all required GPU memory at the beginning of training, and reuses the allocated memory as appropriate. This significantly reduces the overhead of dynamically managing GPU memory"; and page 5, right column, Section 3.2, "1F1B ensures that no GPU is idle in steady state and that we make forward progress in learning from each minibatch"), such that one would buffer the data passed from the second computing node to the first computing node in a queue in order to decouple the producing and consuming nodes, allow the consuming node to retrieve its next input without waiting, and thereby reduce overhead and improve pipeline throughput in the distributed distillation of (Shao), (Amirguliyev), and (Blakeney). Regarding claim 8, the combination of (Shao), (Amirguliyev), (Blakeney), and (Harlap) teaches the method according to claim 1, and claim 8 is rejected under the same rationale set forth for claim 1 above with respect to those limitations carried forward from claim 1. (Shao) teaches: "the student model further comprises a fifth sub-model connected after the first sub-model" ((Shao), ¶0021, "teacher network 202 is broken down into four teacher network neighborhoods 204, 206, 208, and 210"; ¶0045, "the processing system 102 will combine each selected model corresponding to each neighborhood of the plurality of neighborhoods to form the second neural network … to form the final student network 220"). (Shao) teaches that the student network is formed from a plurality of partial student models corresponding to successive neighborhoods, such that a further partial student model is connected after the first partial student model; the further partial student model connected after the first reads on the student model further comprising a fifth sub-model connected after the first sub-model. "the first intermediate output is used as input in a feedforward process when model distillation is performed on the third sub-model" ((Shao), ¶0030, "The teacher root 304 is a portion of the teacher network that immediately precedes the teacher network neighborhood 306 … [and] will pass output to both the teacher network neighborhood 306 and the candidate student model 308"; ¶0050, "recursively feed new data to the given neighborhood and each candidate student model"). (Shao) teaches that the output of a preceding sub-model is fed forward as the input for the distillation of the succeeding sub-model; the output of the preceding sub-model fed forward into the distillation of a succeeding sub-model reads on the first intermediate output being used as input in a feedforward process when model distillation is performed on the third sub-model. However, (Shao) does not expressly teach that the first computing node is communicatively connected to a third computing node and that the third computing node obtains the first intermediate output. In the same field of endeavor, (Amirguliyev) teaches: "the first computing node is communicatively connected to a third computing node" ((Amirguliyev), ¶0012, "the master device may incorporate information from a plurality of independent heterogenous devices that do not need to be part of a common system or under common control"; ¶0052, "the master device 110 and the slave device 120 are communicatively coupled by at least one network 150"). (Amirguliyev) teaches that the system may include a plurality of communicatively-coupled computing devices beyond the first and second; a further computing device communicatively coupled within the system reads on the first computing node being communicatively connected to a third computing node. "the third computing node obtains the first intermediate output" ((Amirguliyev), ¶0057, "The master device 110 is configured to use the second configuration data 190 to update parameters"; ¶0012, "the master device may incorporate information from a plurality of independent heterogenous devices"). (Amirguliyev) teaches that a computing node obtains data output by another communicatively-coupled node for use in its own processing; a further (third) computing node obtaining the output produced upstream reads on the third computing node obtaining the first intermediate output. (Shao) and (Amirguliyev) are analogous to the claimed invention as both are from the same field of endeavor of distillation and distributed training of neural-network models. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the successive-neighborhood feedforward distillation of (Shao) with the plural communicatively-coupled computing nodes of (Amirguliyev), such that the first computing node is communicatively connected to a third computing node and the third computing node obtains the first intermediate output. The motivation to combine (Shao) and (Amirguliyev) is as recited by (Amirguliyev) ((Amirguliyev), ¶0012, "This provides a flexible and scalable system for implementation in real world production environments"), such that one would distribute the successive-neighborhood distillation of (Shao) across a plurality of computing nodes in order to provide a flexible and scalable training system. The combination of (Shao) and (Amirguliyev) does not expressly teach that the first intermediate output transferred to the third computing node is the output of the preceding sub-model serving as the distillation input. In the same field of endeavor, (Blakeney) teaches: "the first intermediate output is used as input in a feedforward process when model distillation is performed on the third sub-model" ((Blakeney), Section 3.1, page 4, "The activations from the proceeding [preceding] block are used as inputs for the student block as described in equation 3"). (Blakeney) teaches that the activations output by the preceding block are used as the input for the distillation of the succeeding block; the activations of the preceding sub-model used as the distillation input for the succeeding sub-model read on the first intermediate output being used as input in a feedforward process when model distillation is performed on the third sub-model. (Shao), (Amirguliyev), and (Blakeney) are analogous to the claimed invention as all are from the same field of endeavor of distillation and distributed/parallel training of neural-network models. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the distributed successive-neighborhood distillation of (Shao) and (Amirguliyev) with the blockwise feedforward dataflow of (Blakeney), such that the first intermediate output obtained by the third computing node is used as the input for the distillation performed on the succeeding sub-model. The motivation to combine (Shao), (Amirguliyev), and (Blakeney) is as recited by (Blakeney) ((Blakeney), Section 3.1, page 4, "all candidate blocks are trained independently and simultaneously. This not only eliminates most unwanted dependencies but also allows us to reduce the accumulated error from the model"), such that using the preceding sub-model's output as the distillation input would eliminate unwanted dependencies and reduce accumulated error in the distributed distillation of (Shao) and (Amirguliyev). The combination of (Shao), (Amirguliyev), and (Blakeney) does not expressly teach a third queue. In the same field of endeavor, (Harlap) teaches: "the first intermediate output is transferred to a third queue for storing the first intermediate output" ((Harlap), page 7, right column, Section 4, "Intermediate State," "Upon receiving intermediate data from the prior stage … PipeDream copies the intermediate data to GPU memory and places a pointer to the associated buffer in a work queue"; page 7, right column, Section 4, "Intermediate State," "Intermediate data … is released as soon as the ML worker finishes using it, and if necessary, after it is sent to the next stage"). (Harlap) teaches that the intermediate output produced at one stage is sent to the next stage and placed in a work queue for storage; the intermediate output being sent to and stored in a work queue at the downstream stage reads on the first intermediate output being transferred to a third queue for storing the first intermediate output. "the third computing node obtains the first intermediate output from the third queue" ((Harlap), page 7, right column, Section 4, "Once the minibatch is processed, the ML worker indicates the completion of work to PipeDream, and pulls its next work item"). (Harlap) teaches that the downstream stage retrieves the intermediate data from its work queue for processing; the downstream (third) computing node pulling the stored intermediate output from the work queue reads on the third computing node obtaining the first intermediate output from the third queue. (Shao), (Amirguliyev), (Blakeney), and (Harlap) are analogous to the claimed invention as all are from the same field of endeavor of distillation and distributed/parallel training of neural-network models. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the distributed blockwise distillation of (Shao), (Amirguliyev), and (Blakeney) with the inter-stage work queue of (Harlap), such that the first intermediate output is transferred to a third queue and the third computing node obtains the first intermediate output from the third queue. The motivation to combine (Shao), (Amirguliyev), (Blakeney), and (Harlap) is as recited by (Harlap) ((Harlap), page 7, right column, Section 4, "PipeDream allocates all required GPU memory at the beginning of training, and reuses the allocated memory as appropriate. This significantly reduces the overhead of dynamically managing GPU memory"; and page 5, right column, Section 3.2, "1F1B ensures that no GPU is idle in steady state and that we make forward progress in learning from each minibatch"), such that one would buffer the first intermediate output passed to the third computing node in a queue in order to decouple the producing and consuming nodes, allow the third computing node to retrieve its input without waiting, and thereby reduce overhead and improve pipeline throughput in the distributed distillation of (Shao), (Amirguliyev), and (Blakeney). Regarding claim 12, claim 12 recites, in apparatus form, the same additional limitations recited in claim 4. Claim 12 is rejected under the same rationale set forth for claims 4 and 9 above. Regarding claim 13, claim 13 recites, in apparatus form, the same additional limitation recited in claim 5. Claim 13 is rejected under the same rationale set forth for claims 5 and 9 above. Regarding claim 14, claim 14 recites, in apparatus form, the same additional limitations recited in claim 6. Claim 14 is rejected under the same rationale set forth for claims 6 and 9 above. Regarding claim 15, claim 15 recites, in apparatus form, the same additional limitations recited in claim 7. Claim 15 is rejected under the same rationale set forth for claims 7 and 9 above. Regarding claim 16, claim 16 recites, in apparatus form, the same additional limitations recited in claim 8. Claim 16 is rejected under the same rationale set forth for claims 8 and 9 above. Regarding claim 20, claim 20 recites, in non-transitory computer-readable storage medium form, the same additional limitations recited in claim 4. Claim 20 is rejected under the same rationale set forth for claims 4 and 17 above. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to HUNG VAN LE whose telephone number is (571)270-0164. The examiner can normally be reached 8 a.m. - 5 p.m.. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Cesar Paula can be reached at (571) 272-4128. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /HUNG VAN LE/Examiner, Art Unit 2145 /CESAR B PAULA/Supervisory Patent Examiner, Art Unit 2145
Read full office action

Prosecution Timeline

Feb 15, 2024
Application Filed
Mar 27, 2024
Response after Non-Final Action
Jul 17, 2026
Non-Final Rejection mailed — §101, §103 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
Grant Probability
Low
PTA Risk
Based on 0 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month