DETAILED ACTION
This communication is in response to the Application No. 18/356,360 filed on June 24, 2026
in which Claims 1 - 11 are presented for examination.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Amendment
The amendments filed on June 24, 2026 have been considered. Claims 1, 4, 8, 9 and 11 have been amended. Thus, Claims 1 - 11 are pending and presented for examination.
Applicant's arguments filed June 24, 2026 with respect to the rejection of claims 1 – 6 and 11 under 35 U.S.C. 112(b) have been fully considered and are persuasive. The rejection of claims 1 – 6 and 11 under 35 U.S.C. 112(b) is hereby withdrawn in view of Applicant's amendment.
Applicant's arguments filed June 24, 2026 with respect to the rejection of claims 1-11 under 35 U.S.C. 101 have been fully considered but they are not persuasive.
Applicant’s argument on pg. 6 of Arguments/Remarks state:
PNG
media_image1.png
873
874
media_image1.png
Greyscale
Examiner respectfully disagrees. As to Prong One, the claims recite performing a forward propagation, a backward propagation, and a weight update. These are mathematical calculations comprising the weighted sums, gradient computations, and parameter adjustments by which a neural network is trained, and thus recite a mathematical concept. Reciting these operations as a "workflow" to "manage continual-learning weight updates" does not change their character as mathematical calculations.
As to Prong Two, the claims do not recite any specific accelerator design or any technical way that the improved speed or reduced forgetting is actually achieved. The accelerator is described only by what it is supposed to do, such as being "configured to improve parallel processing efficiency" and "to mitigate catastrophic forgetting while accelerating training speed," and not by any structure that does it. Claiming a desired result without saying how that result is reached does not integrate the exception into a practical application. See MPEP 2106.05(f). Applicant's argument that the method changes how the processor works with memory is not supported by the claim, which only recites a generic processor, memory, and accelerator. The claimed faster training and lower overhead do not come from any improvement to the hardware, but from the abstract idea itself, such as reducing weight updates based on an activation tendency measure and scheduling the forward and backward operations to run early through speculation. Making generic hardware perform an abstract process faster is not an improvement to the computer itself. For these reasons, Applicant's argument is not persuasive.
Applicant’s argument on pg. 7 of Arguments/Remarks state:
PNG
media_image2.png
358
970
media_image2.png
Greyscale
Examiner respectfully disagrees. The features Applicant points to are not additional elements that could supply an inventive concept. Checking speculation against a threshold, storing an activation baseline, and running the forward and backward operations asynchronously are all part of the abstract idea itself, such as the mathematical scheme of speculative, activation weighted training, and features that make up the abstract idea cannot also serve as the "significantly more" needed to make the claim eligible. See MPEP 2106.05(f) and 2106.04(d). The only additional elements in the claim are the generic processor, memory, and hardware accelerator, which perform nothing more than their ordinary functions of storing data and executing the recited operations, and reciting generic computer components performing their ordinary functions does not amount to significantly more than the abstract idea. See MPEP 2106.05(f). Considered alone or in combination, these elements do not add beyond the abstract idea and its generic hardware environment, and Applicant's argument is therefore not persuasive.
Thus, the rejection under 35 U.S.C. 101 is maintained.
Applicant's arguments filed June 24, 2026 with respect to the rejections under 35 U.S.C. 103 have been fully considered but are moot in view of the new grounds of rejection set forth below, which were necessitated in part by Applicant's amendment. The previously cited reference Park et al. ("Continual Learning with Speculative Backpropagation and Activation History," 2022) is no longer relied upon. Accordingly, Applicant's arguments directed to disqualification of that reference under 35 U.S.C. 102(b)(1)(A) are moot.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claim 1-11 are rejected under 35 U.S.C. 101 because these claimed inventions are directed to an abstract idea without significantly more.
Regarding Claim 1:
Step 1: Claim 1 is a method type claim. Therefore, Claims 1-6 fall within one of the four statutory
categories (i.e., process, machine, manufacture, or composition of matter).
2A Prong 1: If a claim limitation, under its broadest reasonable interpretation, covers performance
of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. If a claim limitation, under its broadest reasonable
interpretation, covers performance of the limitation by mathematical calculation but for the recitation
of generic computer components, then it falls within the “Mathematical Concepts” grouping of abstract
ideas.
a forward propagation operation of performing a forward propagation (mathematical concept - mathematical calculations used to compute outputs of a neural network from input data – as disclosed in the specification (Paragraph [0022], “In the forward propagation, data is propagated from an input layer to an output layer. Each neuron computes a weighted sum of inputs from connected neurons in its prior layer and then adds a value calculated as shown in Equation 1 with a bias”))
a backward propagation operation of performing a backward propagation (mathematical concept - backward propagation in a neural network involves mathematical calculations such as computing error derivatives, gradients, and propagating those values through layers to adjust model parameters – as disclosed in the specification (Paragraph [0024], “Backward propagation (backpropagation) is used to adjust weights (wij j ) by calculating derivates. The backpropagation starts from the output layer based on, for example, Softmax. The derivative in the output layer is expressed in Equation 4.”))
and wherein the weight update operation of performing a weight update, wherein the forward propagation operation, the backward propagation operation, and the weight update operation are repeatedly performed, and wherein the weight update operation is performed based on an activation tendency […] (mathematical concept - recites mathematical calculations used to iteratively adjust neural network parameters based on activation history values and gradient based optimization– as disclosed in the specification (Paragraph [0025], “In DNN training, weights are adjusted based on errors computed in the backpropagation. Initially, as in Equation 7, Δwij j is calculated by multiplying backpropagation outcome δi l by forward propagation outcome yj j-1 Then, weights (wij j ) are updated according to Equation 8. Here, a learning rate η determines a degree of learning. This process is repeated for all the weights.”))
Step 2A Prong 2 & Step 2B: This judicial exception is not integrated into a practical application.
[…] a processor and a memory, wherein the deep learning model is executed on an artificial intelligence (Al) hardware accelerator integrated […] recited at a high-level of generality (i.e., a generic processor, memory, a communication interface, a user interface and accelerator) such that it amounts to no more than mere instructions to apply the exception using generic computer components)
[…] of each of neurons included in the deep learning model, in a process in which training for the first task proceeds (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of training a deep learning model without significantly more)
For the reasons above, Claim 1 is rejected as being directed to an abstract idea without significantly more. This rejection applies equally to dependent claims 1 - 6. The additional limitations of the dependent claims are addressed below.
Regarding Claim 2:
Step 2A Prong 1: See the rejection of Claim 1 above, which Claim 2 depends on.
wherein the forward propagation operation, the backward propagation operation, and the weight update operation are repeatedly performed until […] converge to a predetermined range (mathematical concept - describes an iterative mathematical optimization process that repeatedly performs calculations until model parameters satisfy a convergence)
Step 2A Prong 2 & Step 2B:
[….] weights of the deep learning model […] (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of using weights of the deep learning model without significantly more)
Accordingly, under Step 2A Prong 2 and Step 2B, this additional element does not integrate the
abstract idea into practical application because it does not impose any meaningful limits on practicing
the abstract idea, as discussed above in the rejection of claim 1. The claim does not include additional
elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception.
Regarding Claim 3:
Step 2A Prong 1: See the rejection of Claim 1 above, which Claim 3 depends on.
wherein the weight update operation limits weight update for a neuron of which activation tendency has a value greater than a predetermined value in the process in which training for the first task proceeds (mathematical concept - involves comparing a numerical activation tendency value with a predetermined threshold and adjusting the parameter update calculation accordingly; “activation tendency” is interpreted as a numerical value representing the historical activation behavior of a neuron within a predetermined range, as described in the specification - paragraph [0037])
Step 2A Prong 2 & Step 2B:
Accordingly, under Step 2A Prong 2 and Step 2B, there are no additional elements that integrate the
abstract idea into practical application. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial
exception.
Regarding Claim 4:
Step 2A Prong 1: See the rejection of Claim 1 above, which Claim 4 depends on.
obtaining a gradient (mathematical concept – requires performing calculations to determine the gradient of an objective function with respect to model weights, e.g., Algorithm 3 step 2: “g ← ∇₍w₎ f(w) (get gradients with objective function)”, which represents a mathematical optimization calculation)
for a neuron of which activation tendency has a value greater than a predetermined value in the process in which training for the first task proceeds, reducing an updated amount of weight by multiplying the obtained gradient by a predetermined constant (r) (mathematical concept - requires performing mathematical calculations including comparing a numerical activation tendency with a threshold and reducing the gradient value by multiplying it with a constant (r) to adjust the parameter update)
and updating the weight using the reduced updated amount (mental process - updating the weight using the reduced weight may be performed manually by a user by observing the reduced value and applying the corresponding arithmetic update to the weight)
Step 2A Prong 2 & Step 2B:
Accordingly, under Step 2A Prong 2 and Step 2B, there are no additional elements that integrate the
abstract idea into practical application. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial
exception.
Regarding Claim 5:
Step 2A Prong 1: See the rejection of Claim 1 above, which Claim 5 depends on.
Step 2A Prong 2 & Step 2B:
wherein, in a training process for the second task, the forward propagation operation and the backward propagation operation that are repeatedly performed proceed in parallel at least once (Field of Use – limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that the forward propagation operation and the backward propagation operation that are repeatedly performed proceed in parallel does not integrate the exception into a practical application nor amount to significantly more – See MPEP 2106.05(h))
Accordingly, under Step 2A Prong 2 and Step 2B, this additional element does not integrate the
abstract idea into practical application because it does not impose any meaningful limits on practicing
the abstract idea, as discussed above in the rejection of claim 1. The claim does not include additional
elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception.
Regarding Claim 6:
Step 2A Prong 1: See the rejection of Claim 1 above, which Claim 6 depends on.
wherein the backward propagation operation that proceeds in parallel speculates a forward propagation outcome at a current time based on a result of the forward propagation operation of at least a previous time and performs backward propagation based on the speculated result (mathematical concept – requires performing mathematical calculations to estimate a forward propagation outcome based on prior results and to perform backward propagation computations using the estimated values)
Step 2A Prong 2 & Step 2B:
Accordingly, under Step 2A Prong 2 and Step 2B, there are no additional elements that integrate the
abstract idea into practical application. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial
exception.
Regarding Claim 7:
Step 1: Claim 7 is a method type claim. Therefore, Claims 7-11 fall within one of the four statutory
categories (i.e., process, machine, manufacture, or composition of matter).
2A Prong 1: If a claim limitation, under its broadest reasonable interpretation, covers performance
of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. If a claim limitation, under its broadest reasonable
interpretation, covers performance of the limitation by mathematical calculation but for the recitation
of generic computer components, then it falls within the “Mathematical Concepts” grouping of abstract
ideas.
a forward propagation operation of performing a forward propagation (mathematical concept - mathematical calculations used to compute outputs of a neural network from input data)
a backward propagation operation of performing a backward propagation (mathematical concept - backward propagation in a neural network involves mathematical calculations such as computing error derivatives, gradients, and propagating those values through layers to adjust model parameters)
and a weight update operation of performing a weight update (mathematical concept - involves performing iterative mathematical calculations to perform a weight update)
wherein the forward propagation operation and the backward propagation operation that are repeatedly performed after an initial execution proceed in parallel at least once via speculative backpropagation using an accumulated history of past forward outcomes to concurrently run the backward propagation operation with a subsequent forward propagation operation (mathematical concept - involves performing iterative mathematical calculations including forward propagation, backward propagation, and parameter updates)
Step 2A Prong 2 & Step 2B:
[…] a processor and a memory, wherein the computing device executes the deep learning model via an integrated hardware training accelerator […] recited at a high-level of generality (i.e., a generic processor, memory, a communication interface, a user interface and accelerator) such that it amounts to no more than mere instructions to apply the exception using generic computer components)
wherein the forward propagation operation, the backward propagation operation, and the weight update operation are repeatedly performed (Field of Use – limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying the forward propagation operation, the backward propagation operation, and the weight update operation are repeatedly performed does not integrate the exception into a practical application nor amount to significantly more – See MPEP 2106.05(h))
For the reasons above, Claim 7 is rejected as being directed to an abstract idea without significantly more. This rejection applies equally to dependent claims 7 - 11. The additional limitations of the dependent claims are addressed below.
Regarding Claim 8:
Step 2A Prong 1: See the rejection of Claim 7 above, which Claim 8 depends on.
wherein the backward propagation operation performed in parallel at a time t=i is performed using a forward propagation outcome at a time t=i that is speculated using a forward propagation outcome at a time t=(i-1) and an accumulated history of previous forward propagation outcomes (mathematical concept – requires performing calculations to speculate the forward propagation outcome at time t=i using the forward propagation outcome at time t=(i) and the forward propagation outcome at time t=(i−1) and using the speculated forward propagation outcome to perform the backward propagation operation)
Step 2A Prong 2 & Step 2B:
Accordingly, under Step 2A Prong 2 and Step 2B, there are no additional elements that integrate the
abstract idea into practical application. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial
exception.
Regarding Claim 9:
Step 2A Prong 1: See the rejection of Claim 8 above, which Claim 9 depends on.
wherein the forward propagation outcome speculated at the time t=i is speculated by assigning a greater weight to […] than […] at the time t=(i-1) (mental process - assigning a greater weight may be performed by a user observing and analyzing the results and using judgment or evaluation to assign a greater weight to one result than the other)
Step 2A Prong 2 & Step 2B:
[…] an accumulated result of the deep learning model […] a result of the deep learning model [..] (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of using a result of the deep learning model without significantly more)
Accordingly, under Step 2A Prong 2 and Step 2B, this additional element does not integrate the
abstract idea into practical application because it does not impose any meaningful limits on practicing
the abstract idea, as discussed above in the rejection of claim 8. The claim does not include additional
elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception.
Regarding Claim 10:
Step 2A Prong 1: See the rejection of Claim 8 above, which Claim 10 depends on.
Step 2A Prong 2 & Step 2B:
wherein an activation status of each of the neurons at the time t=i is speculated based on the activation tendency of each of the neurons included in the deep learning model up to the time t=i-1 (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of using the neurons included in the deep learning model without significantly more)
Accordingly, under Step 2A Prong 2 and Step 2B, this additional element does not integrate the
abstract idea into practical application because it does not impose any meaningful limits on practicing
the abstract idea, as discussed above in the rejection of claim 8. The claim does not include additional
elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception.
Regarding Claim 11:
Step 2A Prong 1: See the rejection of Claim 7 above, which Claim 11 depends on.
Step 2A Prong 2 & Step 2B:
wherein, in a training process for a second task performed after training for a first task is completed, the weight update operation is performed based on an activation tendency of […] (Field of Use – limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that in a training process for a second task performed after training for a first task is completed, the weight update operation is performed based on an activation tendency does not integrate the exception into a practical application nor amount to significantly more – See MPEP 2106.05(h))
[…] each of neurons included in the deep learning model in a process in which training for the first task proceeds (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of using the neurons included in the deep learning model without significantly more)
Accordingly, under Step 2A Prong 2 and Step 2B, this additional element does not integrate the
abstract idea into practical application because it does not impose any meaningful limits on practicing
the abstract idea, as discussed above in the rejection of claim 7. The claim does not include additional
elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-6 and 11 are rejected under 35 U.S.C. 103 as being unpatentable over Butvinik et al. (hereafter Butvinik) (US 20220261633), in view of Park et al. (hereinafter Park, a non-patent literature reference titled “Speculative backpropagation for CNN parallel training”) and further in view of Paik et al. (hereinafter Paik, a non-patent literature reference titled “Overcoming catastrophic forgetting by neuron-level plasticity control”).
Regarding Claim 1, Butvinik teaches:
A continual learning method of a deep learning model performed by a computing device comprising at least a processor and a memory […] to mitigate catastrophic forgetting […] for continual learning for a second task and an nᵗʰ task for the deep learning model trained for a first task, (Butvinik, Par. [0004], “when a neural network is used to learn a sequence of tasks, the learning of the later tasks may degrade the performance of the models learned for the earlier tasks. Our human brains, however, seem to have this remarkable ability to learn a large number of different tasks without any of them negatively interfering with each other. OIL algorithms try to achieve this same ability for neural networks and to solve the catastrophic forgetting problem. Thus, in essence, continual learning performs incremental learning of new tasks”, & Par. [0031], “Given a sequence of supervised learning tasks T=(T.sub.1, T.sub.2, . . . T.sub.N), embodiments of the invention sequentially train those tasks in the given sequence such that the learning of each new task will not forget the models learned for the previous tasks”, & Par. [0149], “Computing device 100 may include a controller or computer processor 105 that may be, for example, a central processing unit processor (CPU), a chip or any suitable computing device, an operating system 115, a memory 120”, thus a continual learning method of a deep learning model performed by a computing device comprising at least a processor and a memory, for continual learning of a second task and an nᵗʰ task after training a first task, while mitigating catastrophic forgetting, is disclosed because Butvinik teaches a deep neural network trained sequentially over tasks T1 through TN while retaining knowledge of earlier tasks, with the training performed by a computing device having a processor and memory), the continual learning method comprising:
[…] included in the deep learning model, in a process in which training for the first task proceeds (Butvinik, Par. [0030], “In conventional artificial neural networks, all neurons in the hidden layer are initially activated,” & Par. [0031], “Given a sequence of supervised learning tasks T=(T1, T2, . . . TN), embodiments of the invention sequentially train those tasks in the given sequence such that the learning of each new task will not forget the models learned for the previous tasks,” thus neurons included in the deep learning model and training for a first task are disclosed because Butvinik teaches neurons within the neural network and sequential training beginning with task T1 before subsequent tasks T2 through TN)
Butvinik does not explicitly teach […] wherein the deep learning model is executed on an artificial intelligence (AI) hardware accelerator integrated within the computing device […] while accelerating training speed […], a forward propagation operation of performing a forward propagation, a backward propagation operation of performing a backward propagation, a weight update operation of performing a weight update, wherein the forward propagation operation, the backward propagation operation, and the weight update operation are repeatedly performed, and wherein the weight update operation is performed […].
However, Park teaches:
[…] wherein the deep learning model is executed on an artificial intelligence (AI) hardware accelerator integrated within the computing device […] while accelerating training speed […] (Park, Page 1 – Abstract & Section 1, “We implemented the proposed parallel model with CNNs in both software and hardware” and “The parallel training reduces the training time by 34% in CIFAR-100,” & Page 5 – Section V-A, “The UltraScale+ is composed of the processing system (PS) and programmable logic (PL) sections: The PS has a quad-core Cortex-A53 processor operating at 1.5 GHz. The PL is configured with the hardware accelerator in our work,” & Page 5 – Section V-B, “We designed two C functions performing the forward propagation and the speculative backpropagation. Then, the C functions were translated to hardware by the SDSoC,” thus execution of the deep learning model on an AI hardware accelerator integrated within the computing device while accelerating training speed is disclosed because Park implements neural network training operations in hardware within the programmable logic portion of a processor based SoC and performs the training operations in parallel, thereby reducing the neural network training time)
a forward propagation operation of performing a forward propagation (Park, Page 2 – Section III, “For the neural network training, three sequential operations are performed: forward propagation, backpropagation, and weight update. In the forward propagation process, input data is propagated from the input layer to the output layer”, thus a forward propagation operation of performing a forward propagation is disclosed, because Park identifies forward propagation as one of the operations performed during neural network training. Park’s propagation of input data from the input layer through the neural network to the output layer corresponds to performing the forward propagation operation because the data is processed in the forward direction through the network to generate an output)
a backward propagation operation of performing a backward propagation (Park, Page 3 – Section III, “The backpropagation is used to adjust the weights (wl ij) by calculating derivatives. This phase begins from the output layer, which is based on Softmax in our work. The derivative in the output layer is expressed in Eq. (4)”, thus a backward propagation operation of performing a backward propagation is disclosed, because Park describes backpropagation as a neural network training operation that begins at the output layer and propagates error derived information backward through the network to calculate derivatives used for adjusting the weights)
a weight update operation of performing a weight update (Park, Page 3 – Section III, “In the ANN training, the weights are adjusted based on the errors computed in the backpropagation. First, wl ij is calculated by multiplying the result of the backpropagation δl i and the result of the forward propagation yl−1 j , as shown in Eq. (7). Weights (wl ij) are then updated according to the learning rate η, which determines the degree of learning in Eq. (8). This process is repeated for all the weights”, thus a weight update operation of performing a weight update is disclosed, because Park teaches calculating a weight change amount from the backpropagation and forward propagation results and then updating the network weights according to the calculated amount and the learning rate)
wherein the forward propagation operation, the backward propagation operation, and the weight update operation are repeatedly performed (Park, Page 2 – Section III, “For the neural network training, three sequential operations are performed: forward propagation, backpropagation, and weight update”, & Page 3 – Section III, “Weights (wl ij) are then updated according to the learning rate η, which determines the degree of learning in Eq. (8). This process is repeated for all the weights”, thus the forward propagation operation, the backward propagation operation, and the weight update operation are repeatedly performed is disclosed, because Park teaches neural network training as a recurring sequence that includes forward propagation, backpropagation, and weight update, and further teaches repeating the weight update process for the network weights during training)
wherein the weight update operation is performed […] (Park, Page 3 – Section III, “In the ANN training, the weights are adjusted based on the errors computed in the backpropagation. First, wl ij is calculated by multiplying the result of the backpropagation δl i and the result of the forward propagation yl−1 j , as shown in Eq. (7). Weights (wl ij) are then updated according to the learning rate η, which determines the degree of learning in Eq. (8)”, thus Park teaches performing the weight update operation because Park calculates a weight-change amount from the forward and backward propagation results and updates the neural network weights according to the calculated amount and learning rate)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to combine Butvinik with Park by incorporating Park’s parallel hardware accelerated neural network training into Butvinik’s continual learning system. Butvinik teaches sequentially training a neural network on a plurality of tasks while retaining knowledge learned from previous tasks to mitigate catastrophic forgetting, while also recognizing that as tasks accumulate, the associated training data can result in prohibitively time consuming training sessions and identifying a need to efficiently train neural networks. Park teaches accelerating neural network training by performing forward and backward propagations in parallel using speculative backpropagation implemented with a hardware training accelerator. Therefore, a POSITA would have been motivated to incorporate Park’s parallel hardware accelerated training into Butvinik’s continual learning system so that the successive neural network tasks can be trained more efficiently and in less time while maintaining Butvinik’s continual learning approach for retaining knowledge from previous tasks, thereby accelerating the training process while mitigating catastrophic forgetting (Butvinik, Pars. [0005]-[0006], “As tasks accumulate, however, their associated training data also accumulates, resulting in prohibitively large amount of training data and prohibitively time-consuming training sessions that train based on all past and current training data. Accordingly, there is therefore a need in the art to overcome this limitation and efficiently train neural networks to maintain expertise on tasks which they have not experienced for a long time,” & Park, Page 1 – Section I, “It enables performing the forward and backward propagations in parallel. The core idea is speculating the forward outcomes for backpropagation. This paper shows that the speculative backpropagation speeds up the training time without the prediction accuracy loss,” & Page 2 – Section II, “the training accelerator in both software and hardware. The experiments show that the parallel training reduces the training time by up to 38%”)
Butvinik combined with Park does not explicitly teach […] based on an activation tendency of each of neurons […].
However, Paik teaches:
[…] based on an activation tendency of each of neurons […] (Paik, Page 3 – Section 3.1, “We measure the importance of each neuron by the normalized Taylor criterion as shown in eq. (2) and (3). Then, we take their moving averages as eq. (4) to reduce the fluctuation of the measurements and to improve the learning stability” & Page 4 – Section 3.1, “
C
i
t
is the importance value of the i-th neuron at training step t, ni denotes the activation of the i-th neuron. L is the loss, and N layer is the number of nodes on the layer, and δ is a hyperparameter,” & Page 4 – Section 3.2, “We control the plasticity of each neuron by applying different learning rates
η
i
for each neuron,” thus being based on an activation tendency of each neuron is disclosed because Paik determines a neuron specific importance value using the activation of that neuron, accumulates the value over successive training steps using a moving average, and uses the resulting neuron specific importance value to control the learning rate for that neuron. Paik’s accumulated activation-derived importance value corresponds to an activation tendency because it represents the neuron’s activation related behavior accumulated during training)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to further combine Butvinik and Park with Paik by incorporating Paik’s neuron level learning rate control into the continual learning system of Butvinik as accelerated by Park. Butvinik teaches sequentially training a neural network on multiple tasks while retaining knowledge learned from previous tasks to mitigate catastrophic forgetting, and Park teaches accelerating the neural network training process using parallel forward and backward propagation. Paik also addresses catastrophic forgetting during sequential neural network training and teaches determining an importance value for each neuron and preserving knowledge from previous tasks by applying lower learning rates to important neurons. Therefore, a POSITA would have been motivated to incorporate Paik’s neuron level learning rate control into the Butvinik/Park system so that neurons important to previously learned tasks are subjected to lower learning rates during subsequent task training, thereby reducing the overwriting of previously learned information and improving the retention of prior task knowledge while maintaining the accelerated training provided by Park (Paik, Page 1 – Abstract, “To address the issue of catastrophic forgetting in neural networks, we propose a novel, simple, and effective solution called neuron-level plasticity control (NPC). While learning a new task, the proposed method preserves the knowledge for the previous tasks by controlling the plasticity of the network at the neuron level. NPC estimates the importance value of each neuron and consolidates important neurons by applying lower learning rates,” & Page 8 – Section 5, “NPC is effective in preserving old knowledge since it consolidates all the incoming pathways to important neurons”)
Regarding Claim 2, Butvinik and Park combined with Paik teaches all the limitations of claim 1 as cited above and Park further teaches:
wherein the forward propagation operation, the backward propagation operation, and the weight update operation are repeatedly performed […] (Park, Page 2 – Section III, “For the neural network training, three sequential operations are performed: forward propagation, backpropagation, and weight update,” & Page 3 – Section III, “Weights (wᶫᵢⱼ) are then updated according to the learning rate η, which determines the degree of learning in Eq. (8). This process is repeated for all the weights,” thus the forward propagation operation, the backward propagation operation, and the weight update operation are repeatedly performed because Park teaches neural-network training using forward propagation, backpropagation, and weight update, with the weight-update process repeated during training)
Butvinik further teaches:
[…] until weights of the deep learning model converge to a predetermined range (Butvinik, Par. [0058], “/Learning 1.sup.st task 4: for n = 0,..., until convergence do 5: sample minibatch from T.sub.1 //training DG 6: Minimize ψ.sub.wae and update DS; //ψ.sub.wae Wasserstein auto-encoder //training DPP (dynamic parameter propagator) add C //p.sub.n an C.sub.n are batch of generated parameters and adapted computers respectively 7: Compute p.sub.n using eq.1 - eq.6 8: Construct C.sub.n using p.sub.n and φ.sub.0 9: Minimize ψ.sub.ce and update g(.Math.) and C end for //Learning sub sequent tasks 10: for i = 2; i <= N; i++ do 11: generate replayed samples x.sub.m′ by DG (data generator) 12: for n = 0, ..., until convergence do”, & Par. [0147], “Model parameters may refer to weights of a neural network, hyper-parameters such as an activation function, or more generally to any other model parameters, or other explicit and implicit parameters”, thus performing training until the weights of the deep learning model converge is disclosed because Butvinik teaches repeating the training procedure until convergence and identifies neural-network weights as model parameters)
Regarding Claim 3, Butvinik and Park combined with Paik teaches all the limitations of claim 1 as cited above and Paik further teaches:
wherein […the weight update operation…] limits weight update for a neuron of which activation tendency has a value greater than a predetermined value in the process in which training for […the first task…] proceeds (Paik, Page 3 – Section 3.1, “We measure the importance of each neuron by the normalized Taylor criterion as shown in eq. (2) and (3). Then, we take their moving averages as eq. (4) to reduce the fluctuation of the measurements and to improve the learning stability,” & Page 4 – Section 3.1, “C(t) is the importance value of the i-th neuron at training step t, ni denotes the activation of the i-th neuron. L is the loss, and Nlayer is the number of nodes on the layer, and δ is a hyperparameter. If a node is shared in multiple positions, e.g. the convolution filter in CNN, we average the importance values from all positions before computing its absolute value,” & Page 5 – Section 3.2, “In case of
C
i
>
β
,
l
η
i
is strictly increasing w.r.t.
η
i
, which leads to
η
i
*
=
0
. Note that
η
i
*
=
0
at
C
i
=
β
in eq. (9), which makes the two functions continuously connected” and “a larger
C
i
draws a smaller
η
i
*
, thereby consolidating the important neurons in the subsequent learning,” & Page 6 – Section 4, “First, we trained a CNN for Task 1 of iCIFAR100 for 30 epochs and recorded the neuron activation values of the second top neurons(just before the final classifier) extracted from randomly chosen 256 samples. (512 neurons × 256 samples = 131,072 data points in total.) Then, we trained the CNN for Task 2 for another 30 epochs,” thus limiting weight update for a neuron having an activation tendency greater than a predetermined value in the process in which training for the first task proceeds is disclosed because Paik determines a neuron specific importance value
C
i
from the activation
n
i
of the neuron and accumulates the importance value using a moving average during training. Paik further compares
C
i
with the predetermined hyperparameter
β
, and when
C
i
>
β
, sets the neuron specific learning rate
η
i
*
to zero, thereby limiting the corresponding weight update. Paik further teaches training the CNN for Task 1 and recording neuron activation values before subsequently training Task 2, such that the activation related importance determined for neurons associated with Task 1 is used to protect important neurons during subsequent learning)
Regarding Claim 4, Butvinik and Park combined with Paik teaches all the limitations of claim 1 as cited above and Park further teaches:
obtaining a gradient (Park, Page 3 – Section III, “The backpropagation is used to adjust the weights (wl ij) by calculating derivatives. This phase begins from the output layer, which is based on Softmax in our work. The derivative in the output layer is expressed in Eq. (4). It calculates the difference between the forward propagation outcome(yo z) and target output (tz). The derivative of the error with respect to the weight is calculated with Eq. (5),” thus obtaining a gradient is disclosed because Park calculates the derivative of the error with respect to the neural network weight during backpropagation, which corresponds to obtaining the gradient used for updating the weight)
reducing an update amount of a weight by multiplying the obtained gradient by a predetermined constant (r) (Park, Page 3 – Section III, “Weights (
w
i
j
l
) are then updated according to the learning rate
η
, which determines the degree of learning in Eq. (8),” where Eq. (8) provides
w
i
j
l
=
w
i
j
l
-
Δ
w
i
j
l
η
, thus, reducing an update amount of a weight by multiplying the obtained gradient by a predetermined constant is disclosed because Park multiplies the weight change value obtained from backpropagation by the learning rate factor
η
when updating the weight. Park’s learning rate factor
η
corresponds to the predetermined constant
r
because it controls the magnitude of the weight update)
updating the weight using the reduced updated amount (Park, Page 3 – Section III, “Weights (
w
i
j
l
) are then updated according to the learning rate
η
, which determines the degree of learning in Eq. (8),” where Eq. (8) provides
w
i
j
l
=
w
i
j
l
-
Δ
w
i
j
l
η
, thus updating the weight using the reduced update amount is disclosed because Park updates the existing weight by subtracting the weight change amount after that amount has been scaled by the learning rate factor)
Paik further teaches:
for a neuron of which activation tendency has a value greater than a predetermined value in the process in which training for the first task proceeds (Paik, Page 4 – Section 3.1, “C(t) is the importance value of the i-th neuron at training step t, ni denotes the activation of the i-th neuron. L is the loss, and Nlayer is the number of nodes on the layer, and δ is a hyperparameter. If a node is shared in multiple positions, e.g. the convolution filter in CNN, we average the importance values from all positions before computing its absolute value,” & Page 5 – Section 3.2, “In case of
C
i
>
β
,
l
η
i
is strictly increasing w.r.t.
η
i
, which leads to
η
i
*
=
0
. Note that
η
i
*
=
0
at
C
i
=
β
in eq. (9), which makes the two functions continuously connected” and “a larger
C
i
draws a smaller
η
i
*
, thereby consolidating the important neurons in the subsequent learning,”, & Page 6 – Section 4, “First, we trained a CNN for Task 1 of iCIFAR100 for 30 epochs and recorded the neuron activation values of the second top neurons(just before the final classifier) extracted from randomly chosen 256 samples. (512 neurons × 256 samples = 131,072 data points in total.) Then, we trained the CNN for Task 2 for another 30 epochs,” thus a neuron having an activation related value greater than a predetermined value during training associated with the first task is disclosed because Paik determines a neuron specific importance value from the neuron activation, compares the importance value
C
i
with the predetermined hyperparameter
β
, and identifies the
C
i
>
β
condition for controlling learning of that neuron. Paik further records neuron activation values from Task 1 before subsequent Task 2 training)
Regarding Claim 5, Butvinik and Park combined with Paik teaches all the limitations of claim 1 as cited above and Park further teaches:
wherein, in a training process for the second task, the forward propagation operation and the backward propagation operation that are repeatedly performed proceed in parallel at least once (Page 3 – Section IV, “In our proposed method shown in Figure 1(b), the forward and backward computations occur at the same time. The backward propagation in Figure 1(b) is based on the accumulated previous forward outcomes,” and Figure 1(b), “Simultaneous execution of forward and backward propagations with the speculative backpropagation,” thus the forward propagation operation and the backward propagation operation proceeding in parallel at least once is disclosed because Park performs the forward and backward computations simultaneously using speculative backpropagation)
Regarding Claim 6, Butvinik and Park combined with Paik teaches all the limitations of claim 5 as cited above and Park further teaches:
wherein the backward propagation operation that proceeds in parallel speculates a forward propagation outcome at a current time based on a result of the forward propagation operation of at least a previous time and performs backward propagation based on the speculated result (Park, Page 3 – Section IV, “if it is feasible to speculate the forward outcomes in advance, the backpropagation can be performed simultaneously with the forward computation. We have found one interesting behavior in the ANN training that makes the speculation possible; The Softmax and ReLU outcomes for the same labels in the temporally near-forward propagations tend to be very similar. Thus, the previous forward outcomes can be used for the current backpropagation,” Page 4 – Section IV-A, “These are used for speculative backpropagation at time t(i), instead of using current forward outcomes,” and “At time t(i), the forward propagation processes a current input image. At the same time, the stored data (red neurons in Figure 3) for the same label is used for the backpropagation,” thus speculating a forward propagation outcome at a current time based on a result of the forward propagation operation of at least a previous time and performing backward propagation based on the speculated result is disclosed because Park uses stored and accumulated previous forward propagation outcomes to estimate the current forward outcome at time t(i) and performs the concurrent backward propagation using that speculated outcome)
Regarding Claim 11, Butvinik combined with Park teaches all the limitations of claim 7 as cited above:
Butvinik combined with Park does not explicitly teach wherein, in a training process for a second task performed after training for a first task is completed, […the weight update operation…] is performed based on an activation tendency of each of neurons included in […the deep learning model…] in a process in which training for […the first task…] proceeds.
However, Paik teaches:
wherein, in a training process for a second task performed after training for a first task is completed, […the weight update operation…] is performed based on an activation tendency of each of neurons included in […the deep learning model…] in a process in which training for […the first task…] proceeds (Paik, Page 3 – Section 3.1, “We measure the importance of each neuron by the normalized Taylor criterion as shown in eq. (2) and (3). Then, we take their moving averages as eq. (4),” & Page 4 – Section 3.1, “Ci(t) is the importance value of the i-th neuron at training step t,
n
i
denotes the activation of the i-th neuron,” & Page 5 – Section 3.2, “a larger
C
i
draws a smaller
η
i
*
, thereby consolidating the important neurons in the subsequent learning,” & Page 6 – Section 4, “we trained a CNN for Task 1 of iCIFAR100 for 30 epochs and recorded the neuron activation values of the second top neurons(just before the final classifier) extracted from randomly chosen 256 samples. (512 neurons × 256 samples = 131,072 data points in total.) Then, we trained the CNN for Task 2 for another 30 epochs,” thus performing the weight update operation during training of the second task based on an activation tendency of neurons determined during training of the first task is disclosed because Paik measures a neuron specific importance value using the neuron’s activation, accumulates the importance value over training steps using a moving average, and uses the accumulated importance value to control the neuron-specific learning rate during subsequent learning, such that neurons determined to be more important during the first-task training receive smaller weight updates during the subsequent second-task training. Paik further teaches training Task 1 and recording neuron activation values before subsequently training Task 2)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to further combine Butvinik and Park with Paik by incorporating Paik’s neuron level learning rate control into the continual learning system of Butvinik as accelerated by Park. Butvinik teaches sequentially training a neural network on multiple tasks while retaining knowledge learned from previous tasks to mitigate catastrophic forgetting, and Park teaches accelerating the neural network training process using parallel forward and backward propagation. Paik also addresses catastrophic forgetting during sequential neural network training and teaches determining an importance value for each neuron and preserving knowledge from previous tasks by applying lower learning rates to important neurons. Therefore, a POSITA would have been motivated to incorporate Paik’s neuron level learning rate control into the Butvinik/Park system so that neurons important to previously learned tasks are subjected to lower learning rates during subsequent task training, thereby reducing the overwriting of previously learned information and improving the retention of prior task knowledge while maintaining the accelerated training provided by Park (Paik, Page 1 – Abstract, “To address the issue of catastrophic forgetting in neural networks, we propose a novel, simple, and effective solution called neuron-level plasticity control (NPC). While learning a new task, the proposed method preserves the knowledge for the previous tasks by controlling the plasticity of the network at the neuron level. NPC estimates the importance value of each neuron and consolidates important neurons by applying lower learning rates,” & Page 8 – Section 5, “NPC is effective in preserving old knowledge since it consolidates all the incoming pathways to important neurons”)
Claims 7-10 are rejected under 35 U.S.C. 103 as being unpatentable over Butvinik et al. (hereafter Butvinik) (US 20220261633), in view of Park et al. (hereinafter Park, a non-patent literature reference titled “Speculative backpropagation for CNN parallel training”).
Regarding Claim 7, Butvinik teaches:
A method of training a deep learning model performed by a computing device comprising at least a processor and a memory, wherein the computing device executes the deep learning model via an integrated hardware training accelerator configured to improve parallel processing efficiency (Butvinik, Par. [0149], “Computing device 100 may include a controller or computer processor 105 that may be, for example, a central processing unit processor (CPU), a chip or any suitable computing device, an operating system 115, a memory 120,” Page 5 – Section V-A, “The UltraScale+ is composed of the processing system (PS) and programmable logic (PL) sections: The PS has a quad-core Cortex-A53 processor operating at 1.5 GHz. The PL is configured with the hardware accelerator in our work,” & Page 5 – Section V-B, “We designed two C functions performing the forward propagation and the speculative backpropagation. Then, the C functions were translated to hardware by the SDSoC” and “#pragma async to generate two different hardware operating in parallel,” thus a method of training a deep learning model performed by a computing device comprising a processor and memory and executing the deep learning model via an integrated hardware training accelerator configured to improve parallel processing efficiency is disclosed because Butvinik provides a computing device having a processor and memory for training a machine-learning model, while Park implements neural-network forward propagation and speculative backpropagation in a hardware accelerator integrated with a processor based SoC and configures separate hardware to operate the training computations in parallel) the method comprising:
Butvinik does not explicitly teach a forward propagation operation of performing a forward propagation, a backward propagation operation of performing a backward propagation, a weight update operation of performing a weight update, wherein the forward propagation operation, the backward propagation operation, and the weight update operation are repeatedly performed, and wherein the forward propagation operation and the backward propagation operation that are repeatedly performed after an initial execution proceed in parallel at least once via speculative backpropagation using an accumulated history of past forward outcomes to concurrently run the backward propagation operation with a subsequent forward propagation operation.
However, Park teaches:
a forward propagation operation of performing a forward propagation (Park, Page 2 – Section III, “For the neural network training, three sequential operations are performed: forward propagation, backpropagation, and weight update. In the forward propagation process, input data is propagated from the input layer to the output layer”, thus a forward propagation operation of performing a forward propagation is disclosed, because Park identifies forward propagation as one of the operations performed during neural network training. Park’s propagation of input data from the input layer through the neural network to the output layer corresponds to performing the forward propagation operation because the data is processed in the forward direction through the network to generate an output)
a backward propagation operation of performing a backward propagation (Park, Page 3 – Section III, “The backpropagation is used to adjust the weights (wl ij) by calculating derivatives. This phase begins from the output layer, which is based on Softmax in our work. The derivative in the output layer is expressed in Eq. (4)”, thus a backward propagation operation of performing a backward propagation is disclosed, because Park describes backpropagation as a neural network training operation that begins at the output layer and propagates error derived information backward through the network to calculate derivatives used for adjusting the weights)
a weight update operation of performing a weight update (Park, Page 3 – Section III, “In the ANN training, the weights are adjusted based on the errors computed in the backpropagation. First, wl ij is calculated by multiplying the result of the backpropagation δl i and the result of the forward propagation yl−1 j , as shown in Eq. (7). Weights (wl ij) are then updated according to the learning rate η, which determines the degree of learning in Eq. (8). This process is repeated for all the weights”, thus a weight update operation of performing a weight update is disclosed, because Park teaches calculating a weight change amount from the backpropagation and forward propagation results and then updating the network weights according to the calculated amount and the learning rate)
wherein the forward propagation operation, the backward propagation operation, and the weight update operation are repeatedly performed (Park, Page 2 – Section III, “For the neural network training, three sequential operations are performed: forward propagation, backpropagation, and weight update”, & Page 3 – Section III, “Weights (wl ij) are then updated according to the learning rate η, which determines the degree of learning in Eq. (8). This process is repeated for all the weights”, thus the forward propagation operation, the backward propagation operation, and the weight update operation are repeatedly performed is disclosed, because Park teaches neural network training as a recurring sequence that includes forward propagation, backpropagation, and weight update, and further teaches repeating the weight update process for the network weights during training)
wherein the forward propagation operation and the backward propagation operation that are repeatedly performed after an initial execution proceed in parallel at least once via speculative backpropagation using an accumulated history of past forward outcomes to concurrently run the backward propagation operation with a subsequent forward propagation operation (Park, Page 3 – Section IV-A, “the past result could be used in the current backpropagation. To take advantage of this behavior, we store the Softmax output per label when the forward propagation finishes,” & Page 4 – Section IV-A, “Parallel execution of forward and backward computations with speculative backpropagation, which uses the accumulated previous forward outcomes,” and “the forward propagation processes a current input image. At the same time, the stored data (red neurons in Figure 3) for the same label is used for the backpropagation”, & “Eq. (9) considering both the most recent forward outcome and the history of the forward computations helps reduce the difference. In other words, Eq. (9) accumulates all the Softmax outcomes so far,” thus the forward propagation operation and the backward propagation operation that are repeatedly performed after an initial execution proceeding in parallel at least once via speculative backpropagation using an accumulated history of past forward outcomes is disclosed because Park first stores forward propagation outcomes after the forward propagation finishes, accumulates the previous forward outcomes over successive training steps, and then uses those stored and accumulated outcomes for speculative backpropagation. Park further performs the current forward propagation and the backpropagation at the same time, such that the backward propagation based on the accumulated past forward outcomes concurrently executes with a subsequent forward propagation operation)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to combine Butvinik with Park by incorporating Park’s parallel hardware accelerated neural network training into Butvinik’s continual learning system. Butvinik teaches sequentially training a neural network on a plurality of tasks while retaining knowledge learned from previous tasks to mitigate catastrophic forgetting, while also recognizing that as tasks accumulate, the associated training data can result in prohibitively time consuming training sessions and identifying a need to efficiently train neural networks. Park teaches accelerating neural network training by performing forward and backward propagations in parallel using speculative backpropagation implemented with a hardware training accelerator. Therefore, a POSITA would have been motivated to incorporate Park’s parallel hardware accelerated training into Butvinik’s continual learning system so that the successive neural network tasks can be trained more efficiently and in less time while maintaining Butvinik’s continual learning approach for retaining knowledge from previous tasks, thereby accelerating the training process while mitigating catastrophic forgetting (Butvinik, Pars. [0005]-[0006], “As tasks accumulate, however, their associated training data also accumulates, resulting in prohibitively large amount of training data and prohibitively time-consuming training sessions that train based on all past and current training data. Accordingly, there is therefore a need in the art to overcome this limitation and efficiently train neural networks to maintain expertise on tasks which they have not experienced for a long time,” & Park, Page 1 – Section I, “It enables performing the forward and backward propagations in parallel. The core idea is speculating the forward outcomes for backpropagation. This paper shows that the speculative backpropagation speeds up the training time without the prediction accuracy loss,” & Page 2 – Section II, “the training accelerator in both software and hardware. The experiments show that the parallel training reduces the training time by up to 38%”)
Regarding Claim 8, Butvinik combined with Park teaches all the limitations of claim 7 as cited above and Park further teaches:
wherein the backward propagation operation performed in parallel at a time t=i is performed using a forward propagation outcome at a time t=i that is speculated using a forward propagation outcome at a time t=(i-1) and an accumulated history of previous forward propagation outcomes (Park, Page 4 – Section IV-A, Figure 3, “The red neurons are accumulated Softmax outcomes until time The red neurons are accumulated Softmax outcomes until time t(i − 1), and the blue neurons are previous ReLU outcomes. These are used for speculative backpropagation at time t(i), instead of using current forward outcomes (yellow and green neurons),” & “At time t(i), the forward propagation processes a current input image. At the same time, the stored data (red neurons in Figure 3) for the same label is used for the backpropagation,” & “Eq. (9) con sidering both the most recent forward outcome and the history of the forward computations helps reduce the difference. In other words, Eq. (9) accumulates all the Softmax outcomes so far,” thus performing the parallel backward propagation at time t=i using a speculated forward propagation outcome based on the forward propagation outcome at time t=(i−1) and an accumulated history of previous forward propagation outcomes is disclosed because Park’s speculative backpropagation at time t(i) uses stored forward outcomes accumulated through time t(i−1), and Eq. (9) forms the speculated outcome using both the most recent forward outcome, corresponding to time t(i−1), and the accumulated history of earlier forward outcomes)
Regarding Claim 9, Butvinik combined with Park teaches all the limitations of claim 8 as cited above and Park further teaches:
wherein the forward propagation outcome speculated at the time t=i is speculated by assigning a greater weight to an accumulated result of the deep learning model than a result of the deep learning model at the time t=(i-1) (Park, Page 4 – Section IV-A, “Eq. (9) considering both the most recent forward outcome and the history of the forward computations helps reduce the difference. In other words, Eq. (9) accumulates all the Softmax outcomes so far. α and β are weights for the most recent outcome and the previous history, respectively,” & “In the MNIST, a more weight on the accumulated Softmax outcome (yo) achieves roughly 0.2% smaller mean-difference (α = 1/3, β = 2/3),” the forward propagation outcome speculated at time t=i by assigning a greater weight to the accumulated result than to the result at time t=(i−1) is disclosed because Park calculates the speculative forward outcome using both the most recent forward outcome and the accumulated history of previous forward outcomes, where α corresponds to the weight assigned to the most recent result and β corresponds to the weight assigned to the accumulated history. Park further teaches an example in which α=1/3 and β=2/3, thereby assigning a greater weight to the accumulated result than to the most recent forward propagation result)
Regarding Claim 10, Butvinik combined with Park teaches all the limitations of claim 8 as cited above and Park further teaches:
wherein an activation status of each of the neurons at the time t=i is speculated based on the activation tendency of each of the neurons included in the deep learning model up to the time t=i-1 (Park, Page 4 – Section IV-B, “While training ANN, we found a characteristic that activated neurons in the past are more likely to be activated in the future for inputs with the same label,” and “Table 1 - 4 show our experiment outcomes showing the probability that the same neurons are activated in the previous and current forward propagations,” & Page 5 – Section IV-B, “ Based on this observation and the high activation probability of the same neurons, Eq.(6) can be computed speculatively. We store the ReLU outputs (f(ul j))per label in the hidden layers when the forward propagation finishes. Then, the speculative backpropagation is performed using the stored values,” thus an activation status of each neuron at time t=i being speculated based on the activation tendency of the neurons up to time t=i−1 is disclosed because Park observes from prior forward propagations that neurons activated in the past are likely to be activated again, determines the probability that the same neurons remain activated across previous and current forward propagations, stores the prior ReLU activation outputs, and uses those stored values to speculate whether each neuron is activated or non-activated during the current speculative backpropagation)
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to MAHLIET ADMASU whose telephone number is (571)272-0034. The examiner can normally be reached Mon-Fri, 8am-5pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Alexey Shmatov can be reached at (571)270-3428. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/M.T.A./
Examiner, Art Unit 2123
/ALEXEY SHMATOV/Supervisory Patent Examiner, Art Unit 2123