Prosecution Insights
Last updated: October 01, 2026
Application No. 18/187,030

METHOD AND DEVICE WITH INFERENCE-BASED DIFFERENTIAL CONSIDERATION

Non-Final OA §101§103§112
Filed
Mar 21, 2023
Priority
Mar 22, 2022 — RE 10-2022-0035448
Examiner
FEITL, LEAH M
Art Unit
2147
Tech Center
2100 — Computer Architecture & Software
Assignee
Samsung Electronics Co., Ltd.
OA Round
3 (Non-Final)
23%
Grant Probability
At Risk
3-4
OA Rounds
9m
Est. Remaining
28%
With Interview

Examiner Intelligence

Grants only 23% of cases
23%
Career Allowance Rate
21 granted / 93 resolved
-32.4% vs TC avg
Moderate +6% lift
Without
With
+5.8%
Interview Lift
resolved cases with interview
Typical timeline
4y 3m
Avg Prosecution
24 currently pending
Career history
127
Total Applications
across all art units

Statute-Specific Performance

§101
29.9%
-10.1% vs TC avg
§103
47.6%
+7.6% vs TC avg
§102
7.0%
-33.0% vs TC avg
§112
13.9%
-26.1% vs TC avg
Black line = Tech Center average estimate • Based on career data from 93 resolved cases

Office Action

§101 §103 §112
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Continued Examination Under 37 CFR 1.114 A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 06/23/2026 has been entered. Status of Claims This action is in response to the amendments filed 06/23/2026. Claims 1, 10, and 18 have been amended, claim 20 has been added. Claims 1-20 are currently pending. Response to Arguments Applicant’s arguments regarding the 101 rejection have been fully considered but they are not persuasive. Applicant argues that the limitations directed to generating differential data of activation data in a forward propagation without storing activation data of intermediate neural network layers and obtaining differential data without a separate backward propagation pass do not merely recite judicial exceptions are integrated into a practical application directed to an improvement to the functioning of a computer. Applicant argues that the claimed method is directed to a specific architecture by the elimination of the backward pass, which improves the performance of inference and produces a reduces memory size. Examiner respectfully notes that even if the claimed method does not perform a backpropagation step, that the broadest reasonable interpretation of the claimed forward propagation step in which differential data is generated still includes a mathematical calculation. Applicant’s specification repeatedly describes these steps as calculations, particularly paragraphs [0058]-[0060], [0082]-[0088], and [0091]-[0094]. The reduced memory size and argued improvements to inference speed are merely the byproduct of these mathematical calculations. MPEP 2106.05(a) clearly states that a judicial exception cannot form the sole basis for an improvement. Applicant has not clearly shown the additional elements in the claim which provide the basis for the improvement; therefore, this argument is unpersuasive. Examiner also notes that the additional elements as shown in the rejection of claim 1 do not integrate the claimed judicial exceptions into an abstract idea or amount to significantly more than the claimed judicial exceptions. With regards to newly added claim 20, Examiner notes that the broadest reasonable interpretation of concatenating input activation data and differential data into a single matrix, applying a weight and a bias to the matrix, and applying an activation function to the matrix and applying a derivate to the activation function to obtain differential data all include additional mathematical calculations, particularly in light of paragraphs [0095]-[0096] of Applicant’s specification, rather than particular technical computer implemented operations. The 101 rejections have been updated to include the amended limitations and to clarify the reasoning given for the limitations that were not amended. Applicant’s arguments regarding the prior art rejection have been fully considered but are moot because of the new ground(s) of rejection. Applicant argues that the Bai reference does not anticipate each and every limitation of the claims. Examiner notes that given further consideration in light of Applicant’s amendments, the 102 rejection has been withdrawn but that the claims are now rejected under 35 U.S.C. 103 as being unpatentable over Bai in view of the Chen reference. The prior art rejections have been updated to include the amended limitations and to clarify the reasoning given for the limitations that were not amended. Claim Rejections - 35 USC § 112 The following is a quotation of the first paragraph of 35 U.S.C. 112(a): (a) IN GENERAL.—The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor or joint inventor of carrying out the invention. The following is a quotation of the first paragraph of pre-AIA 35 U.S.C. 112: The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor of carrying out his invention. Claim 20 is rejected under 35 U.S.C. 112(a) or 35 U.S.C. 112 (pre-AIA ), first paragraph, as failing to comply with the written description requirement. The claim(s) contains subject matter which was not described in the specification in such a way as to reasonably convey to one skilled in the relevant art that the inventor or a joint inventor, or for applications subject to pre-AIA 35 U.S.C. 112, the inventor(s), at the time the application was filed, had possession of the claimed invention. Claim 20 recites the limitations “applying a weight of the corresponding layer and an augmented bias to the single matrix in a single operation to obtain a combined result. . . separating the combined result into a first portion and a second portion; and applying an activation function of the corresponding layer to the first portion to obtain the activation data of the corresponding layer, and applying a derivative of the activation function to the second portion to obtain the differential data of the activation data of the corresponding layer”. Applicant has merely pointed to paragraphs [0095]-[0096] to support these amended limitations without clearly explaining how these paragraphs teach an “augmented bias”, “combined results”, “first portion”, or “second portion”. Examiner notes that while paragraph [0095] of Applicant’s specification describes concatenating differential input data and a Jacobian matrix in the equation “xi−1′ = concat(xi−1, J(xi−1)( x ~ 0 )) and b′=(b, 0, 0 . . . 0), y′ = concat(y, J(y)( x ~ 0 )) = W×xi−1′+b′ ”, Applicant has not clearly shown the support for an “augmented bias”, “combined result”, or “separating the combined result into a first portion and second portion” such that one of ordinary skill in the art would understand how to apply an activation function to the first portion and a derivative of the activation to the second function to obtain differential data of the activation data. For purposes of examination, Examiner is interpreting that input activation data and differential data can be concatenated into a single matrix, an activation function can be applied to layer data, and a derivative of the activation function can be applied to obtain differential data. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-20 are rejected under 35 U.S.C. 101. Claims 1-8 and 20 are directed to a method, claim 9 is directed to a non-transitory computer readable medium, claims 10-17 are directed to a system, and claims 18-19 are directed to a separate method; therefore, claims 1-19 fall within one of the four statutory categories (i.e., process, machine, manufacture, or composition of matter). However, claims 1-19 fall within the judicial exception of an abstract idea, specifically the abstract ideas of “Mental Processes” (including observation, evaluation, and opinion) and “Mathematical Concepts (including mathematical calculations and relationships)”. Claim 1: Step 1: Claim 1 is directed to a method; therefore, the claim does fall within one of the four statutory categories (i.e., process, machine, manufacture, or composition of matter). Step 2A Prong 1: Claim 1 recites the following abstract ideas: generate differential data of the activation data of the corresponding layer with respect to input data by performing forward propagation of the corresponding layer without performing backpropagation, such that the differential data of the activation data is generated together with the activation data in the forward propagation (Examiner notes that the broadest reasonable interpretation of “generating” differential data by performing forward propagation to generate differential data together with activation data without performing backpropagation includes a mathematical calculation as supported by at least paragraph [0092] of Applicant’s specification); and generate differential data of output data of the neural network with respect to the input data, based on the generated differential data of each layer, during the forward propagation of each layer of the neural network, and without storing activation data of intermediate layers among the plurality of layers (Examiner notes that the broadest reasonable interpretation of “generating” differential data by performing forward propagation without storing activation data of intermediate layers includes a mathematical calculation as supported by at least paragraph [0093] of Applicant’s specification); wherein a memory size for the generating of the differential data is determined based on a number of elements of differential input data and a maximum value of dimension of an output of performing forward propagation for each of the plurality of layers such that memory usage is reduced and execution speed is enhanced (mental step directed to observation, evaluation – a person could determine a required or optimal memory size for inference of a neural network based on observed or determined elements of differential input data and an observed or determined maximum value of dimensions of an output from a forward propagation calculation. Examiner notes that the broadest reasonable interpretation of determining this memory size includes a mathematical calculation as supported by at least paragraph [0096] of Applicant’s specification. Examiner also notes that “such that memory usage is reduced and execution speed is enhanced” is interpreted as the intended use or necessary outcome of determining the size of the memory and does not provide additional patentable weight to this limitation (see MPEP 2103)). Claim 1 recites the following additional elements: for each layer of a plurality of layers of a neural network for an input data provided to the neural network: obtain activation data of a corresponding layer of the plurality of layers, resulting from an inference operation of the corresponding layer, and wherein the differential data of the output data is obtained upon completion of the forward propagation of the neural network without a separate backward propagation pass. Step 2A Prong 2: Obtaining activation data resulting from an inference operation for each layer of a plurality of layers and obtaining differential output data without a separate backward propagation pass are each interpreted as insignificant extra-solution activity directed to mere data gathering in the technological environment in which the claimed abstract ideas are performed. These additional elements, when considered with the aforementioned claimed abstract ideas, do not integrate those claimed abstract idea into a practical application or amount to significantly more than the abstract idea (see MPEP 2106.05(g) and MPEP 2106.05(h)). Step 2B: Obtaining activation data resulting from an inference operation for each layer of a plurality of layers and obtaining differential output data without a separate backward propagation pass are each interpreted as well-understood, routine, conventional activity directed to receiving data over a network. These additional elements, when considered with the aforementioned claimed abstract ideas, do not amount to significantly more than the claimed abstract ideas (see MPEP 2106.05(d)(II)). Claim 9 is a non-transitory computer readable medium claim which performs the inference method of claim 1. The only difference is that claim 9 requires a non-transitory computer readable medium, which is interpreted as a generic computer component merely used to apply, or perform, the abstract ideas as identified in the analysis of claim 1 (see MPEP 2106.05(f)). Therefore, claim 9 is rejected for the same reasons as claim 1. Claim 10 is a system claim and its limitation is included in claim 1. The only difference is that claim 10 requires a system, which is interpreted as a generic computer component merely used to apply the abstract ideas as identified in the analysis of claim 1 (see MPEP 2106.05(f)). Therefore, claim 10 is rejected for the same reasons as claim 1. Claim 18: Step 1: Claim 18 is directed to a method; therefore, the claim does fall within one of the four statutory categories (i.e., process, machine, manufacture, or composition of matter). Step 2A Prong 1: Claim 18 recites the following abstract ideas: generating differential data of output data of a neural network based on respective differential data of each layer of a plurality of layers of the neural network, generated during corresponding forward propagation operations of the neural network without performing backpropagation, the respective differential data of each layer being generated together with activation data of the layer in the forward propagation operations and without storing activation data of intermediate layers of the neural network (mental step directed to observation, evaluation – a person could generate differential output data in their mind based on observed or determined differential data generated with activation data during forward propagation operations for a neural network without performing a backpropagation calculation or storing activation data of intermediate layers. Examiner notes that the broadest reasonable interpretation of “generating” differential data by performing forward propagation also includes the mathematical calculations associated with forward propagation operations as supported by at least paragraphs [0092]-[0095] of Applicant’s specification); wherein a memory size for the generating of the differential data is determined based on a number of elements of differential input data and a maximum value of dimension of an output of performing forward propagation for each of the plurality of layers such that memory usage is reduced and execution speed is enhanced (mental step directed to observation, evaluation – a person could determine a required or optimal memory size for inference of a neural network based on observed or determined elements of differential input data and an observed or determined maximum value of dimensions of an output from a forward propagation calculation. Examiner notes that the broadest reasonable interpretation of determining this memory size includes a mathematical calculation as supported by at least paragraph [0096] of Applicant’s specification. Examiner also notes that “such that memory usage is reduced and execution speed is enhanced” is interpreted as the intended use or necessary outcome of determining the size of the memory and does not provide additional patentable weight to this limitation (see MPEP 2103)). Claim 18 recites the following additional elements: wherein the differential data of output data is obtained based on a Jacobian matrix for input data of a layer of the plurality of layers. Obtaining differential output data based on a Jacobian matrix for input data of a given layer is interpreted as interpreted as insignificant extra-solution activity directed to mere data gathering. This additional element, when considered with the aforementioned claimed abstract ideas, does not integrate those claimed abstract idea into a practical application or amount to significantly more than the abstract idea (see MPEP 2106.05(g) and MPEP 2106.05(h)). Step 2B: Obtaining differential output data based on a Jacobian matrix for input data of a given layer is interpreted as interpreted as well-understood, routine, conventional activity directed to receiving data over a network. This additional element, when considered with the aforementioned claimed abstract ideas, does not amount to significantly more than the claimed abstract ideas (see MPEP 2106.05(d)(II)). The independent claims are not patent eligible. Dependent claims 2-8, 11-17, and 19-20 when analyzed as a whole are held to be patent ineligible under 35 U.S.C. 101 because the additional recited limitations fail to establish that the claims are not directed to an abstract idea, as they recite further embellishment of the judicial exception. Claim 2 recites wherein the generating of the differential data comprises: for each layer of the layers, calculating a Jacobian matrix with respect to the input data (mental step directed to observation, evaluation – a person could generate differential data in their mind by calculating a Jacobian matrix based on observed input data for each layer. Examiner notes that the broadest reasonable interpretation of “generating” differential data and calculating a Jacobian matrix also includes the mathematical calculations associated with forward propagation operations as supported by at least paragraphs [0092]-[0095] of Applicant’s specification and the mathematical calculations associated with the Jacobian matrix as supported by at least [0098]-[0099] of Applicant’s specification). Claim 3 recites wherein the generating of the differential data comprises: calculating a Jacobian matrix of the corresponding layer with respect to the input data by performing the inference operation of the corresponding layer (mental step directed to observation, evaluation – a person could generate differential data in their mind by calculating a Jacobian matrix as part of an inference operation based on observed input data. Examiner notes that the broadest reasonable interpretation of “generating” differential data and calculating a Jacobian matrix also includes the mathematical calculations associated with forward propagation operations as supported by at least paragraphs [0092]-[0095] of Applicant’s specification and the mathematical calculations associated with the Jacobian matrix as supported by at least [0098]-[0099] of Applicant’s specification). Claim 4 recites for each layer, calculating a Jacobian matrix of the corresponding layer with respect to the input data without performing backpropagation (mental step directed to observation, evaluation – a person could generate differential output data in their mind by calculating a Jacobian matrix based on observed input data without performing backpropagation operations. Examiner notes that the broadest reasonable interpretation of “generating” differential data and calculating a Jacobian matrix also includes the mathematical calculations associated with forward propagation operations as supported by at least paragraphs [0092]-[0095] of Applicant’s specification and the mathematical calculations associated with the Jacobian matrix as supported by at least [0098]-[0099] of Applicant’s specification). Claim 5 recites for each layer, performing the inference operation of the corresponding layer to generate the activation data of the corresponding layer; and generating output data of the neural network based on the generated activation data of each of the layers (mental step directed to observation, evaluation – a person could perform inference operations to generate activation data corresponding to a layer of a neural network in their mind and generate output data in their mind based on the generated activation data. Examiner notes that the broadest reasonable interpretation of “generating” differential data also includes the mathematical calculations associated with forward propagation, or inference operations as supported by at least paragraphs [0092]-[0095] of Applicant’s specification). Claim 6 recites generating differential input data comprising one or more elements for a differential value among a plurality of elements of the input data (mental step directed to observation, evaluation – a person could generate differential input data comprising elements for differential values of observed input data in their mind. Examiner notes that the broadest reasonable interpretation of “generating” differential data also includes the mathematical calculations associated with forward propagation operations as supported by at least paragraphs [0092]-[0095] of Applicant’s specification). Claim 7 recites wherein the generating of the differential data comprises: for each layer, calculating a Jacobian matrix of the corresponding layer with respect to the differential input data (mental step directed to observation, evaluation – a person could generate differential data in their mind by calculating a Jacobian matrix based on observed differential input data. Examiner notes that the broadest reasonable interpretation of “generating” differential data and calculating a Jacobian matrix also includes the mathematical calculations associated with forward propagation operations as supported by at least paragraphs [0092]-[0095] of Applicant’s specification and the mathematical calculations associated with the Jacobian matrix as supported by at least [0098]-[0099] of Applicant’s specification). Claim 8 recites wherein a memory size for inference of the neural network is determined based on a number of elements of the differential input data and a maximum value of dimensions of each Jacobian matrix of the plurality of layers with respect to the differential input data (mental step directed to observation, evaluation – a person could determine a required or optimal memory size for inference of a neural network based on observed or determined elements of differential input data and an observed or determined maximum value of dimensions of a Jacobian matrix). Claim 11 is a system claim and its limitation is included in claim 2. Claim 11 is rejected for the same reasons as claim 2. Claim 12 is a system claim and its limitation is included in claim 3. Claim 12 is rejected for the same reasons as claim 3. Claim 13 is a system claim and its limitation is included in claim 4. Claim 13 is rejected for the same reasons as claim 4. Claim 14 is a system claim and its limitation is included in claim 5. Claim 14 is rejected for the same reasons as claim 5. Claim 15 is a system claim and its limitation is included in claim 6. Claim 15 is rejected for the same reasons as claim 6. Claim 16 is a system claim and its limitation is included in claim 7. Claim 16 is rejected for the same reasons as claim 7. Claim 17 is a system claim and its limitation is included in claim 8. Claim 17 is rejected for the same reasons as claim 8. Claim 19 recites wherein the differential data of the output data of the neural network is obtained with respect to the input data, based on differential data of an output activation of a corresponding layer with respect to the input data (obtaining output data of a neural network based on differential activation output data corresponding to input data of a given layer is interpreted as well-understood, routine, conventional activity directed to receiving data over a network and as aspects of the technological environment in which the claimed abstract ideas are performed. These additional elements do not integrate the abstract idea into a practical application or amount to significantly more than the abstract idea (see MPEP 2106.05(d)(II) and MPEP 2106.05(h)). Claim 20 recites wherein the generating of the differential data of the activation data comprises: concatenating input activation data of the corresponding layer and differential data of the input activation data into a single matrix; applying a weight of the corresponding layer and an augmented bias to the single matrix in a single operation to obtain a combined result, the augmented bias comprising a bias of the corresponding layer for a portion corresponding to the input activation data and zero values for a portion corresponding to the differential data of the input activation data; separating the combined result into a first portion and a second portion; and applying an activation function of the corresponding layer to the first portion to obtain the activation data of the corresponding layer, and applying a derivative of the activation function to the second portion to obtain the differential data of the activation data of the corresponding layer. Claim 20 is interpreted in light of the 112(a) rejection such that input activation data and differential data can be concatenated into a single matrix, an activation function can be applied to layer data, and a derivative of the activation function can be applied to obtain differential data. Given this interpretation, concatenating input activation data and differential data into a single matrix, applying an activation function to layer data, and applying a derivative of the activation function to obtain differential data are each interpreted as mathematical calculations in light of paragraphs [0095]-[0096] of Applicant’s specification. Viewed as a whole, these additional claim elements do not provide meaningful limitations to transform the abstract idea into a patent eligible application of the abstract idea such that the claims amount to significantly more than the abstract idea itself. Therefore, the claims are rejected under 35 U.S.C. 101 as being directed to non-statutory subject matter. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-20 are rejected under 35 U.S.C. 103 as being unpatentable over Bai et al** (US 20210042606 A1, herein Bai), in view of Chen et al* (“Neural Ordinary Differential Equations”, herein Chen). *this document was included in the IDS dated 03/21/2023 **this document was included in the IDS dated 08/11/2023 Regarding claim 1, Bai teaches a processor-implemented method for inference-based differential consideration (para. [0008] recites “In accordance with a first aspect of the invention, a computer-implemented method and corresponding system are provided for training a neural network”), comprising: for each layer of a plurality of layers of a neural network for an input data provided to the neural network: obtain activation data of a corresponding layer of the plurality of layers, resulting from an inference operation of the corresponding layer (para. [0009] recites “providing a neural network which comprises an iterative function (z[i+ 1] = f(z[i], θ, c(x)). Such an iterative function is known in the field of machine learning to be representable by a stack of layers which have mutually shared weights. In such a stack of layers, each layer except for the first layer may receive, as input, i) an output of the previous layer and ii) (a part of) an input to the stack of layers, being either the original input (x) to the neural network or a transformation of that input (c(x)), for example by one or more previous layers preceding the stack of layers in the neural network. The first layer of the stack of layers may receive an initial activation as input, which may for example be an output of yet another layer of the neural network” (i.e., obtaining activation data for a plurality of layers during a forward pass, or inference operation)); generate differential data of the activation data of the corresponding layer with respect to input data by performing forward propagation of the corresponding layer without performing backpropagation, such that the differential data of the activation data is generated together with the activation data in the forward propagation (para. [0009] recites “In such a stack of layers, each layer except for the first layer may receive, as input, i) an output of the previous layer and ii) (a part of) an input to the stack of layers, being either the original input (x) to the neural network or a transformation of that input (c(x)), for example by one or more previous layers preceding the stack of layers in the neural network. The first layer of the stack of layers may receive an initial activation as input, which may for example be an output of yet another layer of the neural network” (i.e., generating activation data for a given layer with respect to the input data). Para. [0022] recites “backpropagation of the backward gradient through the stack of layers may be replaced by solving the above linear system which may involve using one step of matrix multiplications that involves the Jacobian at equilibrium. Herein, the vector Jacobian product may be efficiently computed via automatic differentiation tools for any x, without having to explicitly write out the Jacobian matrix” (i.e., calculating the Jacobian vector product, which one of ordinary skill in the art would recognize as the projection of the Jacobian matrix into a vector during the forward pass, for a given layer instead of performing backpropagation. Examiner notes that the broadest reasonable interpretation of “differential” data includes an interpretation data that can be input to a differentiation technique, or “differentiable” data)), wherein a memory size for the generating of the differential data is determined based on a number of elements of differential input data and a maximum value of dimension of an output of performing forward propagation for each of the plurality of layers such that memory usage is reduced and execution speed is enhanced (Bai para. [0003] recites “the training itself of a deep neural network requires a large amount of memory, since in addition to the weights per layer, also a large amount of temporary data has to be stored for the forward passes ('forward propagation') and backward passes ('backward propagation') during the training. For example, the layer output of each individual layer ('hidden state') during forward propagation may need to be stored as temporary data as it may be used in the backward propagation. This way, the training of a deep neural network may require many gigabytes of memory, with the memory requirements being expected to further increase as the complexity of models increases”. Bai para. [0011] recites “As described in the background section, during training, the iterative execution of a stack of layers still requires a sizable memory footprint, since the layer output of each individual layer (even if weight-tied) during forward propagation may need to be stored as temporary data as it may be used in the subsequent backward propagation” (i.e., the amount of memory required is determined using at least the size of the data generated during the forward passes, which includes the input data described in at least para. [0009] and the output of the corresponding Jacobian matrix for each layer as described in at least para. [0022]. Examiner notes that “such that memory usage is reduced and execution speed is enhanced” is interpreted as the intended use or necessary outcome of determining the size of the memory and does not provide additional patentable weight to this limitation (see MPEP 2103))). However, while Bai teaches that backpropagation operations can be replaced and that intermediate layer outputs do not need to be stored in memory (see at least paragraphs [0020]-[0022]), Bai does not explicitly teach generating differential data of output data of the neural network with respect to the input data, based on the generated differential data of each layer during the forward propagation of each layer of the neural network and without storing activation data of intermediate layers among the plurality of layers, wherein the differential data of the output data is obtained upon completion of the forward propagation of the neural network without a separate backward propagation pass. Chen teaches generating differential data of output data of the neural network with respect to the input data, based on the generated differential data of each layer during the forward propagation of each layer of the neural network and without storing activation data of intermediate layers among the plurality of layers, wherein the differential data of the output data is obtained upon completion of the forward propagation of the neural network without a separate backward propagation pass (section I para. 3 recites “we show how to compute gradients of a scalar-valued loss with respect to all inputs of any ODE solver, without backpropagating through the operations of the solver. Not storing any intermediate quantities of the forward pass allows us to train our models with constant memory cost as a function of depth, a major bottleneck of training deep models” (i.e., completing a forward pass without a backward pass and without storing intermediate layer data, which would include intermediate layer activation data, in memory). Algorithm 1 and section 2 para. 3-7 recite “Consider optimizing a scalar-valued loss function L(), whose input is the result of an ODE solver: (EQ3)To optimize L, we require gradients with respect to ϴ. The first step is to determining how the gradient of the loss depends on the hidden state z(t) at each instant. This quantity is called the adjoint a(t) = ∂ L / ∂ z ( t ) . Its dynamics are given by another ODE, which can be thought of as the instantaneous analog of the chain rule: (EQ4). We can compute ∂ L / ∂ z ( t 1 ) by another call to an ODE solver. This solver must run backwards, starting from the initial value of ∂ L / ∂ z ( t 1 ) . One complication is that solving this ODE requires the knowing value of z(t) along its entire trajectory. However, we can simply recompute z(t) backwards in time together with the adjoint, starting from its final value z(t1). Computing the gradients with respect to the parameters ϴ requires evaluating a third integral, which depends on both z(t) and a(t): (EQ5). The vector-Jacobian products a(t)T ∂ f / ∂ z and a(t)T ∂ f / ∂ ϴ in (EQ4) and (EQ5) can be efficiently evaluated by automatic differentiation, at a time cost similar to that of evaluating f” (i.e., generating differential output data)). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine these teachings by substituting the linear equation solver for the Jacobian matrix from Bai with the ordinary differential equation solver from Chen. Bai and Chen are both directed to methods of improving the computation time and memory usage of neural network computations; accordingly, one of ordinary skill in the art would recognize that the method of calculating the Jacobian matrix from Bai could be solved with the automatic differentiation method from Chen to yield a predictable result, as Bai states in paragraph [0022], “the vector-Jacobian product may be efficiently computed via automatic differential tools”. Regarding claim 2, the combination of Bai and Chen teaches the method of claim 1, wherein the generating of the differential data comprises: for each layer of the layers, calculating a Jacobian matrix with respect to the input data (Bai para. [0017]-[0020] recite “performing the backward propagation part comprises: computing respective partial derivatives of the iterative function with respect to the weights and the part of the input; computing a gradient of the input of the iterative execution layer for the weights and/or the part of the input as a function of the respective partial derivative and a backpropagated gradient at the output of the iterative execution layer. The derivatives which are indicated above may be implemented via their analytic equations or computed, e.g., implemented via their automatic differentiation tools”. Bai para. [0022] recites “backpropagation of the backward gradient through the stack of layers may be replaced by solving the above linear system which may involve using one step of matrix multiplications that involves the Jacobian at equilibrium” (i.e., differential output data obtained based on a Jacobian matrix calculated from the input data of a given layer)). Regarding claim 3, the combination of Bai and Chen teaches the method of claim 1, wherein the generating of the differential data comprises: calculating a Jacobian matrix of the corresponding layer with respect to the input data by performing the inference operation of the corresponding layer ( Bai para. [0009] recites “providing a neural network which comprises an iterative function (z[i+ 1] = f(z[i], θ, c(x)). Such an iterative function is known in the field of machine learning to be representable by a stack of layers which have mutually shared weights. In such a stack of layers, each layer except for the first layer may receive, as input, i) an output of the previous layer and ii) (a part of) an input to the stack of layers, being either the original input (x) to the neural network or a transformation of that input (c(x)), for example by one or more previous layers preceding the stack of layers in the neural network. Bai para. [0017]-[0020] recite “performing the backward propagation part comprises: computing respective partial derivatives of the iterative function with respect to the weights and the part of the input; computing a gradient of the input of the iterative execution layer for the weights and/or the part of the input as a function of the respective partial derivative and a backpropagated gradient at the output of the iterative execution layer. The derivatives which are indicated above may be implemented via their analytic equations or computed, e.g., implemented via their automatic differentiation tools”. Bai para. [0022] recites “backpropagation of the backward gradient through the stack of layers may be replaced by solving the above linear system which may involve using one step of matrix multiplications that involves the Jacobian at equilibrium” (i.e., differential output data obtained based on a Jacobian matrix calculated from the input data of a given layer obtained during the inference, or forward pass)). Regarding claim 4, the combination of Bai and Chen teaches the method of claim 1, wherein the generating of the differential data comprises: for each layer, calculating a Jacobian matrix of the corresponding layer with respect to the input data without performing backpropagation (Bai para. [0022] recites “backpropagation of the backward gradient through the stack of layers may be replaced by solving the above linear system which may involve using one step of matrix multiplications that involves the Jacobian at equilibrium. Herein, the vector Jacobian product may be efficiently computed via automatic differentiation tools for any x, without having to explicitly write out the Jacobian matrix” (i.e., calculating the Jacobian vector product, which one of ordinary skill in the art would recognize as the projection of the Jacobian matrix into a vector, for a given layer instead of performing backpropagation)). Regarding claim 5, the combination of Bai and Chen teaches the method of claim 1, further comprising: for each layer, performing the inference operation of the corresponding layer to generate the activation data of the corresponding layer (Bai para. [0009] recites “providing a neural network which comprises an iterative function (z[i+ 1] = f(z[i], θ, c(x)). Such an iterative function is known in the field of machine learning to be representable by a stack of layers which have mutually shared weights. In such a stack of layers, each layer except for the first layer may receive, as input, i) an output of the previous layer and ii) (a part of) an input to the stack of layers, being either the original input (x) to the neural network or a transformation of that input (c(x)), for example by one or more previous layers preceding the stack of layers in the neural network. The first layer of the stack of layers may receive an initial activation as input, which may for example be an output of yet another layer of the neural network” (i.e., obtaining activation data for a plurality of layers during a forward pass, or inference operation)); and generating output data of the neural network based on the generated activation data of each of the layers (Bai para. [0022] recites “backpropagation of the backward gradient through the stack of layers may be replaced by solving the above linear system which may involve using one step of matrix multiplications that involves the Jacobian at equilibrium”. Bai para. [0024] recites “outputting the trained neural network comprises representing the stack of layers in the trained neural network by at least i) a data representation of a layer of the stack of layers, and ii) a hyperparameter defining a number of layers of the stack of layers at which the output of the stack of layers reaches or to a selected degree approximates the equilibrium point during forward propagation” (i.e., generating output data of the neural network based on the stack of layers, which includes the activation data as described in at least para. [0009])). Regarding claim 6, the combination of Bai and Chen teaches the method of claim 1, further comprising: generating differential input data comprising one or more elements for a differential value among a plurality of elements of the input data (Bai para. [0009] recites “providing a neural network which comprises an iterative function (z[i+ 1] = f(z[i], θ, c(x)). Such an iterative function is known in the field of machine learning to be representable by a stack of layers which have mutually shared weights. In such a stack of layers, each layer except for the first layer may receive, as input, i) an output of the previous layer and ii) (a part of) an input to the stack of layers, being either the original input (x) to the neural network or a transformation of that input (c(x)), for example by one or more previous layers preceding the stack of layers in the neural network. The first layer of the stack of layers may receive an initial activation as input, which may for example be an output of yet another layer of the neural network” (i.e., generating input data comprising one or more elements with values capable of being differentiable as described in at least para. [0017]-[0020])). Regarding claim 7, the combination of Bai and Chen teaches the method of claim 6, wherein the generating of the differential data comprises: for each layer, calculating a Jacobian matrix of the corresponding layer with respect to the differential input data (Bai para. [0017]-[0020] recite “performing the backward propagation part comprises: computing respective partial derivatives of the iterative function with respect to the weights and the part of the input; computing a gradient of the input of the iterative execution layer for the weights and/or the part of the input as a function of the respective partial derivative and a backpropagated gradient at the output of the iterative execution layer. The derivatives which are indicated above may be implemented via their analytic equations or computed, e.g., implemented via their automatic differentiation tools”. Bai para. [0022] recites “backpropagation of the backward gradient through the stack of layers may be replaced by solving the above linear system which may involve using one step of matrix multiplications that involves the Jacobian at equilibrium” (i.e., differential output data obtained based on a Jacobian matrix calculated from the differentiable input data)). Regarding claim 8, the combination of Bai and Chen teaches the method of claim 7, wherein a memory size for inference of the neural network is determined based on a number of elements of the differential input data and a maximum value of dimensions of each Jacobian matrix of the plurality of layers with respect to the differential input data (Bai para. [0003] recites “the training itself of a deep neural network requires a large amount of memory, since in addition to the weights per layer, also a large amount of temporary data has to be stored for the forward passes ('forward propagation') and backward passes ('backward propagation') during the training. For example, the layer output of each individual layer ('hidden state') during forward propagation may need to be stored as temporary data as it may be used in the backward propagation. This way, the training of a deep neural network may require many gigabytes of memory, with the memory requirements being expected to further increase as the complexity of models increases”. Bai para. [0011] recites “As described in the background section, during training, the iterative execution of a stack of layers still requires a sizable memory footprint, since the layer output of each individual layer (even if weight-tied) during forward propagation may need to be stored as temporary data as it may be used in the subsequent backward propagation” (i.e., the amount of memory required is determined using the size of the data generated during the forward and backward passes, which includes the input data described in at least para. [0009] and the corresponding Jacobian matrix for each layer as described in at least para. [0022])). Claim 9 is a non-transitory computer readable medium claim which performs the inference method of claim 1. The only difference is that claim 9 requires a non-transitory computer readable medium (Bai para. [0008] recites “a computer-readable medium is provided comprising transitory or non-transitory data representing model data defining a trained neural network”). Therefore, claim 9 is rejected for the same reasons as claim 1. Claim 10 is a system claim and its limitation is included in claim 1. The only difference is that claim 10 requires a system (Bai para. [0008] recites “In accordance with a first aspect of the invention, a computer-implemented method and corresponding system are provided for training a neural network”). Therefore, claim 10 is rejected for the same reasons as claim 1. Claim 11 is a system claim and its limitation is included in claim 2. Claim 11 is rejected for the same reasons as claim 2. Claim 12 is a system claim and its limitation is included in claim 3. Claim 12 is rejected for the same reasons as claim 3. Claim 13 is a system claim and its limitation is included in claim 4. Claim 13 is rejected for the same reasons as claim 4. Claim 14 is a system claim and its limitation is included in claim 5. Claim 14 is rejected for the same reasons as claim 5. Claim 15 is a system claim and its limitation is included in claim 6. Claim 15 is rejected for the same reasons as claim 6. Claim 16 is a system claim and its limitation is included in claim 7. Claim 16 is rejected for the same reasons as claim 7. Claim 17 is a system claim and its limitation is included in claim 8. Claim 17 is rejected for the same reasons as claim 8. Regarding claim 18, Bai teaches a processor-implemented method (para. [0008] recites “In accordance with a first aspect of the invention, a computer-implemented method and corresponding system are provided for training a neural network”), comprising: wherein the differential data of output data is obtained based on a Jacobian matrix for input data of a layer of the plurality of layers (para. [0017]-[0020] recite “performing the backward propagation part comprises: computing respective partial derivatives of the iterative function with respect to the weights and the part of the input; computing a gradient of the input of the iterative execution layer for the weights and/or the part of the input as a function of the respective partial derivative and a backpropagated gradient at the output of the iterative execution layer. The derivatives which are indicated above may be implemented via their analytic equations or computed, e.g., implemented via their automatic differentiation tools”. Para. [0022] recites “backpropagation of the backward gradient through the stack of layers may be replaced by solving the above linear system which may involve using one step of matrix multiplications that involves the Jacobian at equilibrium” (i.e., differential output data obtained based on a Jacobian matrix)); and wherein a memory size for the generating of the differential data is determined based on a number of elements of differential input data and a maximum value of dimension of an output of performing forward propagation for each of the plurality of layers such that memory usage is reduced and execution speed is enhanced (Bai para. [0003] recites “the training itself of a deep neural network requires a large amount of memory, since in addition to the weights per layer, also a large amount of temporary data has to be stored for the forward passes ('forward propagation') and backward passes ('backward propagation') during the training. For example, the layer output of each individual layer ('hidden state') during forward propagation may need to be stored as temporary data as it may be used in the backward propagation. This way, the training of a deep neural network may require many gigabytes of memory, with the memory requirements being expected to further increase as the complexity of models increases”. Bai para. [0011] recites “As described in the background section, during training, the iterative execution of a stack of layers still requires a sizable memory footprint, since the layer output of each individual layer (even if weight-tied) during forward propagation may need to be stored as temporary data as it may be used in the subsequent backward propagation” (i.e., the amount of memory required is determined using at least the size of the data generated during the forward passes, which includes the input data described in at least para. [0009] and the output of the corresponding Jacobian matrix for each layer as described in at least para. [0022]. Examiner notes that “such that memory usage is reduced and execution speed is enhanced” is interpreted as the intended use or necessary outcome of determining the size of the memory and does not provide additional patentable weight to this limitation (see MPEP 2103))). However, while Bai teaches backpropagation operations can be replaced and that intermediate layer outputs do not need to be stored in memory (see at least paragraphs [0020]-[0022]), Bai does not explicitly teach generating differential data of output data of a neural network based on respective differential data of each layer of a plurality of layers of the neural network generated during corresponding forward propagation operations of the neural network without performing backpropagation and without storing activation data of intermediate layers of the neural network. Chen teaches generating differential data of output data of a neural network based on respective differential data of each layer of a plurality of layers of the neural network generated during corresponding forward propagation operations of the neural network without performing backpropagation and without storing activation data of intermediate layers of the neural network (section I para. 3 recites “we show how to compute gradients of a scalar-valued loss with respect to all inputs of any ODE solver, without backpropagating through the operations of the solver. Not storing any intermediate quantities of the forward pass allows us to train our models with constant memory cost as a function of depth, a major bottleneck of training deep models” (i.e., completing a forward pass without a backward pass and without storing intermediate layer data, which would include intermediate layer activation data, in memory). Algorithm 1 and section 2 para. 3-7 recite “Consider optimizing a scalar-valued loss function L(), whose input is the result of an ODE solver: (EQ3)To optimize L, we require gradients with respect to ϴ. The first step is to determining how the gradient of the loss depends on the hidden state z(t) at each instant. This quantity is called the adjoint a(t) = ∂ L / ∂ z ( t ) . Its dynamics are given by another ODE, which can be thought of as the instantaneous analog of the chain rule: (EQ4). We can compute ∂ L / ∂ z ( t 1 ) by another call to an ODE solver. This solver must run backwards, starting from the initial value of ∂ L / ∂ z ( t 1 ) . One complication is that solving this ODE requires the knowing value of z(t) along its entire trajectory. However, we can simply recompute z(t) backwards in time together with the adjoint, starting from its final value z(t1). Computing the gradients with respect to the parameters ϴ requires evaluating a third integral, which depends on both z(t) and a(t): (EQ5). The vector-Jacobian products a(t)T ∂ f / ∂ z and a(t)T ∂ f / ∂ ϴ in (EQ4) and (EQ5) can be efficiently evaluated by automatic differentiation, at a time cost similar to that of evaluating f” (i.e., generating differential output data)). See claim 1 for motivation to combine. Regarding claim 19, the combination of Bai and Chen teaches the method of claim 18, wherein the differential data of the output data of the neural network is obtained with respect to the input data, based on differential data of an output activation of a corresponding layer with respect to the input data (Bai para. [0022] recites “backpropagation of the backward gradient through the stack of layers may be replaced by solving the above linear system which may involve using one step of matrix multiplications that involves the Jacobian at equilibrium”. Para. [0024] recites “outputting the trained neural network comprises representing the stack of layers in the trained neural network by at least i) a data representation of a layer of the stack of layers, and ii) a hyperparameter defining a number of layers of the stack of layers at which the output of the stack of layers reaches or to a selected degree approximates the equilibrium point during forward propagation” (i.e., generating differential output data of the neural network based on the stack of layers, which includes the input data and the activation data as described in at least para. [0009])). Regarding claim 20, the combination of Bai and Chen teaches the method of claim 1 as mentioned above, wherein the generating of the differential data of the activation data comprises: concatenating input activation data of the corresponding layer and differential data of the input activation data into a single matrix; applying a weight of the corresponding layer and an augmented bias to the single matrix in a single operation to obtain a combined result, the augmented bias comprising a bias of the corresponding layer for a portion corresponding to the input activation data and zero values for a portion corresponding to the differential data of the input activation data; separating the combined result into a first portion and a second portion; and applying an activation function of the corresponding layer to the first portion to obtain the activation data of the corresponding layer, and applying a derivative of the activation function to the second portion to obtain the differential data of the activation data of the corresponding layer (Examiner notes that claim 20 is interpreted in light of the 112(a) rejection such that input activation data and differential data can be concatenated into a single matrix, an activation function can be applied to layer data, and a derivative of the activation function can be applied to obtain differential data. Given this interpretation, Chen teaches concatenating input activation data and differential data in at least Algorithm 1 and section 2 para. 7. Bai teaches applying an activation function to layer data in at least para. [0009]-[0011]. Chen teaches applying a derivative of an activation function to obtain differential data in at least Algorithm 1 and section 2 para. 3-7). Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. US 20220317332 A1 (Bowden, Jr. et al) teaches a method for constructing a neural ordinary differential equation network for physics-constrained modeling. US 20200233920 A1 (Meeds et al) teaches a method for modeling ordinary differential equations using a variational autoencoder. US 20200218999 A1 (Eleftheriadis et al) teaches a method for calculating a natural gradient descent to optimize a variational Gaussian process by utilizing forward-mode differentiation. Any inquiry concerning this communication or earlier communications from the examiner should be directed to LEAH M FEITL whose telephone number is (571) 272-8350. The examiner can normally be reached on M-F 0900-1700 EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Viker Lamardo can be reached on (571) 270-5871. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /L.M.F./ Examiner, Art Unit 2147 /VIKER A LAMARDO/Supervisory Patent Examiner, Art Unit 2147
Read full office action

Prosecution Timeline

Show 2 earlier events
Feb 19, 2026
Applicant Interview (Telephonic)
Feb 19, 2026
Examiner Interview Summary
Feb 27, 2026
Response Filed
Jun 05, 2026
Final Rejection mailed — §101, §103, §112
Jul 08, 2026
Response after Non-Final Action
Aug 07, 2026
Request for Continued Examination
Aug 10, 2026
Response after Non-Final Action
Sep 24, 2026
Non-Final Rejection mailed — §101, §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12705248
PREDICTION OF AN EVENT AFFECTING A PHYSICAL SYSTEM
7y 7m to grant Granted Aug 11, 2026
Patent 12682009
CONTENT TARGETING USING CONTENT CONTEXT AND USER PROPENSITY
5y 6m to grant Granted Jul 14, 2026
Patent 12670304
METHODS AND APPARATUSES FOR RESOURCE-OPTIMIZED FERMIONIC LOCAL SIMULATION ON QUANTUM COMPUTER FOR QUANTUM CHEMISTRY
3y 6m to grant Granted Jun 30, 2026
Patent 12619874
Stochastic Gradient Boosting For Deep Neural Networks
2y 2m to grant Granted May 05, 2026
Patent 12572720
METHODS AND APPARATUSES FOR RESOURCE-OPTIMIZED FERMIONIC LOCAL SIMULATION ON QUANTUM COMPUTER FOR QUANTUM CHEMISTRY
5y 0m to grant Granted Mar 10, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
23%
Grant Probability
28%
With Interview (+5.8%)
4y 3m (~9m remaining)
Median Time to Grant
High
PTA Risk
Based on 93 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month