Prosecution Insights
Last updated: October 02, 2026
Application No. 18/343,303

METHOD AND APPARATUS WITH NEURAL NETWORK TRAINING

Final Rejection §103§112
Filed
Jun 28, 2023
Priority
Jan 10, 2023 — RE 10-2023-0003628
Examiner
CHEN, KUANG FU
Art Unit
2143
Tech Center
2100 — Computer Architecture & Software
Assignee
Samsung Electronics Co., Ltd.
OA Round
2 (Final)
80%
Grant Probability
Favorable
3-4
OA Rounds
0m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 80% — above average
80%
Career Allowance Rate
216 granted / 271 resolved
+24.7% vs TC avg
Strong +69% interview lift
Without
With
+69.0%
Interview Lift
resolved cases with interview
Typical timeline
2y 11m
Avg Prosecution
30 currently pending
Career history
298
Total Applications
across all art units

Statute-Specific Performance

§101
16.4%
-23.6% vs TC avg
§103
50.8%
+10.8% vs TC avg
§102
10.9%
-29.1% vs TC avg
§112
15.5%
-24.5% vs TC avg
Black line = Tech Center average estimate • Based on career data from 271 resolved cases

Office Action

§103 §112
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Amendment The Amendment filed 6/17/2026 has been entered. Claims 1 and 9-10 have been amended. Claims 1-17 are pending in the application. Specification The disclosure is objected to because of the following informalities: [0040] designates element 100 as a computing device, while [0085] and [0089] designate the same element as an electronic device; [0078], [0085], [0086] and [0088] designate element 505 as a communication interface, while [0089] designates the same element as a communications interface. Appropriate correction is required. Claim Objections Claims 9, 10 and 13 are objected to because of the following informalities: Claim 9 recites "second neural network , wherein" which carries an extraneous space before the comma and should read "second neural network, wherein"; Claim 10 recites "the forward propagation of the first neural network," which should read "the forward propagation process of the first neural network" for consistency with the "forward propagation process of the first neural network" recited in claim 9; and Claim 13 recites "the forward propagation of the first neural network," which should read "the forward propagation process of the first neural network" for the same reason. Appropriate correction is required. Claim Rejections - 35 U.S.C. 112(b) The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 9-17 are rejected under 35 U.S.C. 112(b) or pre-AIA 35 U.S.C. 112, second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Claim 9 recites train the first neural network and the second neural network by updating the shared parameters based on the estimated output data and the output differential value without using the differential value of the output differential value. The negative limitation does not set forth the boundary of the subject matter it excludes. A differential value is a derivative taken with respect to a variable. The claim states that variable for the output differential value itself, which is recited as a differential value of the output data with respect to the input data, but it states no variable for the further differential value that it excludes. Two different quantities therefore answer to the excluded term on the face of the claim: a differential value of the output differential value taken with respect to the input data, and a differential value of the output differential value taken with respect to the parameters that the same limitation requires to be updated. The two are not the same quantity and they do not give the claim the same scope. The specification's own term second-order differential denotes the derivative taken with respect to a parameter, while the claim term itself states no variable of differentiation. [0041] states that the gradient of the second loss function is a second order term and identifies the second-order differential that may be needed to be calculated as the derivative of the output differential value taken with respect to a parameter of the first neural network. [0067] states that this second-order differential function may be effectively calculated through the calculation of that gradient, and that the neural network is trained without requiring the calculating of the Hessian matrix for calculating a second-order differential with respect to the gradient. [0068] states that the derivative of the output of the second neural network taken with respect to a parameter may be calculated through a backpropagation process and may be calculated to update the parameters of the first neural network and the second neural network. Read in that sense, claim 9 excludes the very quantity by which the specification carries out the updating of the shared parameters that the same limitation requires, and no disclosed embodiment falls within the claim. The claim accordingly fails to inform a person of ordinary skill in the art, with reasonable certainty, of the scope of the invention. See Nautilus, Inc. v. Biosig Instruments, Inc., 572 U.S. 898, 901, 910 (2014); In re Packard, 751 F.3d 1307 (Fed. Cir. 2014). A negative limitation complies with 35 U.S.C. 112(b) only where the boundaries of the protection sought are set forth definitely, albeit negatively (MPEP 2173.05(i); In re Wakefield, 422 F.2d 897, 899, 904 (CCPA 1970); In re Barr, 444 F.2d 588 (CCPA 1971)), and those boundaries are not set forth here. Two further parts of the record confirm the uncertainty rather than removing it. Claim 15, which depends from claim 9 through claim 14, recites calculating a second differential value of the second differential data with respect to a parameter of the layer of the first neural network, which is the same kind of quantity that one sense of the exclusion in claim 9 forbids. The reply filed 06/17/2026 describes the excluded quantity as a differential taken with respect to the network weights, which is the sense in which the disclosed training falls outside the claim. For examination, this limitation is treated as excluding only a direct calculation of a second-order differential by way of a Hessian matrix, consistent with [0042] and [0067]; a derivative of the output differential value taken with respect to the shared parameters and obtained by ordinary backpropagation, as described at [0067] and [0068], is within the claim. The remaining limitations of claim 9 are given their ordinary and customary meaning to a person of ordinary skill in the art. Applicant may overcome this rejection by amending claim 9 to state the variable with respect to which the excluded differential value is taken, or by cancelling the negative limitation. Claims 10-17 depend from claim 9, directly or through claim 14 or claim 16; they incorporate and do not cure this defect, and are rejected for the same reason. This ground of rejection is necessitated by Applicant's amendment filed 06/17/2026, which added the negative limitation to claim 9. See MPEP 706.07(a). Claim Rejections - 35 U.S.C. 112(d) The following is a quotation of 35 U.S.C. 112(d): (d) REFERENCE IN DEPENDENT FORMS.—Subject to subsection (e), a claim in dependent form shall contain a reference to a claim previously set forth and then specify a further limitation of the subject matter claimed. A claim in dependent form shall be construed to incorporate by reference all the limitations of the claim to which it refers. The following is a quotation of pre-AIA 35 U.S.C. 112, fourth paragraph: Subject to the following paragraph [i.e., the fifth paragraph of pre-AIA 35 U.S.C. 112], a claim in dependent form shall contain a reference to a claim previously set forth and then specify a further limitation of the subject matter claimed. A claim in dependent form shall be construed to incorporate by reference all the limitations of the claim to which it refers. Claims 4 and 11 are rejected under 35 U.S.C. 112(d) as being of improper dependent form for failing to further limit the subject matter of the claim upon which they depend, or for failing to include all the limitations of the claim upon which they depend. Claim 4 adds to claim 1 only that "an activation function of a layer of the second neural network comprises a function that multiplies a differential value of an activation function of a layer of the first neural network that corresponds to the layer of the second neural network." Claim 1, as amended on 06/17/2026, already requires that "an activation function of each layer of the second neural network is defined as a function that multiplies a derivative of an activation function of a corresponding layer of the first neural network." A condition that claim 1 imposes on each layer of the second neural network is necessarily met at any one layer of that network, so the single layer recited in claim 4 is already within what claim 1 requires. The transitional phrase "comprises" in claim 4 is open ended and is therefore satisfied wherever the narrower "is defined as" of claim 1 is satisfied. A differential value of an activation function is the derivative of that activation function, as the specification confirms at [0061], which identifies the differential value of an activation function of a layer as the first derivative of that activation function with respect to the pre-activation input of the layer. Claim 4 therefore restates in different words a requirement that claim 1 already imposes, and it adds no further limitation. Claim 11 stands in the identical relation to claim 9, which was amended on 06/17/2026 to add the same activation function requirement. This ground of rejection is newly applied and is necessitated by Applicant's amendment to claims 1 and 9 filed 06/17/2026. Applicant may cancel the claim(s), amend the claim(s) to place the claim(s) in proper dependent form, rewrite the claim(s) in independent form, or present a sufficient showing that the dependent claim(s) complies with the statutory requirements. Claims 6 and 12 are rejected under 35 U.S.C. 112(d) as being of improper dependent form for failing to further limit the subject matter of the claim upon which they depend, or for failing to include all the limitations of the claim upon which they depend. Claim 6 adds to claim 1 only that "parameters of the first neural network for the estimation of the output data are the same as parameters of the second neural network." Claim 1, as amended on 06/17/2026, already requires that "parameters of the second neural network are shared with parameters of the first neural network," and further requires training the two networks "by updating the shared parameters." Parameters that are shared between the first neural network and the second neural network are one and the same parameters, and are accordingly the same as one another. The specification uses the two expressions for a single relationship, stating at [0061] that parameters of the second neural network "may be the same as parameters of the first neural network" and, in the same sentence, that "the parameters may be shared." The qualifier "for the estimation of the output data" in claim 6 identifies no different parameters, because claim 1 already recites a first neural network "that estimates output data from the input data." Neither claim states how many of the parameters stand in the recited relationship, so claim 6 does not narrow claim 1 in that respect either. Claim 6 therefore restates in different words a requirement that claim 1 already imposes, and it adds no further limitation. Claim 12 stands in the identical relation to claim 9, which was amended on 06/17/2026 to add the same shared parameter requirement. This ground of rejection is newly applied and is necessitated by Applicant's amendment to claims 1 and 9 filed 06/17/2026. Applicant may cancel the claim(s), amend the claim(s) to place the claim(s) in proper dependent form, rewrite the claim(s) in independent form, or present a sufficient showing that the dependent claim(s) complies with the statutory requirements. Claim Rejections - 35 USC 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-6 and 8-16 are rejected under 35 U.S.C. 103 as being unpatentable over Czarnecki et al. (hereinafter Czarnecki) “Sobolev Training for Neural Networks” (2017) in view of Bishop (hereinafter Bishop) “Exact Calculation of the Hessian Matrix for the Multi-layer Perceptron” (1992). Bishop was disclosed in an IDS dated 6/28/2023. Regarding independent claim 1, Czarnecki teaches a processor-implemented method, the method comprising (Czarnecki: page 5, footnote 2, "All experiments were performed using TensorFlow [2] and the Sonnet neural network library [1]"; the training procedure Czarnecki sets out was carried out on those machine learning libraries, which are run by a processor, so the procedure Czarnecki describes is a processor-implemented method): the first neural network that estimates output data from the input data (Czarnecki: page 1, Abstract, "use neural networks…training them to produce outputs from inputs in emulation of a ground truth function or data creation process"; the model (a first neural network) is trained to produce an output (output data) from a given input (input data) in emulation of the ground truth function, which is an estimation of the output data from the input data; page 3, Section 2, "In this situation, the derivative information can easily be incorporated into training a neural network model of f by making derivatives of the neural network match the ones given by f"; the neural network model of the target function named there is the network whose outputs approximate that function's values at the training points), generating, using a second neural network, an output differential value of the output data with respect to the input data (Czarnecki: page 2, Figure 1(a), "Diamond nodes m and f indicate parameterised functions, where m is trained to approximate f. Green nodes receive supervision"; Figure 1(a) places beside the model node a further node of the compute graph (a second neural network), reached from the model node by an arrow labelled with the partial derivative with respect to the input and supervised by its own loss, and the quantity that node emits is the derivative of the model's output with respect to the model's input (an output differential value of the output data with respect to the input data)), wherein parameters of the second neural network are shared with parameters of the first neural network (Czarnecki: page 1, Abstract, "By optimising neural networks to not only approximate the function's outputs but also the function's derivatives we encode additional information about the target function within the parameters of the neural network"; the outputs and the derivatives are both encoded within one set of parameters, that of the network itself, so the derivative-emitting node holds no parameters of its own and operates on the parameters of the model node; page 2, Figure 1(a), "Solid lines indicate connections through which error signal from loss l, l1, and l2 are backpropagated through to train m"; the error signal of the derivative losses is returned to the same model node that the value loss trains, which is available only because the derivative node is parameterised by the parameters of the model node), and training the first neural network and the second neural network by updating the shared parameters based on ground truth data of the output data and ground truth data of the output differential value (Czarnecki: page 3, Section 2, Equation 1, "This causes the neural network to encode derivatives of the target function in its own derivatives. Such a model can still be trained using backpropagation and off-the-shelf optimisers."; the objective of Equation 1 adds, to the error between the model's output and the target function's value at each training point (ground truth data of the output data), a further error term for each derivative order between the model's derivative with respect to the input and the target function's derivative of that order (ground truth data of the output differential value), and minimising that combined objective by backpropagation and an off-the-shelf optimiser updates the single parameter set that carries both the model output and the derivative emitted beside it). Czarnecki does not expressly teach generating respective first neural network differential data by differentiating a respective output of each layer of a first neural network with respect to input data provided to the first neural network, by a forward propagation process of the first neural network; using the respective first neural network differential data; and an activation function of each layer of the second neural network is defined as a function that multiplies a derivative of an activation function of a corresponding layer of the first neural network. However, Bishop teaches generating respective first neural network differential data by differentiating a respective output of each layer of a first neural network with respect to an internal pre-activation of that network, by a forward propagation process of the first neural network (Bishop: page 2, "Using equations 1 and 2 we then obtain the forward propagation equation"; Equation 7 defines, for each unit of the network, the derivative of that unit's input with respect to the input of a chosen unit, and Equation 11 forms that quantity for a unit from the corresponding quantities already formed for the units that feed it; page 1, "Consider a feed-forward network in which the activation zi of the ith unit is a non-linear function of the input to the unit"; a unit's output is the activation function applied to that unit's input, so the derivative of a unit's output is the derivative of that unit's activation function multiplied by the quantity Equation 7 defines; page 2, "The second derivatives can now be written in the form"; Equation 9 forms exactly that product, the derivative of the activation function evaluated at a unit multiplied by that same unit's forward-propagated quantity, so the propagation delivers, for the output of every layer of the feed-forward network (a first neural network), a quantity of differential data (respective first neural network differential data); page 3, "The remaining elements of gli can then be found by forward propagation using equation 11"; the quantities for the successive layers are obtained by running that equation forward through the network, which is a forward propagation process of that network), using the respective first neural network differential data (Bishop: page 2, "where the sum runs over all units r which send connections to unit l"; the quantity formed at a unit is a sum, over the units that send connections to it, of the quantities already carried forward at those units, so each stage of the propagation consumes the differential data produced at the preceding layers and the quantity delivered at the output layer is built from all of them), and an activation function of each layer of the second neural network is defined as a function that multiplies a derivative of an activation function of a corresponding layer of the first neural network (Bishop: page 2, "Using equations 1 and 2 we then obtain the forward propagation equation"; each term of Equation 11 is the product of the derivative of the activation function evaluated at a feeding unit, the weight from that unit, and the quantity already carried forward at that unit, and substituting Equation 11 into the product Equation 9 forms shows that the output derivative carried at a unit is the derivative of that unit's own activation function multiplied by the weighted sum of the output derivatives already carried at the units that feed it, so the operation the propagation performs at each layer is a weighted sum of the quantities delivered by the preceding layer followed by a multiplication by the derivative of the activation function of that same layer of the network whose outputs are being differentiated, which is the function that plays, in the propagation of the differential data, the part the activation function plays in the propagation of the activations; page 4, "Before using the above equations in a software implementation, the appropriate expressions for the derivatives of the activation function should be substituted"; Bishop directs that the derivative of the activation function actually in use be substituted into those equations, Equation 21 supplying that derivative for the sigmoid, so the multiplier applied at each layer is the derivative of that layer's own activation function). Czarnecki and Bishop are analogous art. Both address the computation of derivative quantities of a multi-layer feed-forward neural network by propagation through that network. Accordingly, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to apply Bishop's forward-propagation derivative recursion to the method of Czarnecki, with a reasonable expectation of success, by initializing that recursion at the units that receive the input data instead of at a hidden unit, to teach generating respective first neural network differential data by differentiating a respective output of each layer of a first neural network with respect to input data provided to the first neural network, by a forward propagation process of the first neural network; using the respective first neural network differential data; and an activation function of each layer of the second neural network is defined as a function that multiplies a derivative of an activation function of a corresponding layer of the first neural network. So initialized, the quantity carried forward from layer to layer is the derivative of each layer's output with respect to the input data, and the quantity delivered at the output layer is the derivative of the output data with respect to the input data that Czarnecki's objective is written in terms of. Each stage of that propagation reuses the weights of the corresponding layer of the first neural network and multiplies by the derivative of the activation function of that same layer. The reasonable expectation of success rests on the form of the recursion. Equation 11 is a linear recursion whose only inputs are the network's own synaptic weights and its activation-function derivatives, each of which is available at the moment the ordinary forward evaluation reaches a unit. Bishop further provides expressly for a unit whose activation function is the identity and whose activation-function derivative is accordingly unity (Bishop: page 4, "For the case of linear output units, we have"). Starting the recursion at a unit that receives the input data therefore changes the initial condition and nothing else. This modification would have been motivated by the desire to obtain the derivatives of the network's output with respect to its input, which Czarnecki's training matches against the derivatives of the target function, layer by layer for a network of arbitrary feed-forward topology and in a form that can readily be implemented in software (Bishop: page 1, "In Section 2 we derive the algorithm for a network of arbitrary feed-forward topology, in a form which can readily be implemented in software"). Regarding dependent claim 2, Czarnecki, in view of Bishop, teach the method of claim 1, further comprising generating, by a layer of the second neural network, second differential data obtained by differentiating, with respect to the input data, an output of a layer of the first neural network (Bishop: page 2, "Using equations 1 and 2 we then obtain the forward propagation equation"; Equation 11 is applied once for each unit in turn, so the propagation is organised as a succession of stages, one for each layer of the network being differentiated, and the stage belonging to a layer (a layer of the second neural network) produces the differential quantity for that layer's output (second differential data), and as modified in the manner set out in the reason to combine above that quantity is the derivative of that layer's output with respect to the input data) from first differential data, of the respective first neural network differential data, obtained by differentiating an output of another layer previous to the layer of the first neural network, with respect to the input data (Bishop: page 2, "where the sum runs over all units r which send connections to unit l"; the quantities summed are those already formed at the units that send connections to the unit in question, which are the units of the preceding layer (another layer previous to the layer of the first neural network), and those quantities are themselves differential data of the same kind (first differential data, of the respective first neural network differential data), taken, as modified in the manner set out in the reason to combine above, with respect to the input data), based on parameters of the first neural network and a first differential value of an activation function of the layer of the first neural network (Bishop: page 2, "Using equations 1 and 2 we then obtain the forward propagation equation"; each term of Equation 11 carries the quantity already formed at a feeding unit through the weight of the connection traversed (parameters of the first neural network), page 2, "The second derivatives can now be written in the form"; the output derivative carried at a unit is the product Equation 9 forms of that unit's own activation-function derivative and the unit's forward-propagated quantity, so substituting Equation 11 into that product gives the output derivative at a layer as the weighted sum, over the units of the preceding layer, of the output derivatives already carried there, taken through the weights of the connections traversed and then multiplied by the derivative of the activation function of that later layer itself (a first differential value of an activation function of the layer of the first neural network), the later layer being the one whose output is being differentiated; page 1, "where wij is the synaptic weight from unit j to unit i"; the weights the quantities are carried through are the synaptic weights of the network itself rather than a separate set). Regarding dependent claim 3, Czarnecki, in view of Bishop, teach the method of claim 2, wherein a second differential value of the second differential data with respect to a parameter of the layer of the first neural network (Bishop: page 2, "The second derivatives can now be written in the form"; the output derivative at a layer is the product Equation 9 forms of that layer's activation-function derivative and the quantity Equation 11 carries, and that quantity is a sum of terms each of which is linear in one synaptic weight of that layer, so the derivative of the output derivative with respect to one such weight (a second differential value of the second differential data with respect to a parameter of the layer of the first neural network) follows directly from those equations; this differentiation is reasoning applied here to Bishop's equations and is not asserted to be a step Bishop itself performs) is calculated by multiplying the first differential data by the first differential value (Bishop: page 2, "where the sum runs over all units r which send connections to unit l"; differentiating with respect to the weight of one traversed connection removes that weight from the sum and leaves exactly the output derivative already carried forward at the feeding unit of the preceding layer (the first differential data) multiplied by the derivative of the activation function of the later layer itself (the first differential value), which is the product the claim recites). Regarding dependent claim 4, Czarnecki, in view of Bishop, teach the method of claim 1, wherein an activation function of a layer of the second neural network comprises a function that multiplies a differential value of an activation function of a layer of the first neural network that corresponds to the layer of the second neural network (Bishop: page 2, "Using equations 1 and 2 we then obtain the forward propagation equation"; Equation 11 forms the quantity at a unit from the quantities already formed at the units that feed it; page 2, "The second derivatives can now be written in the form"; the output derivative carried at a unit is the product Equation 9 forms of that unit's own activation-function derivative and the unit's forward-propagated quantity, and substituting Equation 11 into that product shows the operation performed at a layer of the differential propagation to be a weighted sum of the output derivatives carried forward from the preceding layer followed by a multiplication by the derivative of the activation function of that same layer, so the function applied at that stage of the propagation (an activation function of a layer of the second neural network) is one that multiplies, for the layer being differentiated, the derivative of the activation function (a differential value of an activation function of a layer of the first neural network that corresponds to the layer of the second neural network); page 4, "Before using the above equations in a software implementation, the appropriate expressions for the derivatives of the activation function should be substituted"; the multiplier used at each stage is the derivative of the activation function actually in use at the corresponding layer, Equation 21 supplying it for the sigmoid). Regarding dependent claim 5, Czarnecki, in view of Bishop, teach the method of claim 1, further comprising storing a respective differential value of a corresponding activation function for each layer of the first neural network (Bishop: page 4, "Before using the above equations in a software implementation, the appropriate expressions for the derivatives of the activation function should be substituted"; the derivative of the activation function is evaluated for the units of every layer and held for use in the propagation equation, Equation 21 supplying the expression to be evaluated for the sigmoid, so a respective differential value of the corresponding activation function is retained for each layer; that retention necessarily follows rather than merely probably follows, because each successive stage of Equation 11 consumes the activation-function derivative of each feeding unit, and a value produced when the forward evaluation reaches a unit and consumed at a later stage of the same propagation must be held between those stages) in a forward propagation process of the first neural network (Bishop: page 1, "Consider a feed-forward network in which the activation zi of the ith unit is a non-linear function of the input to the unit"; every unit's activation is computed from that unit's weighted input as the network is evaluated from its inputs onward, and the derivative of the activation function is a function of that same unit input, so its value is available at the moment the forward evaluation reaches the unit; page 5, "the {gli} are obtained by forward propagation using equation 11"; the differential quantities that consume those activation derivatives are themselves obtained by forward propagation, so the derivative values are evaluated and held in the course of that forward propagation of the network rather than in a later pass). Regarding dependent claim 6, Czarnecki, in view of Bishop, teach the method of claim 1, wherein parameters of the first neural network for the estimation of the output data are the same as parameters of the second neural network (Czarnecki: page 1, Abstract, "By optimising neural networks to not only approximate the function's outputs but also the function's derivatives we encode additional information about the target function within the parameters of the neural network"; the same parameters of the network carry both the approximation of the function's outputs, which is the estimation of the output data the claim recites, and the approximation of the function's derivatives, so the parameters used for the estimation and the parameters of the derivative-emitting node are one and the same set; page 2, Figure 1(a), "Solid lines indicate connections through which error signal from loss l, l1, and l2 are backpropagated through to train m."; Figure 1(a) draws no parameter set for the derivative node apart from that of the model node, and returns the derivative losses to the model node itself). Regarding dependent claim 8, Czarnecki, in view of Bishop, teach a non-transitory computer-readable storage medium storing instructions that, when executed by a processor, cause the processor to perform the method of claim 1 (Czarnecki: page 2, "derivative information about the desired function in a way that can easily be incorporated into any training pipeline using modern machine learning libraries"; a training pipeline built on machine learning libraries is a body of stored program instructions (instructions) that a processor runs, and incorporating the described training into such a pipeline puts those instructions on the storage from which the libraries and the pipeline are loaded (a non-transitory computer-readable storage medium); page 5, footnote 2, "All experiments were performed using TensorFlow [2] and the Sonnet neural network library [1]"; the described procedure was in fact carried out by running such stored instructions on a processor). The substantive limitations of claim 8 are the non-transitory computer-readable storage medium form counterparts of the limitations of claim 1 and are taught for the same reasons set forth above, and the motivation to combine is the same. Regarding independent claim 9, Czarnecki teaches an apparatus, comprising: a processor configured to (Czarnecki: page 5, footnote 2, "All experiments were performed using TensorFlow [2] and the Sonnet neural network library [1]"; the described procedure was carried out on those machine learning libraries, which are run by a processor of a computer, so the arrangement that carries it out is an apparatus comprising a processor so configured; page 2, "derivative information about the desired function in a way that can easily be incorporated into any training pipeline using modern machine learning libraries"; incorporating the procedure into a training pipeline configures the processor that runs the pipeline to carry out the steps of that procedure): estimate output data with respect to input data (Czarnecki: page 1, Abstract, "training them to produce outputs from inputs in emulation of a ground truth function or data creation process"; the model produces an output (output data) for a given input (input data) in emulation of the ground truth function, which is an estimate of the output for that input), generate an output differential value of the output data with respect to the input data (Czarnecki: page 2, Figure 1(a), "Diamond nodes m and f indicate parameterised functions, where m is trained to approximate f. Green nodes receive supervision"; the compute graph carries, beside the model node, a node reached from it by an arrow labelled with the partial derivative with respect to the input, and that node emits the derivative of the model's output with respect to the model's input (an output differential value of the output data with respect to the input data)) through forward propagation of a second neural network (Czarnecki: page 3, Section 2, "Figure 1 illustrates compute graphs for non-stochastic and stochastic Sobolev Training of order 2"; the derivative node is a node of a compute graph, and a compute graph delivers a quantity by being evaluated from its inputs onward through its nodes, so the derivative node emits its quantity by propagation forward through that compute graph (a second neural network) rather than by any return pass over it), wherein parameters of the second neural network are shared with parameters of the first neural network (Czarnecki: page 1, Abstract, "By optimising neural networks to not only approximate the function's outputs but also the function's derivatives we encode additional information about the target function within the parameters of the neural network"; the outputs and the derivatives are both encoded within one set of parameters, that of the network itself, so the derivative-emitting graph holds no parameters of its own; page 2, Figure 1(a), "Solid lines indicate connections through which error signal from loss l, l1, and l2 are backpropagated through to train m"; the error signal of the derivative losses is returned to the same model node that the value loss trains, which is available only because that graph is parameterised by the parameters of the model node; page 5, Section 4.1, "use a first order Sobolev method"; the same single parameter set carries both quantities in the first-order embodiment of Section 4.1, which is the embodiment applied here), and train the first neural network and the second neural network by updating the shared parameters based on the estimated output data and the output differential value without using the differential value of the output differential value (Czarnecki: page 3, Section 2, Equation 1, "This causes the neural network to encode derivatives of the target function in its own derivatives. Such a model can still be trained using backpropagation and off-the-shelf optimisers"; the objective of Equation 1 is formed from the model's own output (the estimated output data) and, at the first derivative order, the derivative of the model's output with respect to the model's input (the output differential value), each compared against the corresponding quantity of the target function, and it is minimised by ordinary backpropagation and an off-the-shelf optimiser, which updates the single parameter set that carries both quantities; page 5, Section 4.1, "use a first order Sobolev method (second order derivatives of ReLU networks with a linear output layer are constant, zero)"; the embodiment applied here is that first-order one, whose objective carries only the value term and the first-derivative term, so its training forms and fits no derivative of the model's own derivative, performs no direct calculation of a second-order differential by way of a Hessian matrix, and equally forms no differential value of the output differential value with respect to the input data, which satisfies the recitation under either of the interpretations stated above). Czarnecki does not expressly teach by differentiating a respective output of each layer of a first neural network with respect to input data provided to the first neural network by a forward propagation process of the first neural network; using respective first neural network differential data; and an activation function of each layer of the second neural network is defined as a function that multiplies a derivative of an activation function of a corresponding layer of the first neural network. However, Bishop teaches by differentiating a respective output of each layer of a first neural network with respect to an internal pre-activation of that network by a forward propagation process of the first neural network (Bishop: page 2, "Using equations 1 and 2 we then obtain the forward propagation equation"; Equation 7 defines, for each unit of the network, the derivative of that unit's input with respect to the input of a chosen unit, and Equation 11 forms that quantity for a unit from the corresponding quantities already formed at the units that feed it; page 1, "Consider a feed-forward network in which the activation zi of the ith unit is a non-linear function of the input to the unit"; a unit's output is the activation function applied to that unit's input, so the derivative of a unit's output is the derivative of that unit's activation function multiplied by the quantity Equation 7 defines; page 2, "The second derivatives can now be written in the form"; Equation 9 forms exactly that product, so the propagation delivers a differential quantity for the output of every layer of the feed-forward network (a first neural network); page 3, "The remaining elements of gli can then be found by forward propagation using equation 11"; those quantities for the successive layers are obtained by running that equation forward through the network, which is a forward propagation process of that network); and using respective first neural network differential data (Bishop: page 2, "where the sum runs over all units r which send connections to unit l"; the quantity formed at a unit is a sum, over the units that send connections to it, of the quantities already carried forward at those units, so the quantity delivered at the output layer is built by using the differential data produced at each preceding layer), and an activation function of each layer of the second neural network is defined as a function that multiplies a derivative of an activation function of a corresponding layer of the first neural network (Bishop: page 2, "Using equations 1 and 2 we then obtain the forward propagation equation"; each term of Equation 11 is the product of the derivative of the activation function evaluated at a feeding unit, the weight from that unit, and the quantity already carried forward at that unit, and substituting Equation 11 into the product Equation 9 forms shows that the output derivative carried at a unit is the derivative of that unit's own activation function multiplied by the weighted sum of the output derivatives already carried at the units that feed it, so the operation the propagation performs at each layer is a weighted sum of the quantities delivered by the preceding layer followed by a multiplication by the derivative of the activation function of that same layer of the network being differentiated; page 4, "Before using the above equations in a software implementation, the appropriate expressions for the derivatives of the activation function should be substituted"; the multiplier applied at each layer is the derivative of that layer's own activation function, Equation 21 supplying it for the sigmoid). Czarnecki and Bishop are analogous art. Both address the computation of derivative quantities of a multi-layer feed-forward neural network by propagation through that network. Accordingly, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to apply Bishop's forward-propagation derivative recursion to the apparatus of Czarnecki, with a reasonable expectation of success, by initializing that recursion at the units that receive the input data instead of at a hidden unit, to teach by differentiating a respective output of each layer of a first neural network with respect to input data provided to the first neural network by a forward propagation process of the first neural network; using respective first neural network differential data; and an activation function of each layer of the second neural network is defined as a function that multiplies a derivative of an activation function of a corresponding layer of the first neural network. So initialized, the quantity carried forward from layer to layer is the derivative of each layer's output with respect to the input data, and the quantity delivered at the output layer is the derivative of the output data with respect to the input data that Czarnecki's objective is written in terms of. Each stage of that propagation reuses the weights of the corresponding layer of the first neural network and multiplies by the derivative of the activation function of that same layer. Each unit's activation and its differential quantity are formed as the single forward evaluation reaches that unit. The estimating is therefore carried out by the differentiating, and both by one forward propagation process of the first neural network. The reasonable expectation of success rests on the form of the recursion. Equation 11 is a linear recursion whose only inputs are the network's own synaptic weights and its activation-function derivatives, each of which is available at the moment the ordinary forward evaluation reaches a unit. Bishop further provides expressly for a unit whose activation function is the identity and whose activation-function derivative is accordingly unity (Bishop: page 4, "For the case of linear output units, we have"). Starting the recursion at a unit that receives the input data therefore changes the initial condition and nothing else. This modification would have been motivated by the desire to obtain the derivatives of the network's output with respect to its input, which Czarnecki's training matches against the derivatives of the target function, layer by layer for a network of arbitrary feed-forward topology and in a form that can readily be implemented in software (Bishop: page 1, "In Section 2 we derive the algorithm for a network of arbitrary feed-forward topology, in a form which can readily be implemented in software"). Regarding dependent claim 10, Czarnecki, in view of Bishop, teach the apparatus of claim 9, further comprising a memory configured to store a respective differential value of an activation function for each layer of the first neural network (Bishop: page 4, "Before using the above equations in a software implementation, the appropriate expressions for the derivatives of the activation function should be substituted"; the derivative of the activation function is evaluated for the units of every layer and retained for use in the propagation equation, and each such value, produced when the forward evaluation reaches its layer and consumed at a later stage of the forward propagation Equation 11 carries out, must necessarily be held between those stages, so the software implementation (a memory configured to store a respective differential value of an activation function for each layer) into which Bishop directs the expressions be substituted necessarily holds each of those per-layer values in storage while the propagation runs) wherein the respective differential value is obtained through the forward propagation of the first neural network (Bishop: page 1, "Consider a feed-forward network in which the activation zi of the ith unit is a non-linear function of the input to the unit"; every unit's activation is computed from that unit's weighted input as the network is evaluated from its inputs onward, and the derivative of the activation function is a function of that same unit input, so its value is produced at the moment the forward evaluation reaches the unit; page 5, "the {gli} are obtained by forward propagation using equation 11"; the differential quantities that consume those activation derivatives are themselves obtained by forward propagation, so the value held for each layer is obtained in the course of that forward propagation of the network). Regarding dependent claim 11, the rejection of claim 4 is applied in the same manner to the corresponding apparatus form recitation of claim 11. Regarding dependent claim 12, the rejection of claim 6 is applied in the same manner to the corresponding apparatus form recitation of claim 12. Regarding dependent claim 13, Czarnecki, in view of Bishop, teach the apparatus of claim 9, wherein the first neural network and the second neural network are trained based on a first loss function and a second loss function (Czarnecki: page 3, Section 2, Equation 1, "This causes the neural network to encode derivatives of the target function in its own derivatives"; the objective of Equation 1 is a sum of two kinds of error term, one taken on the model's output and one taken on the model's derivative, and training minimises that sum, so the training proceeds on a first loss function and a second loss function), wherein the first loss function is based on ground truth of the output data and an estimated value of the output data that is output from the forward propagation of the first neural network (Czarnecki: page 1, Abstract, "training them to produce outputs from inputs in emulation of a ground truth function or data creation process"; the first error term of Equation 1 is taken between the model's output (an estimated value of the output data that is output from the forward propagation of the first neural network), which the model produces for a training point by evaluating its layers in the forward direction, and the value of the ground truth function at that point (ground truth of the output data)), and the second loss function is based on the output differential value and ground truth data of the output differential value with respect to the input data (Czarnecki: page 2, "Sobolev Training exploits this property, and tries to match not only the output of the function being trained but also its derivatives."; the second error term of Equation 1 is taken between the derivative of the model's output with respect to the model's input (the output differential value) and the corresponding derivative of the ground truth function (ground truth data of the output differential value with respect to the input data), which is the matching of derivatives Czarnecki describes). Regarding dependent claim 14, the rejection of claim 2 is applied in the same manner to the corresponding apparatus-form recitation of claim 14. Claim 14 additionally recites that the second neural network comprises a layer defined to output second differential data. Each stage of the forward propagation of the differential data set out in the rejection of claim 2 above is the stage belonging to one layer of the network being differentiated and is defined by Bishop's equations to deliver the output derivative for that layer, so the propagation comprises, for each such layer, a stage defined to output that layer's second differential data (Bishop: page 3, "The remaining elements of gli can then be found by forward propagation using equation 11"). Regarding dependent claim 15, the rejection of claim 3 is applied in the same manner to the corresponding apparatus-form recitation of claim 15. Regarding dependent claim 16, Czarnecki, in view of Bishop, teach the apparatus of claim 9, wherein the processor is configured to calculate differential data obtained by differentiating the output of each layer of the first neural network with respect to the input data (Bishop: page 3, "The remaining elements of gli can then be found by forward propagation using equation 11"; the propagation equation is applied in turn to every unit of the network, so a differential quantity is produced for the output of each layer, and as modified in the manner set out in the reason to combine above that quantity is the derivative of each layer's output with respect to the input data; page 3, "The number of forward passes needed to evaluate all elements of {gli} will depend on the network topology, but will typically scale like the number of (hidden plus output) units in the network"; the passes enumerated there are the ones that produce the differential quantities for the whole network, so the calculation Bishop sets out covers the output of each layer rather than of the output layer alone). Claims 7 and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Czarnecki in view of Bishop, as applied in the rejections of claims 1 and 16, respectively above, and further in view of Baydin et al. (hereinafter Baydin) "Automatic Differentiation in Machine Learning: a Survey" (2018). Regarding dependent claim 7, Czarnecki, in view of Bishop, teach all the elements of claim 1. Czarnecki and Bishop do not expressly teach wherein the generating of the respective first neural network differential data further comprises: determining select input data among plural input data, for which a calculation of differential value is determined to be needed; and for each of the select input data, storing corresponding respective first neural network differential data obtained by respectively differentiating the outputs of each layer of the first neural network with a corresponding select input data. However, Baydin teaches determining select input data among plural input data, for which a calculation of differential value is determined to be needed (Baydin: page 25, Section 5.2, "For procedures coded in ANSI C, the ADIC tool (Bischof et al., 1997) implements AD as a source code transformation after the specification of dependent and independent variables"; the independent variables are specified before the derivative computation is generated, so which of the program's several inputs (plural input data) are the ones whose derivatives are wanted (select input data, for which a calculation of differential value is determined to be needed) is settled by that specification); and for each of the select input data, storing corresponding respective first neural network differential data obtained by respectively differentiating the outputs of each layer of the first neural network with a corresponding select input data (Baydin: page 9, Section 3.1, "each forward pass of AD is initialized by setting only one of the variables ... and setting the rest to zero"; a pass is run for one selected independent variable at a time and carries, alongside each intermediate value, the derivative of that value with respect to that one input, which is the differential data obtained by differentiating the outputs with respect to that corresponding selected input; page 10, "giving us one column of the Jacobian matrix"; the derivative values a pass carries are retained as one column of the Jacobian matrix, one column for each selected input; page 10, "Thus, the full Jacobian can be computed in n evaluations"; a further column is produced only by a further evaluation for a further selected input, so what is computed and retained is the derivative data for the selected inputs alone). Because Czarnecki, in view of Bishop, and Baydin are analogous art with all three addressing the computation of the derivatives of a layered numerical computation with respect to its inputs, accordingly, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to apply Baydin's specification of the independent variables to the method of Czarnecki in view of Bishop, with a reasonable expectation of success, by designating in advance the inputs for which a derivative is wanted and running the forward propagation of the differential data once for each designated input, each such run carrying and retaining only the differential data belonging to that input, that designation requiring no change to the propagation itself and so supplying the reasonable expectation of success, to teach wherein the generating of the respective first neural network differential data further comprises: determining select input data among plural input data, for which a calculation of differential value is determined to be needed; and for each of the select input data, storing corresponding respective first neural network differential data obtained by respectively differentiating the outputs of each layer of the first neural network with a corresponding select input data. This modification would have been motivated by the desire to obtain a wanted derivative in a single forward pass instead of evaluating the whole Jacobian matrix (Baydin: page 10). Regarding dependent claim 17, Czarnecki, in view of Bishop, teach all the elements of claim 16. Czarnecki and Bishop do not expressly teach wherein, in the calculating of the differential data, the processor is configured to: determine select input data among plural input data, for which a calculation of a differential value is determined to be needed; and for each of the select input data, calculate the differential data corresponding respectively to first neural network differential data obtained by respectively differentiating the outputs of each layer of the first neural network with a corresponding select input data. However, Baydin teaches determine select input data among plural input data, for which a calculation of a differential value is determined to be needed (Baydin: page 25, Section 5.2, "For procedures coded in ANSI C, the ADIC tool (Bischof et al., 1997) implements AD as a source code transformation after the specification of dependent and independent variables"; the independent variables are specified before the derivative computation is generated, so which of the program's several inputs (plural input data) are the ones whose derivatives are wanted (select input data, for which a calculation of differential value is determined to be needed) is settled by that specification), and for each of the select input data, calculate the differential data corresponding respectively to first neural network differential data obtained by respectively differentiating the outputs of each layer of the first neural network with a corresponding select input data (Baydin: page 9, Section 3.1, "each forward pass of AD is initialized by setting only one of the variables ... and setting the rest to zero"; a pass is run for one selected independent variable at a time and carries, alongside each intermediate value, the derivative of that value with respect to that one input, which is the differential data obtained by differentiating the outputs with respect to that corresponding selected input; page 10, "giving us one column of the Jacobian matrix"; the derivative values a pass carries are retained as one column of the Jacobian matrix, one column for each selected input; page 10, "Thus, the full Jacobian can be computed in n evaluations"; a further column is produced only by a further evaluation for a further selected input, so what is computed and retained is the derivative data for the selected inputs alone). Because Czarnecki, in view of Bishop, and Baydin are analogous art with all three addressing the computation of the derivatives of a layered numerical computation with respect to its inputs, accordingly, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to apply Baydin's specification of the independent variables to the apparatus of Czarnecki in view of Bishop, with a reasonable expectation of success, by designating in advance the inputs for which a derivative is wanted and running the forward propagation of the differential data once for each designated input, each such run carrying and retaining only the differential data belonging to that input, that designation requiring no change to the propagation itself and so supplying the reasonable expectation of success, to teach wherein, in the calculating of the differential data, the processor is configured to: determine select input data among plural input data, for which a calculation of a differential value is determined to be needed; and for each of the select input data, calculate the differential data corresponding respectively to first neural network differential data obtained by respectively differentiating the outputs of each layer of the first neural network with a corresponding select input data. This modification would have been motivated by the desire to obtain a wanted derivative in a single forward pass instead of evaluating the whole Jacobian matrix (Baydin: page 10). Response to Arguments Applicant’s amendment to the Specification filed 6/17/2026 that replaces the title with Method and Apparatus with Neural Network Training with Differential Data is persuasive. Thus, the title objection set forth in the Office Action dated 3/18/2026 is hereby withdrawn. Applicant’s claim amendments and Remarks filed 6/17/2026 traversing the 35 U.S.C. 101 rejections are persuasive. Thus, the 35 U.S.C. 101 rejections set forth in the Office Action dated 3/18/2026 are hereby withdrawn. Applicant’s claim amendments and Remarks filed 6/17/2026 traversing the 35 U.S.C. 103 rejections have been fully considered but they are not persuasive. Regarding Applicant's argument that Czarnecki alone does not teach the differential data limitations, Remarks page 15, Applicant argues that Czarnecki does not disclose generating the respective first neural network differential data by a forward propagation process, and therefore cannot teach generating the output differential value using that data. Examiner respectfully disagrees. Applicant’s argument attacks a reference individually where the rejection rests on a combination. Office Action dated 3/18/2026 stated in terms that Czarnecki does not expressly teach that limitation and turned to Bishop for it, and this action does the same. One cannot show nonobviousness by attacking references individually where the rejections are based on combinations of references. In re Keller, 642 F.2d 413, 426 (CCPA 1981); In re Merck and Co., 800 F.2d 1091, 1097 (Fed. Cir. 1986); MPEP 2145, subsection IV. The test is what the combined teachings would have suggested to one of ordinary skill in the art, which is the proposition Applicant itself quotes at pages 16 and 17 of the Remarks. The argument is not persuasive and the rejection is maintained. Regarding Applicant's argument, Remarks page 15, that Bishop is directed to computing the Hessian matrix rather than to avoiding it, Examiner respectfully disagrees. Three answers are given, each sufficient on its own: First, a reference is prior art for everything it teaches and not only for the result its author was after; the disclosure is evaluated for what it fairly teaches one of ordinary skill in the art. MPEP 2123, subsection I; In re Heck, 699 F.2d 1331, 1332 to 1333 (Fed. Cir. 1983). Bishop is relied upon for a discrete teaching, namely the layer by layer forward propagation of first order per-layer derivative quantities and the multiplication by the derivative of the activation function of the corresponding layer. Second, what the rejection takes from Bishop is that forward propagation, and not Bishop's later assembly of second derivatives with respect to the weights, so the premise that the borrowed teaching requires the differential of the output differential value is not correct. Third, the argument is aimed at a limitation that the argued claim does not contain. At page 16 of the Remarks Applicant frames the point as a contention that Czarnecki and Bishop do not teach training the first neural network and the second neural network without using the differential value of the output differential value as recited by claim 1. Claim 1 recites no such limitation. Claim 1 as amended ends by reciting training the first neural network and the second neural network by updating the shared parameters based on ground truth data of the output data and ground truth data of the output differential value. Arguments must be directed to the claims as they stand; limitations not appearing in the claims cannot be relied upon for patentability. MPEP 2145, subsection VI; In re Self, 671 F.2d 1344, 1348 (CCPA 1982). The limitation does appear in claim 9, where it is addressed both in the rejection under 35 U.S.C. 112(b) above and in the rejection under 35 U.S.C. 103 above. The argument is not persuasive as to claims 1 to 8. Regarding Applicant's argument, Remarks pages 16-17, that the references teach away from the claimed combination and that there is no motivation to combine them, Examiner respectfully disagrees. A reference teaches away only where it criticizes, discredits or otherwise discourages the solution claimed. In re Fulton, 391 F.3d 1195, 1201 (Fed. Cir. 2004); MPEP 2145, subsection X, paragraph D. Neither reference does so. Bishop nowhere says that forward propagated first order derivative quantities should not be used to obtain an input derivative; it is simply directed to a different objective, and a reference does not teach away merely because its author had a different goal or because it presents the claimed approach as one alternative among others. In re Gurley, 27 F.3d 551, 553 (Fed. Cir. 1994). The factual premise is also incorrect. Czarnecki's objective runs over derivative orders, and at first order the objective matches only first derivatives and the model is trained by ordinary backpropagation with off-the-shelf optimizers; it is Czarnecki's second order variant that involves second derivatives. Finally, the reason to combine that Applicant addresses is not the reason stated in this action. The rejections above are modified to reach the limitations added on 06/17/2026, and the reason to combine is restated accordingly: it rests on obtaining the input derivatives that Czarnecki's objective already requires, accurately and layer by layer in a forward pass, and the clause of the prior rationale concerned with evaluating all elements of the Hessian matrix exactly is not carried forward. The argument is not persuasive. Regarding Applicant's argument, Remarks pages 16-17, that the proposed modification would change the principle of operation of the primary reference or render it inoperable, Examiner respectfully disagrees. Applicant recites the standards of MPEP 2143.01 but identifies no principle of operation that the modification would change and no purpose of Czarnecki that would be defeated. An allegation unsupported by a showing is entitled to little weight. MPEP 2145; In re Geisler, 116 F.3d 1465, 1470 (Fed. Cir. 1997). On the merits the modification preserves Czarnecki's principle of operation, which is to augment the training objective so that the network's derivatives with respect to its inputs match the derivatives of the target function. That principle requires the input derivative to be available during training. Supplying Bishop's exact layer by layer forward computation of that derivative changes how a quantity Czarnecki already needs is obtained; it does not change what Czarnecki optimizes or why, and Czarnecki remains operable for its intended purpose. Neither the objective nor the operating principle of the primary reference is disturbed. The argument is not persuasive. Regarding Applicant's argument, Remarks page 17, that the limitations added to claim 1 are not taught, Examiner respectfully disagrees. Each limitation added to claim 1 on 06/17/2026 is the subject matter of a dependent claim that the prior Office action already rejected over the same two references, and the findings carry over. The recitation that parameters of the second neural network are shared with parameters of the first neural network is the subject matter of claim 6, and of claim 12 in apparatus form, both of which were rejected as taught by Czarnecki's single shared parameter set and by Bishop's reuse of the network's own weights in the derivative calculation. The recitation that an activation function of each layer of the second neural network is defined as a function that multiplies a derivative of an activation function of a corresponding layer of the first neural network is the subject matter of claim 4, and of claim 11 in apparatus form, both of which were rejected as taught by Bishop's per-layer recursion, which is applied at every layer in the forward pass. The recitation of training by updating the shared parameters is taught by Czarnecki, which trains the one parameter set that both of its nodes share. Moving the subject matter of a rejected dependent claim into the claim from which it depends does not disturb the rejection. The argument is not persuasive. One consequence of that movement is addressed separately: because claims 4, 6, 11 and 12 now recite subject matter that the claims from which they depend already require, they are rejected above under 35 U.S.C. 112(d). Regarding Applicant's argument directed to independent claim 9 and to the dependent claims. The only sentence in the Remarks directed to independent claim 9 argues that claim over Hall, Reisteter and Dhuse. None of those three references is of record in this application, and none is applied against any claim. Claim 9 is rejected over Czarnecki in view of Bishop. The sentence appears to have been carried over from a reply in another application, and it presents no contention addressed to the references actually applied that this Office can answer on the merits. Applicant is advised of the apparent error so that it may be corrected in any further reply. Because no argument reaches claim 9 as rejected, the rejection of claim 9 stands on the record made, and the limitation added to claim 9 on 06/17/2026 is nevertheless addressed on its merits in the rejections above, both under 35 U.S.C. 112(b) and under 35 U.S.C. 103, as MPEP 707.07(f) requires. As to the dependent claims, the assertion that they are directed to patentable subject matter by virtue of their dependency and by virtue of the additional features recited identifies no feature and advances no reason. A bare assertion that dependent claims are allowable, unaccompanied by an explanation of why the additional limitations distinguish over the applied art, is not a separate argument for patentability. In re Lovin, 652 F.3d 1349, 1357 (Fed. Cir. 2011); MPEP 2145. Each dependent claim is separately addressed in the rejections above. Applicant’s arguments are not persuasive and the rejections under 35 U.S.C. 103 are maintained as modified above. Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to KUANG FU CHEN whose telephone number is (571)272-1393. The examiner can normally be reached M-F 9:00-5:30pm ET. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Jennifer Welch can be reached on (571) 272-7212. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /KC CHEN/Primary Patent Examiner, Art Unit 2143
Read full office action

Prosecution Timeline

Jun 28, 2023
Application Filed
Mar 18, 2026
Non-Final Rejection mailed — §103, §112
Jun 12, 2026
Applicant Interview (Telephonic)
Jun 12, 2026
Examiner Interview Summary
Jun 17, 2026
Response Filed
Sep 15, 2026
Final Rejection mailed — §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12749010
MACHINE LEARNING MODEL FOR PREDICTING DRIVER-VEHICLE COMPATIBILITY
4y 8m to grant Granted Sep 29, 2026
Patent 12737658
MICROWAVE PHOTONIC QUANTUM PROCESSOR
3y 5m to grant Granted Sep 15, 2026
Patent 12737604
ANALOGUE ARITHMETIC UNIT AND NUROMORPHIC DEVICE
3y 5m to grant Granted Sep 15, 2026
Patent 12725036
EXPLAINABLE DEEP INTERPOLATION
3y 9m to grant Granted Sep 01, 2026
Patent 12718125
APPLICATION OF LOCAL INTERPRETABLE MODEL-AGNOSTIC EXPLANATIONS ON DECISION SYSTEMS WITHOUT TRAINING DATA
5y 4m to grant Granted Aug 25, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
80%
Grant Probability
99%
With Interview (+69.0%)
2y 11m (~0m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 271 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month