Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 3/30/2026 has been entered.
Remarks
This Office Action is responsive to Applicants' Amendment filed on March 30, 2026, in which claims 1-7, 11-19, 21-26, and 28 are currently amended. Claims 1-30 are currently pending.
Response to Arguments
Applicant’s arguments with respect to rejection of claims 1-30 under 35 U.S.C. 101 based on amendment have been considered, however, are not persuasive.
With respect to Applicant’s arguments on p. 13 of the Remarks submitted 3/30/2026 that the additional elements “generating an inference using the neural network with the one or more refined parameters” and “generating an inference based on the first and second output data tensors” integrate the judicial exception into a practical application, Examiner respectfully disagrees. First, Examiner notes that the recited claim limitations appear to be explicitly directed towards outputting data, which is insignificant extra-solution activity (See MPEP 2106.05(g)) which is well-understood, routine, and conventional in the art (See MPEP 2106.05(d)(II)(i) and MPEP 2106.05(d)(II)(iv)) and does not integrate the judicial exception into a practical application.
With respect to Applicant’s arguments on p. 13 of the Remarks submitted 3/30/2026 that machine learning inference is improved by the judicial exception and the claims are therefore patent eligible, Examiner respectfully disagrees. First, Examiner notes that the claims explicitly tie the neural networks to mathematical computations. The neural networks explicitly operate on, and output tensors. Examiner notes that in view of Example 47 Claim 2 of the 2024 SME guidance, this provides strong support that the claims as a whole are directed towards mathematical calculations and relationships, which is a judicial exception. In other words, the generated output which Applicant alleges the claims improve, is itself math, such that the claims appear to be using generic computer components to improve the judicial exception, not the other way around ( MPEP 2106.07(a)(II) "employing well-known computer functions to execute an abstract idea, even when limiting the use of the idea to one particular environment, does not integrate the exception into a practical application"). Examiner asserts that Applicant is relying entirely on the judicial exception to provide a technical improvement rather than on additional elements (2106.05(a) "It is important to note, the judicial exception alone cannot provide the improvement”). The additional elements in the claim cannot provide the technical improvement as they are insignificant extra-solution activity (MPEP 2106.05(g)) that is well-understood, routine, and conventional in the art (MPEP 2106.05(d)(II)(i) and MPEP 2106.05(d)(II)(iv)) that the MPEP explicitly states does not integrate a judicial exception into a practical application. Examiner asserts that the independent claims are recited at a very high, abstract level, and do not recite the concrete mechanics necessary for Applicant’s asserted technical improvement. For these reasons and those further detailed below, Examiner asserts that it is reasonable and appropriate to maintain the rejection of claims 1-30 under 35 U.S.C. 101.
Applicant’s arguments with respect to rejection of claims 1-29 under 35 U.S.C. 103 based on amendment have been considered. The argument is moot in view of a new ground of rejection set forth below.
Applicant’s arguments with respect to rejection of claim 30 under 35 U.S.C. 103 based on amendment have been considered, however, are not persuasive. First, Applicant appears to have narrowly interpreted what it means for a multiplication operation to be “based on a summation”. Dinh's affine coupling layer generates the second output component by multiplying the second input component by a convolution-derived scale term and summing the result with a convolution derived translation term; thus, the second output is generated using a multiplication operation and a summation involving outputs of convolutional subnetworks applied to the first input tensor. Additionally, in view of Applicant’s narrow interpretation, Dinh's formula y1:D = x1:D * exp(s(x1:d))+t(x1:d) is mathematically equivalent (by factoring out the scale term) to y2 = exp(s(x1:d))*(xd+1:D+t(x1:d)*exp(-s)x1:d)) where x1:d is the first input data tensor, xd+1:D is the second input data tensor, s(x1) is the convolutional scale output, t(x1) is the convolutional translation output, and * is element-wise multiplication. For at least these reasons, Examiner asserts that it is reasonable and appropriate to maintain the rejection of claim 30 in view of Dinh.
Claim Rejections - 35 USC § 101
101 Rejection
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-30 are rejected under 35 USC § 101 because the claimed invention is directed to non-statutory subject matter.
Regarding Claim 1: Claim 1 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1 Analysis: Claim 1 is directed to a method, which is directed to a process, one of the statutory categories.
Step 2A Prong One Analysis: Claim 1 under its broadest reasonable interpretation is a series of mental processes. For example, but for the generic computer components language, the above limitations in the context of this claim encompass neural network processing, including the following:
Separating the first data tensor into a first subtensor and a second subtensor using a tensor splitting operation (observation, evaluation, and judgement),
refining one or more parameters of the first layer of the neural network based at least in part on the stored second subtensor (observation, evaluation, and judgement)
generating an inference using the neural network with the one or more refined parameters of the first layer of the neural network (observation, evaluation, and judgement)
Therefore, claim 1 recites an abstract idea which is a judicial exception.
Step 2A Prong Two Analysis: Claim 1 recites additional elements “computer implemented method of machine learning” and “neural network”. However, these additional features are computer components recited at a high-level of generality, such that they amount to no more than mere instructions to apply the judicial exception using a generic computer component. An additional element that merely recites the words “apply it” (or an equivalent) with the judicial exception, or merely includes instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea, does not integrate the judicial exception into a practical application (See MPEP 2106.05(f)). Claim 1 also recites additional elements “generating a first data tensor as output from a first layer of a neural network”, “storing the second subtensor”, and “providing the first subtensor to a subsequent layer of the neural network” which amounts to gathering and outputting data which is insignificant extra-solution activity (See MPEP 2106.05(g)). Therefore, claim 1 is directed to a judicial exception.
Step 2B Analysis: Claim 1 does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to the lack of integration of the abstract idea into a practical application, the additional elements recited in claim 1 amount to no more than mere instructions to apply the judicial exception using a generic computer component and insignificant extra-solution activity. The gathering and outputting of data is considered well-understood, routine, and conventional in the art (See MPEP 2106.05(d)(II)(i) and 2106.05(d)(II)(iv)).
For the reasons above, claim 1 is rejected as being directed to non-patentable subject matter under §101. This rejection applies equally to independent claims 11 and 21, which recite a system and a computer program product, respectively, as well as to dependent claims 2-10, 12-20, and 22-29.
Independent claim 11 recites additional instructions to apply the judicial exception using generic computer components “A processing system, comprising: a memory comprising computer-executable instructions; and one or more processors configured to execute the computer-executable instructions and cause the processing system to perform an operation comprising”.
Independent claim 21 recites additional instructions to apply the judicial exception using generic computer components: “A non-transitory computer-readable medium comprising computer-executable instructions that, when executed by one or more processors of a processing system, cause the processing system to perform an operation comprising:”.
The additional limitations of the dependent claims are addressed briefly below:
Dependent claims 2, 12, and 22 recite additional observation, evaluation, and judgement “wherein the tensor splitting operation comprises: a downsampling operation to generate a first intermediate data tensor having a reduced spatial dimensionality, as compared to the first data tensor, and an increased channel depth, as compared to the first data tensor; and a bijective function to delineate the first intermediate data tensor into the first and second subtensors.”
Dependent claims 3, 13, and 23 recite additional observation, evaluation, and judgement “the first subtensor corresponds to a first set of channels from the first intermediate data tensor, the second subtensor corresponds to a second set of channels from the first intermediate data tensor, and the first and second sets of channels are non-overlapping.
Dependent claims 4, 14, and 24 recite additional observation, evaluation, and judgement “the first data tensor has dimensionality B×C×H×W, wherein B is a batch size, C is a channel depth, and H and W are spatial dimensions, and the first intermediate data tensor has dimensionality B×4xC×H2×W2”.
Dependent claims 5, 15, and 25 recite additional observation, evaluation, and judgement “the first and second subtensors each have dimensionality B×2xC×H2×W2.
Dependent claims 6 and 16 recite additional observation, evaluation, and judgement “the first subtensor has dimensionality B×1×H2×W2, and the second subtensor has dimensionality B×(C-1)×H2×W2.”.
Dependent claims 7, 17, and 26 recite additional mathematical calculations and relationships “wherein refining the one or more parameters of the first layer of the neural network comprises: recreating the first subtensor using backpropagation through the subsequent layer of the neural network” (Explicitly disclosed as mathematical calculations and relationships in instant specification at [¶0091] “the recreated input tensor may be generated while computing gradients and updating model parameters during backpropagation. Using the chain rule, the machine learning system can thereby iterate backwards through the model layers, refining each in turn.”). Claims 7, 17, and 26 also recite additional observation, evaluation, and judgement “generating a recreated first data tensor by combining the recreated first subtensor and the stored second subtensor; and generating a recreated first data tensor by combining the recreated first subtensor and the stored second subtensor refining the one or more parameters of the first layer using backpropagation of the recreated first data tensor”
Dependent claims 8, 18, and 27 recite additional observation, evaluation, and judgement “performs an invertible operation.” And mere instructions to apply the judicial exception using generic computer components “the first layer of the neural network”
Dependent claims 9, 19, and 28 recite additional observation, evaluation, and judgement “generates the first data tensor based on a first input data tensor and a second input data tensor; the first data tensor comprises the first input data tensor and a third data tensor, and the third data tensor data tensor comprises a non-linear combination of the first and second input data tensors” And mere instructions to apply the judicial exception using generic computer components “the first layer of the neural network”
Dependent claims 10, 20, and 29 recite additional observation, evaluation, and judgement “applying the tensor splitting operation after each layer of a plurality of layers of the neural network”
Regarding Claim 30: Claim 30 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1 Analysis: Claim 30 is directed to a method, which is directed to a process, one of the statutory categories.
Step 2A Prong One Analysis: Claim 30 under its broadest reasonable interpretation is a series of mental processes and mathematical calculations and relationships. For example, but for the generic computer components language, the above limitations in the context of this claim encompass neural network processing, including the following:
Generating, using the layer of the neural network, a first output data tensor and a second output data tensor by processing the first input data tensor and the second input data tensor (observation, evaluation, and judgement),
the first output data tensor is equal to the first input data tensor (observation, evaluation, and judgement)
the second output data tensor is generated by: applying one or more convolution operations to the first input data tensor (mathematical calculations and relationships)
applying a multiplication operation based on a summation of an output of the one or more convolution operations and the second input data tensor to generate the second output data tensor (mathematical calculations and relationships)
the first and second input data tensors can be reconstructed based on the first and second output data tensors (observation, evaluation, and judgement)
generating an inference based on the first and second output data tensors (observation, evaluation, and judgement)
Therefore, claim 30 recites an abstract idea which is a judicial exception.
Step 2A Prong Two Analysis: Claim 30 recites additional elements “using the layer of the neural network”. However, these additional features are computer components recited at a high-level of generality, such that they amount to no more than mere instructions to apply the judicial exception using a generic computer component. An additional element that merely recites the words “apply it” (or an equivalent) with the judicial exception, or merely includes instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea, does not integrate the judicial exception into a practical application (See MPEP 2106.05(f)). Claim 30 also recites additional elements “accessing, at a layer of a neural network, a first input data tensor and a second input data tensor, the second input data tensor being different from the first input data tensor” and “outputting the first and second output data tensors from the layer of the neural network” which amounts to gathering and outputting data which is insignificant extra-solution activity (See MPEP 2106.05(g)). Therefore, claim 30 is directed to a judicial exception.
Step 2B Analysis: Claim 30 does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to the lack of integration of the abstract idea into a practical application, the additional elements recited in claim 30 amount to no more than mere instructions to apply the judicial exception using a generic computer component and insignificant extra-solution activity. The gathering and outputting of data is considered well-understood, routine, and conventional in the art (See MPEP 2106.05(d)(II)(i) and 2106.05(d)(II)(iv)).
For the reasons above, claim 30 is rejected as being directed to non-patentable subject matter under §101.
Therefore, when considering the elements separately and in combination, they do not add significantly more to the inventive concept. Accordingly, claims 1-30 are rejected under 35 U.S.C. § 101.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 1-30 are rejected under U.S.C. §103 as being unpatentable over the combination of Dinh (“DENSITY ESTIMATION USING REAL NVP”, 2017) and Jiang (“Back Razor: Memory-Efficient Transfer Learning by Self-Sparsified Backpropagation”, 2022).
PNG
media_image1.png
310
660
media_image1.png
Greyscale
FIG. 2 of Dinh
Regarding claim 1, Dinh teaches generating a first data tensor as output from a first layer of a neural network;([p. 5] "The squeezing operation reduces the 4 × 4 × 1 tensor (on the left) into a 2 × 2 × 4 tensor (on the right)" [p. 6] "The sequence of coupling-squeezing-coupling operations described above is performed per layer when computing f(i) (Equation 14). At each layer, as the spatial resolution is reduced, the number of hidden layer features in s and t is doubled. All variables which have been factored out at different scales are concatenated to obtain the final transformed output (Equation 16)." [p. 7] "We use the multi-scale architecture described in Section 3.6 and use deep convolutional residual networks in the coupling layers with rectifier nonlinearity and skip-connections" Dinh's architecture is explicitly layered where the claims "first data tensor" maps to the intermediate tensor representation generated at a layer/scale of this multi-scale architecture before it is squeezed, partitioned, and factored.)
separating the first data tensor into a first subtensor and a second subtensor using a tensor splitting operation;([p. 5] "Partitioning can be implemented using a binary mask b, and using the functional form for y […] We use two partitionings that exploit the local correlation structure of images: spatial checkerboard patterns, and channel-wise masking" [p. 6] "Propagating a D dimensional vector through all the coupling layers would be cumbersome, in terms of computational and memory cost, and in terms of the number of parameters that would need to be trained. For this reason we follow the design choice of [57] and factor out half of the dimensions at regular intervals" [p. 6] "Factoring out variables. At each step, half the variables are directly modeled as Gaussians, while the other half undergo further transformation [...] for each channel, it divides the image into subsquares of shape 2 × 2 × c" Dinh discloses partitioning by binary mask and, in the multi-scale architecture, factoring out half the dimensions in recursive intervals. The "first subtensor" maps to the half that "undergo further transformation," and the "second subtensor" maps to the half that is "directly modeled as Gaussians" / factored out.)
storing the second subtensor;([p. 6] "All variables which have been factored out at different scales are concatenated to obtain the final
transformed output (Equation 16)." For the factored out variables to be concatenated and used in the final latent representation they must necessarily be retained/stored)
providing the first subtensor to a subsequent layer of the neural network; and([p. 6] "Factoring out variables. At each step, half the variables are directly modeled as Gaussians, while the other half undergo further transformation." [p. 6] "Propagating a D dimensional vector through all the coupling layers would be cumbersome, in terms of computational and memory cost, and in terms of the number of parameters that would need to be trained. For this reason we follow the design choice of [57] and factor out half of the dimensions at regular intervals (see Equation 14). We can define this operation recursively" The half that undergo further transformation (first subtensor) are provided to a subsequent layer of the neural network)
refining one or more parameters of the first layer of the neural network based at least in part on the stored second subtensor.([p. 3] "we will tackle the problem of learning highly nonlinear models in high-dimensional continuous spaces through maximum likelihood. In order to optimize the log-likelihood, we introduce a more flexible class of architectures that enables the computation of log-likelihood on continuous data using the change of variable formula. Building on our previous work in [17], we define a powerful class of bijective functions which enable exact and tractable density evaluation and exact and tractable inference. Moreover, the resulting cost function does not to rely on a fixed form reconstruction cost such as square error [38, 47], and generates sharper samples as a result." [p. 6] "Gaussianizing and factoring out units in earlier layers has the practical benefit of distributing the loss function throughout the network" [p. 7] "We optimize with ADAM [33] with default hyperparameters and use an L2 regularization on the weight scale parameters with coefficient" Dinh trains the Real NVP model through maximum likelihood optimization and explicitly states that earlier factored-out units distribute the loss throughout the network. Because the factored-out subtensor is retained and concatenated into the final transformed output, it contributes to the loss. That loss is then optimized with ADAM, refining trainable parameters of the network including parameters of earlier coupling layers.)
generating an inference using the neural network with the one or more refined parameters of the first layer of the neural network.([p. 1] "This model can perform efficient and exact inference" [p. 2] "Figure 1: Real NVP learns an invertible, stable, mapping between a data distribution ^pX and a latent distribution pZ (typically a Gaussian). Here we show a mapping that has been learned on a toy 2-d dataset. The function f (x) maps samples x from the data distribution in the upper left into approximate samples z from the latent distribution, in the upper right. This corresponds to exact inference of the latent state given the data").
However, Dinh does not explicitly teach A computer-implemented method, comprising:
refining one or more parameters of the first layer of the neural network based at least in part on the stored second subtensor.
Jiang, in the same field of endeavor, teaches A computer-implemented method, comprising: ([p. 6] "Our experiments are implemented with Pytorch [50] and conducted on 1080 Ti or V100 GPUs. To measure the training memory, by default we report the theoretical memory usage following [1]. We also report the actual memory usage measured on the GPU at Section 4.3")
refining one or more parameters of the first layer of the neural network based at least in part on the stored second subtensor. ([p. 4] "L is the loss between prediction f(x;θ) and label y calculated by loss function L. fi and θi denote the function and parameters of i th layer, respectively. The network is composed of l layers. We further denote the output of i th layer (a.k.a activation) as zi = fi(zi−1,θi) and z1 = f1 (x;θ1). When conducting back-propagation, ∂L ∂zi and ∂L ∂θi are required to be calculated for each activation [...] As shown in Equation 2, the activation zi is employed for computing one of the gradients. This creates the need of saving the activation of the forward process. Otherwise, extra FLOPs are required for re-computing the activation").
Dinh as well as Jiang are directed towards neural network training through gradient propagation. Therefore, Dinh as well as Jiang are reasonably pertinent analogous art. It would have been obvious before the effective filing date of the claimed invention to combine the teachings of Dinh with the teachings of Jiang by storing the activation for backpropagating the gradient. Jiang provides as additional motivation for combination ([p. 4] “We inherit the mathematical forms behind h(z) i (4) and h(θ) i from the normal backward to preserve the gradient flow, but calculate the gradients on weights and activations with the pruned layerwise activation ˜zi. With the above backward, there is no need for saving the dense activation zi. Instead, we store a sparse version ˜zi that is more memory efficient”). This motivation for combination also applies to the remaining claims which depend on this combination.
Regarding claim 2, the combination of Dinh, and Jiang teaches The computer-implemented method of claim 1, wherein the tensor splitting operation comprises: a downsampling operation to generate a first intermediate data tensor having a reduced spatial dimensionality, as compared to the first data tensor, and an increased channel depth, as compared to the first data tensor; and(Dinh [p. 6] "The squeezing operation transforms an sxsxc tensor into an s/2xs/2x4c tensor (See Figure 3)")
a bijective function to delineate the first intermediate data tensor into the first and second subtensors.(Dinh [p. 4 §3.2] "We will build a flexible and tractable bijective function by stacking a sequence of simple bijections. In each simple bijection, part of the input vector is updated using a function which is simple to invert, but which depends on the remainder of the input vector in a complex way. We refer to each of these simple bijections as an affine coupling layer.").
Regarding claim 3, the combination of Dinh, and Jiang teaches The computer-implemented method of claim 2, wherein: the first subtensor corresponds to a first set of channels from the first intermediate tensor, the second subtensor corresponds to a second set of channels from the first intermediate tensor, and the first and second sets of channels are non-overlapping.(Dinh [p. 5] "On the right, a channel-wise masking. The squeezing operation reduces the tensor (on the left) into a tensor (on the right). Before the squeezing operation, a checkerboard pattern is used for coupling layers while a channel-wise masking pattern is used afterward [...] The channel-wise mask b is 1 for the first half of the channel dimensions and 0 for the second half.").
Regarding claim 4, the combination of Dinh, and Jiang teaches The computer-implemented method of claim 2, wherein: the first data tensor has dimensionality B×C×H×W, wherein B is a batch size, C is a channel depth, and H and W are spatial dimensions, and the first intermediate data tensor has dimensionality B×4C×H2×W2.(Dinh [p. 6 §3.6] "We implement a multi-scale architecture using a squeezing operation: for each channel, it divides the image into subsquares of shape 2 × 2 × c, then reshapes them into subsquares of shape 1 × 1 × 4c. The squeezing operation transforms an s × s × c tensor into an s2 ×s2 × 4c tensor (see Figure 3), effectively trading spatial size for number of channels" See FIG. 3 where first data tensor has B=1, C=1, H=4, W=4 and intermediate data tensors have B=1, C=4, H2=2, W2=2).
Regarding claim 5, the combination of Dinh, and Jiang teaches The computer-implemented method of claim 4, wherein the first and second subtensors each have dimensionality B×2C×H2×W2.(Dinh [p. 6 §3.6] "We implement a multi-scale architecture using a squeezing operation: for each channel, it divides the image into subsquares of shape 2 × 2 × c, then reshapes them into subsquares of shape 1 × 1 × 4c. The squeezing operation transforms an s × s × c tensor into an s2 ×s2 × 4c tensor (see Figure 3), effectively trading spatial size for number of channels" FIG. 3 shows two subsets (white and black) of the input tensor, the respective subsets have dimensions (B=1, C=2, H2=2, W2=2)).
Regarding claim 6, the combination of Dinh, and Jiang teaches The method of claim 4, wherein: the first subtensor has dimensionality B×1×H2×W2, and the second subtensor has dimensionality B×(C-1)×H2×W2.(Dinh [p. 6 §3.6] "We implement a multi-scale architecture using a squeezing operation: for each channel, it divides the image into subsquares of shape 2 × 2 × c, then reshapes them into subsquares of shape 1 × 1 × 4c. The squeezing operation transforms an s × s × c tensor into an s2 ×s2 × 4c tensor (see Figure 3), effectively trading spatial size for number of channels" See FIG. 3. C-1 interpreted as zero such that second subset is empty set and any of the remaining subsets are interpreted as the first subtensor.).
Regarding claim 7, the combination of Dinh, and Jiang teaches The computer-implemented method of claim 1, wherein refining the one or more parameters of the first layer of the neural network comprises: recreating the first subtensor using backpropagation through the subsequent layer of the neural network;(Dinh [p. 1] "the architecture presented in this paper enables exact and efficient reconstruction of input images from the hierarchical features extracted by this model" [p. 4] "Computational graphs for forward and inverse propagation. A coupling layer applies a simple invertible transformation consisting of scaling followed by addition of a constant offset to one part x2 of the input vector conditioned on the remaining part of the input vector x1. Because of its simple nature, this transformation is both easily invertible and possesses a tractable determinant" Dinh explicitly discloses inverse propagation through invertible coupling layers. Dinh's FIG. 2 is explicitly directed to "forward and inverse propagation" through coupling layers that are "easily invertible" to obtain "exact and efficient reconstruction".)
generating a recreated first data tensor by combining the recreated first subtensor and the stored second subtensor; and(Dinh [p. 1] "the architecture presented in this paper enables exact and efficient reconstruction of input images from the hierarchical features extracted by this model" [p. 6] "All variables which have been factored out at different scales are concatenated to obtain the final transformed output (Equation 16)." [p. 9] "we have defined a class of invertible functions with tractable Jacobian determinant, enabling exact and tractable log-likelihood evaluation, inference, and sampling" Dinh states that variables are split with half being factored out and half continuing through further transformation. The factored variables are concatenated to obtain the final transformed output. The inverse of this recursive structure combines the factored out variables with the inverse-propagated continuing variables to reconstruct the earlier hierarchical representation. See also FIG. 2(b))
generating a recreated first data tensor by combining the recreated first subtensor and the stored second subtensor(Dinh [p. 4 §3.2] "We will build a flexible and tractable bijective function by stacking a sequence of simple bijections. In each simple bijection, part of the input vector is updated using a function which is simple to invert, but which depends on the remainder of the input vector in a complex way. We refer to each of these simple bijections as an affine coupling layer." See FIG. 2 and eqn. 4 and 5)
refining the one or more parameters of the first layer using backpropagation of the recreated first data tensor.(Jiang [p. 4] "L is the loss between prediction f(x;θ) and label y calculated by loss function L. fi and θi denote the function and parameters of i th layer, respectively. The network is composed of l layers. We further denote the output of i th layer (a.k.a activation) as zi = fi(zi−1,θi) and z1 = f1 (x;θ1). When conducting back-propagation, ∂L ∂zi and ∂L ∂θi are required to be calculated for each activation [...] As shown in Equation 2, the activation zi is employed for computing one of the gradients. This creates the need of saving the activation of the forward process. Otherwise, extra FLOPs are required for re-computing the activation").
Regarding claim 8, the combination of Dinh, and Jiang teaches The computer-implemented method of claim 1, wherein the first layer of the neural network performs an invertible operation.(Dinh [p. 2] "Figure 1: Real NVP learns an invertible, stable, mapping between a data distribution ^pX and a latent distribution pZ (typically a Gaussian)." See FIG. 2(b)).
Regarding claim 9, the combination of Dinh, and Jiang teaches The computer-implemented method of claim 8, wherein: the first layer of the neural network generates the first data tensor based on a first input data tensor and a second input data tensor; (Dinh [p. 4] "coupling layer applies a simple invertible transformation consisting of scaling followed by addition of a constant offset to one part x2 of the input vector conditioned on the remaining part of the input vector x1" Dinh's affine coupling layer partitions an input into two parts, x1:d and xd+1:d. The output tensor y is generated from those two input portions: one portion passes through, and the other is transformed based on functions of the first portion. This corresponds to generating the claimed first data tensor based on a first input tensor and second input tensor.)
the first data tensor comprises the first input data tensor and a third data tensor, (Dinh "y1:d = x1:d (4) yd+1:D = xd+1:D […] + t(x1:d)" Dinh's coupling-layer output includes y1:d = x1:d, so the first input portion is literally included unchanged in the output. The other output portion, yd+1:D, is the claimed "third data tensor" such that the output tensor comprises the first input tensor and a transformed third tensor)
and the third data tensor comprises a non-linear combination of the first and second input data tensors.(Dinh [p. 4] "y1:D = x1:D * exp(s(x1:d))+t(x1:d) […] s and t stand for scale and translation […] Since computing the Jacobian determinant of the coupling layer operation does not involve computing the Jacobian of s or t, those functions can be arbitrarily complex. We will make them deep convolutional neural networks. Note that the hidden layers of s and t can have more features than their input and output layers." Dinh's transformed output portion is xd+1:D multiplied by exp(s(x1:d)) and shifted by t(x1:d). Because s and t are deep convolutional neural networks the transformed portion is a nonlinear function of the first input portion and also directly includes the second input portion.).
Regarding claim 10, the combination of Dinh, and Jiang teaches The computer-implemented method of claim 1, further comprising applying the tensor splitting operation after each layer of a plurality of layers of the neural network.(Dinh [p. 6] "factor out half of the dimensions at regular intervals (see Equation 14). We can define this operation recursively […] The sequence of coupling-squeezing-coupling operations described above is performed per layer when computing f(i) (Equation 14). At each layer, as the spatial resolution is reduced, the number of hidden layer features in s and t is doubled" Dinh applies factor-out recursively in the multi-scale architecture: at each step, half the variables are factored out and the other half continues. Dinh explicitly says the coupling-squeezing-coupling sequence is "performed per layer" and that variables factored out at different scales are concatenated into the final output).
Regarding claim 11, the combination of Dinh, and Jiang teaches same as claim 1(Dinh same as claim 1)
A processing system, comprising: a memory comprising computer-executable instructions; and
one or more processors configured to execute the computer-executable instructions and cause the processing system to perform an operation comprising:(Jiang [p. 6] "Our experiments are implemented with Pytorch [50] and conducted on 1080 Ti or V100 GPUs. To measure the training memory, by default we report the theoretical memory usage following [1]. We also report the actual memory usage measured on the GPU at Section 4.3").
Regarding claims 12-20, claims 12-20 are directed towards a system for performing the method of claims 2-10, respectively. Therefore, the rejections applied to claims 2-10 also apply to claims 12-20.
Regarding claim 21, claim 21 is directed towards a computer program product for performing the method of claim 1. Therefore, the rejection applied to claim 1 also applies to claim 21. Claim 21 also recites additional elements A non-transitory computer-readable medium comprising computer-executable instructions that, when executed by one or more processors of a processing system, cause the processing system to perform an operation comprising:(Jiang [p. 6] "Our experiments are implemented with Pytorch [50] and conducted on 1080 Ti or V100 GPUs. To measure the training memory, by default we report the theoretical memory usage following [1]. We also report the actual memory usage measured on the GPU at Section 4.3").
Similarly, regarding claims 22-29, claims 22-29 are directed towards a computer program product for performing the methods of claims 2-10, respectively. Therefore, the rejection applied to claims 2-10 also apply to claims 22-29.
Regarding claim 30, Dinh teaches accessing, at a layer of a neural network, a first input data tensor and a second input data tensor;([p. 4] "A coupling layer applies a simple invertible transformation consisting of scaling followed by addition of a constant offset to one part x2 of the input vector conditioned on the remaining part of the input vector x1" See FIG. 2 and 4 and Eqn. 5-8 which each take two input tensors into a coupling layer.)
the second input data tensor being different from the first input data tensor([p. 4] "A coupling layer applies a simple invertible transformation consisting of scaling followed by addition of a constant offset to one part x2 of the input vector conditioned on the remaining part of the input vector x1" First part of input vector x1 interpreted as first input tensor, second part of input vector x1 interpreted as second input data tensor. Both input data tensors are explicitly input into the coupling layer.)
generating, using the layer of the neural network, a first output data tensor and a second output data tensor by processing the first input data tensor and the second input data tensor, ([p. 5 §3.5] "coupling layers can be powerful, their forward transformation leaves some components unchanged. This difficulty can be overcome by composing coupling layers in an alternating pattern, such that the components that are left unchanged in one coupling layer are updated in the next (see Figure 4(a))." [p. 6] "In this alternating pattern, units which remain identical in one transformation are modified in the next." See FIG. 2 and FIG. 4a.)
wherein: the first output data tensor is equal to the first input data tensor;([p. 5] "y1:d = x1:d (7)")
the second output data tensor is generated by: applying one or more convolution operations to the first input data tensor;([p. 5] "For the models presented here, both s(·) and t(·) are rectified convolutional networks" [p. 7 §4.1] "We use the multi-scale architecture described in Section 3.6 and use deep convolutional residual networks in the coupling layers with rectifier nonlinearity and skip-connections as suggested by [46]." Dinh explicitly states that the models are convolutional networks such that all operations are interpreted as convolution operations)
applying a multiplication operation based on a summation of an output of the one or more convolution operations and the second input data tensor to generate the second output data tensor; and([p. 4] "y1:D = x1:D * exp(s(x1:d))+t(x1:d) […] s and t stand for scale and translation […] Since computing the Jacobian determinant of the coupling layer operation does not involve computing the Jacobian of s or t, those functions can be arbitrarily complex. We will make them deep convolutional neural networks. Note that the hidden layers of s and t can have more features than their input and output layers." [p. 5] "For the models presented here, both s(·) and t(·) are rectified convolutional networks" First, Dinh's affine coupling layer generates the second output component by multiplying the second input component by a convolution-derived scale term and summing the result with a convolution derived translation term; thus, the second output is generated using a multiplication operation and a summation involving outputs of convolutional subnetworks applied to the first input tensor. Additionally, Dinh's formula y1:D = x1:D * exp(s(x1:d))+t(x1:d) is mathematically equivalent (by factoring out the scale term) to y2 = exp(s(x1:d))*(xd+1:D+t(x1:d)*exp(-s)x1:d)) where x1:d is the first input data tensor, xd+1:D is the second input data tensor, s(x1) is the convolutional scale output, t(x1) is the convolutional translation output, and * is element-wise multiplication.)
the first and second input data tensors can be reconstructed based on the first and second output data tensors; and([p. 4 §3.2] "We will build a flexible and tractable bijective function by stacking a sequence of simple bijections. In each simple bijection, part of the input vector is updated using a function which is simple to invert, but which depends on the remainder of the input vector in a complex way. We refer to each of these simple bijections as an affine coupling layer." See FIG. 2 and 5 and eqn. 4-8)
outputting the first and second output data tensors from the layer of the neural network.(See FIG. 2 and 5 and eqn. 4-8)
and generating an inference based on the first and second output data tensors.([Abstract] "Unsupervised learning of probabilistic models is a central yet challenging problem in machine learning. Specifically, designing models with tractable learning, sampling, inference and evaluation is crucial in solving this task" [p. 3] "See also Figure 1. Exact and efficient inference enables the accurate and fast evaluation of the model" [p. 6] "the number of hidden layer features in s and t is doubled. All variables which have been factored out at different scales are concatenated to obtain the final transformed output").
However, Dinh does not explicitly teach A computer-implemented method, comprising:.
Jiang, in the same field of endeavor, teaches A computer-implemented method, comprising: ([p. 6] "Our experiments are implemented with Pytorch [50] and conducted on 1080 Ti or V100 GPUs. To measure the training memory, by default we report the theoretical memory usage following [1]. We also report the actual memory usage measured on the GPU at Section 4.3").
Dinh as well as Jiang are directed towards neural network training through gradient propagation. Therefore, Dinh as well as Jiang are reasonably pertinent analogous art. It would have been obvious before the effective filing date of the claimed invention to combine the teachings of Dinh with the teachings of Jiang by storing the activation for backpropagating the gradient. Jiang provides as additional motivation for combination ([p. 4] “We inherit the mathematical forms behind h(z) i (4) and h(θ) i from the normal backward to preserve the gradient flow, but calculate the gradients on weights and activations with the pruned layerwise activation ˜zi. With the above backward, there is no need for saving the dense activation zi. Instead, we store a sparse version ˜zi that is more memory efficient”). This motivation for combination also applies to the remaining claims which depend on this combination.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Lugmayr (“SRFlow: Learning the Super-Resolution Space with Normalizing Flow”, 2022) is directed towards normalizing flows with coupling networks that split tensors. Similarly, Kingma (“Glow: Generative Flow with Invertible 1×1 Convolutions”, 2018) explicitly builds upon Dinhs coupling networks.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SIDNEY VINCENT BOSTWICK whose telephone number is (571)272-4720. The examiner can normally be reached M-F 7:30am-5:00pm EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Miranda Huang can be reached on (571)270-7092. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/SIDNEY VINCENT BOSTWICK/Examiner, Art Unit 2124