Prosecution Insights
Last updated: October 02, 2026
Application No. 18/634,466

DATA AUGMENTATION FOR TRAINING NEURAL NETWORKS

Non-Final OA §102§103§112
Filed
Apr 12, 2024
Priority
May 25, 2023 — provisional 63/504,395
Examiner
BASOM, BLAINE T
Art Unit
Tech Center
Assignee
NAVER Corporation
OA Round
1 (Non-Final)
43%
Grant Probability
Moderate
1-2
OA Rounds
2y 0m
Est. Remaining
64%
With Interview

Examiner Intelligence

Grants 43% of resolved cases
43%
Career Allowance Rate
146 granted / 338 resolved
-16.8% vs TC avg
Strong +21% interview lift
Without
With
+20.8%
Interview Lift
resolved cases with interview
Typical timeline
4y 6m
Avg Prosecution
23 currently pending
Career history
369
Total Applications
across all art units

Statute-Specific Performance

§101
8.2%
-31.8% vs TC avg
§103
60.9%
+20.9% vs TC avg
§102
10.8%
-29.2% vs TC avg
§112
13.2%
-26.8% vs TC avg
Black line = Tech Center average estimate • Based on career data from 338 resolved cases

Office Action

§102 §103 §112
DETAILED ACTION This Office Action is responsive to the Applicant’s preliminary amendment filed on April 12, 2024. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Information Disclosure Statement The information disclosure statement (IDS) submitted on February 7, 2025 has been considered by the examiner. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 21, 23, 24, 30, 32 and 33 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Particularly, in claim 21, there is improper antecedent basis for “the training” recited therein. It is unclear as to whether “the training” is intended to refer to “training the data augmentation policy” or “training the neural network,” both of which are recited in claim 1, upon which claim 21 depends. Similarly, in claim 23, there is improper antecedent basis for “the training” recited therein, as it is unclear as to whether “the training” is intended to refer to the recitation in claim 1 of “training the data augmentation policy” or “training the neural network.” As per claim 24, there is no antecedent basis for “the data” recited therein. Claim 24 depends from claim 1, which recites “a data augmentation policy” and “a training dataset” but not “data” per se. As per claim 30, there is no antecedent basis for “the evaluation dataset” recited therein. Claim 30 previously recites “a validation dataset” but not an “evaluation dataset.” As per claim 32, there is no antecedent basis for “the data” recited therein. Claim 32 depends from claim 1, which recites “a data augmentation policy” and “a training dataset” but not “data” per se. Also in claim 32, it is unclear as to whether “the data augmentation” is intended to refer to the recitation of “the training dataset being augmented by an initial augmentation policy” or “the training dataset being augmented by the current data augmentation policy,” both of which are recited in claim 1. Similarly, in claim 33, there is no antecedent basis for “the data” recited therein. Claim 33 depends from claim 1, which recites “a data augmentation policy” and “a training dataset” but not “data” per se. Also in claim 33, it is unclear as to whether “the data augmentation” is intended to refer to the recitation of “the training dataset being augmented by an initial augmentation policy” or “the training dataset being augmented by the current data augmentation policy,” both of which are recited in claim 1. Claim Rejections - 35 USC § 102 The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. Claims 1, 6-8, 17, 24, 25, 27, 28, 33, 38 and 39 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by the article entitled “Improving Auto-Augment via Augmentation-Wise Weight Sharing” by Tian et al. (“Tian”). Regarding claims 1, 38 and 39, Tian generally teaches using “Augmentation-Wise Weight Sharing” (AWS) to form an evaluation process for augmentation polices, and which is used while automatically searching augmentation policies (see e.g. the Abstract). Like claimed, Tian particularly teaches: pretraining a neural network having neural network parameters on a task on a training dataset, the training dataset being augmented by an initial augmentation policy (see e.g. section 3.3 “Our Proxy Task” on pages 4-5: Tian teaches partitioning the augmented training of neural network parameters into two parts: a first part in which a shared augmentation policy is applied to train a neural network and thereby obtain shared weights w s h a r e for the neural network; and a second part in which the neural network model is fine-tuned from the shared weights by a given data augmentation policy so that the policy can evaluated. The first part trains the neural network on a task on a training dataset D t r that is augmented by the shared augmentation policy – see e.g. section 3.3 “Our Proxy Task” on pages 4-5 and the line “Obtain w s h a r e in Equ. 3;” in Algorithm 1 on page 6. The first part is considered pretraining a neural network having network parameters on a task on a training dataset, the training dataset being augmented by an initial augmentation policy, i.e. by the shared augmentation policy.); initializing the data augmentation policy with the initial data augmentation policy to define a current data augmentation policy (see e.g. section 3.3 “Our Proxy Task” on pages 4-5: as noted above, Tian teaches partitioning the augmented training of neural network parameters into two parts: a first part in which a shared augmentation policy is applied to train a neural network and obtain shared weights w s h a r e for the neural network; and a second part in which the neural network is fine-tuned from the shared weights by a given data augmentation policy so that the policy can evaluated. The second part particularly comprises an iterative process in which, for each iteration: (i) the shared weights w s h a r e are loaded; (ii) the neural network is fined-tuned from the shared weights w s h a r e with the training dataset D t r modified by a current data augmentation policy; and (iii) the current data augmentation policy is updated based on the accuracy ACC of the fine-tuned neural network on a validation dataset D v a l – see e.g. section 3 “Method” on pages 3-6, and the following lines in Algorithm 1 on page 6: while T ≤ T m a x do Load w s h a r e ; Fine-tune w s h a r e to get w - θ * ; Use ACC ( w - θ * , D v a l ) to update θ ; end while The augmentation policy used in the first part and prior to being first updated during the iterative process in the second part is considered a current data augmentation policy which is defined by initializing the data augmentation policy with the initial data augmentation policy, i.e. with the shared data augmentation policy.); and iteratively training the data augmentation policy using bilevel optimization, wherein said iteratively training comprises, for each of n rounds, where n ≥ 1 (see e.g. section 3 “Method” on pages 3-6, and Algorithm 1 on page 6: as noted above, Tian describes an iterative process in which, for each iteration: (i) the shared weights w s h a r e are loaded; (ii) the neural network is fined-tuned from the shared weights w s h a r e with the training dataset D t r modified by a current data augmentation policy; and (iii) the current data augmentation policy is updated based on the accuracy ACC of the fine-tuned neural network on a validation dataset D v a l . Each iteration is considered a round like claimed, which is used to train the augmentation policy using bilevel optimization. Tian discloses that there are T m a x iterations – see e.g. section 3.3 “Our Proxy Task” on page 5 and the line “while T ≤ T m a x do” in Algorithm 1 on page 6 – and thus teaches n rounds where n ≥ 1 .): initializing the neural network parameters of the neural network with the neural network parameters trained during said pretraining (see e.g. section 3 “Method” on pages 3-6, and Algorithm 1 on page 6: as noted above, Tian describes an iterative process in which, for each iteration, the shared weights w s h a r e are first loaded. The neural network parameters, i.e. weights, of the neural network are thus initialized in each round with the neural network parameters w s h a r e , which are trained during pretraining.); and over a plurality of steps, training the neural network on the task to update the neural network parameters on the training dataset, the training dataset being augmented by the current data augmentation policy for the current step (see e.g. section 3 “Method” on pages 3-6, and Algorithm 1 on page 6: as noted above, Tian describes an iterative process in which each iteration comprises fine-tuning the neural network from the shared weights w s h a r e with the training dataset D t r modified by a current data augmentation policy. Such fine-tuning occurs over a plurality of epochs – see e.g. section 3.3 “Our Proxy Task” on page 5, and section 4.2 “Implementation Details” on page 6. The fine-tuning in each iteration is considered a step (or more) of training the neural network on the task to update the neural network parameters on the training dataset, wherein the training dataset is augmented by the current data augmentation policy for the current step.); and updating the data augmentation policy based on said updated neural network parameters to define the current data augmentation policy for the next step or the data augmentation policy on the last round (see e.g. section 3 “Method” on pages 3-6, and Algorithm 1 on page 6: as noted above, Tian describes an iterative process in which each iteration comprises updating the current data augmentation policy based on the accuracy ACC of the fine-tuned neural network on a validation dataset D v a l . The updated data augmentation policy is then used in the next step, i.e. within a next iteration, or is returned if T > T m a x – see e.g. section 3 “Method” on pages 3-6, and Algorithm 1 on page 6. Accordingly, Tian teaches updating the data augmentation policy based on the updated neural network parameters, i.e. based on the fine-tuned neural network, to define the current augmentation policy for the next step or the data augmentation policy on the last round, i.e. the last iteration.). Accordingly, Tian teaches a computer-implemented method like that of claim 1, which is for training a data augmentation policy. The method is understandably implemented via executable instructions stored in the memory of a computer system that further comprises a processor (e.g. a GPU) for executing the instructions (see e.g. “Results on computational cost” on page 8 of Tian). Such a computer system comprising a processor, a memory and executable instructions stored in the memory for execution by the processor to perform the above-described tasks taught by Tian is considered a computer-implemented system like that of claim 38. The memory of such a system is considered an apparatus like that of claim 39. As per claim 6, Tian further teaches that updating the data augmentation policy comprises training the data augmentation policy using the trained neural network with the updated neural network parameters on a validation dataset that is separate from the training dataset, without data augmentation, to update data augmentation policy parameters of the data augmentation policy (see e.g. section 3 “Method” on pages 3-6, and Algorithm 1 on page 6: as noted above, Tian describes an iterative process in which each of a plurality of iterations comprises updating a current data augmentation policy based on the accuracy ACC of the fine-tuned neural network on a validation dataset D v a l . The validation dataset D v a l is separate from the training dataset D t r used to fine-tune the neural network, and data augmentation is not applied to the validation dataset D v a l – see e.g. section 3 “Method” on pages 3-6. The data augmentation policy is thus trained, in part, by using the trained neural network with the updated neural network parameters, i.e. by using the fine-tuned neural network, on a validation dataset that is separate from the training dataset, without data augmentation, to update the data augmentation policy parameters of the data augmentation policy.). Accordingly, Tian further teaches a method like that of claim 6. As per claim 7, Tian further teaches that n > 1 (see e.g. section 3.3 “Our Proxy Task” on page 5 and the line “while T ≤ T m a x do” in Algorithm 1 on page 6: as noted above, Tian discloses that there are T m a x iterations, wherein each iteration is considered a round like claimed. Tian further discloses that T m a x   > 1 – see e.g. section 4.2 “Implementation Details.”). Accordingly, Tian further teaches a method like that of claim 7. As per claim 8, Tian further discloses that the initial augmentation policy comprises a uniform policy (see e.g. section 3.3 “Our Proxy Task” on page 5: Tian discloses that the shared augmentation policy comprises a uniform sampling of augmentation transforms. As noted above, the shared augmentation policy is considered an initial augmentation policy like claimed. The uniform sampling of augmentation transforms described by Tian is considered a uniform policy like claimed.). Accordingly, Tian further teaches a method like that of claim 8. As per claim 17, Tian further teaches that iteratively training the data augmentation policy comprises, over an additional plurality of steps, training the neural network on the task to update the neural network parameters on the training dataset without updating the data augmentation policy, the training dataset being augmented by the current data augmentation policy for the current step (see e.g. section 3 “Method” on pages 3-6, and Algorithm 1 on page 6: as noted above, Tian describes an iterative process in which each iteration comprises: (i) loading the shared weights w s h a r e ; (ii) fine-tuning the neural network from the shared weights w s h a r e with the training dataset D t r modified by a current data augmentation policy; and (iii) updating the current data augmentation policy based on the accuracy ACC of the fine-tuned neural network on a validation dataset D v a l . Tian discloses that the fine-tuning can occur over a number of steps, e.g. over a number of epochs – see e.g. section 3.3 “Our Proxy Task” on page 5, and section 4.2 “Implementation Details” on page 6. Iteratively training the data augmentation policy would thus comprise, over an additional plurality of steps, e.g. epochs, fine-tuning the neural network on the task to update the neural network parameters on the training dataset without updating the data augmentation policy, wherein the training dataset is augmented by the current data augmentation policy for the current step.). Accordingly, Tian further teaches a method like that of claim 17. As per claim 24, Tian teaches that the data can comprise image data and that the data augmentation can comprise transforming the image data (see e.g. section 1 “Introduction,” section 3.2 “Auto-Aug Formulation” and section 3.4 “Augmentation Policy Space and Search Pipeline”). Accordingly, Tian further teaches a method like that of claim 24. As per claim 25, Tian teaches that the task can comprise an image classification task (see e.g. section 1 “Introduction” and section 4.1 “Datasets and Comparison Methods”). Accordingly, Tian further teaches a method like that of claim 25. As per claim 27, Tian further teaches generating augmented data from a dataset using the trained data augmentation policy, and training the neural network or a different neural network on the task using the augmented data (see e.g. section 4.1 “Datasets and Comparison Methods,” section 4.2 “Implementation Details,” and section 4.3 “Comparison with state-of-the-arts:” Tian discloses that on a dataset such as CIFAR-10, a neural network such as ResNet-18 can be used to search data augmentation polices according to the above-described process, wherein the searched policy can then be transferred to another model such as Shake-Shake. This would entail generating augmented data from the CIFAR-10 dataset using the trained augmentation policy, i.e. the searched augmentation policy, and training the neural network such as ResNet-18 and/or a different neural network such as Shake-Shake on the task using the augmented data.). Accordingly, Tian further teaches a method like that of claim 27. As per claim 28, Tian teaches that updating the data augmentation policy comprises training the data augmentation policy on a validation dataset that is separate from the training dataset starting from the trained neural network with the updated neural network parameters without data augmentation on the validation dataset, to update data augmentation policy parameters of the data augmentation policy (see e.g. section 3 “Method” on pages 3-6, and Algorithm 1 on page 6: as noted above, Tian describes an iterative process for training a data augmentation policy in which each of a plurality of iterations comprises: (i) loading shared weights w s h a r e ; (ii) fine-tuning the neural network from the shared weights w s h a r e with the training dataset D t r modified by a current data augmentation policy; and (iii) updating the current data augmentation policy based on the accuracy ACC of the fine-tuned neural network on a validation dataset D v a l . The validation dataset D v a l is separate from the training dataset D t r , and data augmentation is not applied to the validation dataset D v a l – see e.g. section 3 “Method” on pages 3-6. Tian thus teaches training the data augmentation policy on a validation dataset D v a l that is separate from the training dataset D t r starting from the trained neural network with the updated neural network parameters, i.e. the fine-tuned neural network, without data augmentation on the validation dataset D v a l , to update data augmentation policy parameters of the data augmentation policy.). Tian further teaches that the training dataset and the validation dataset can be taken from the same dataset (see e.g. section 4.1 “Datasets and Comparison Methods” and section 4.2 “Implementation Details:” Tian teaches that the training dataset and the validation dataset can be taken from the same dataset, e.g. by splitting the CIFAR-10 dataset.). Accordingly, Tian further teaches a method like that of claim 28. As per claim 33, Tian teaches that the data can comprise image data (e.g. from the CIFAR-10 dataset), that the data augmentation can comprise transforming the image data using one or more transformations, and that the task can comprise classifying a visual input (i.e. image classification) (see e.g. section 1 “Introduction,” section 3.2 “Auto-Aug Formulation,” section 3.4 “Augmentation Policy Space and Search Pipeline” and section 4.1 “Datasets and Comparison Methods”). Tian further demonstrates that the above-described process learns the data augmentation policy without using default transformations or hand-selected magnitude ranges (see e.g. section 3.4 “Augmentation Policy Space and Search Pipeline:” Tian teaches that the data augmentation policy is learned as a probability distribution over possible augmentation operations. The policy is thus learned without using default transformations or hand-selected magnitude ranges.). Accordingly, Tian further teaches a method like that of claim 33. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention. Claims 2, 4, 13 and 21 are rejected under 35 U.S.C. 103 as being unpatentable over the article by Tian described above, and also over the article entitled, “Proximal Policy Optimization Algorithms” by Schulman et al. (“Schulman”). Regarding claim 2, Tian teaches a method like that of claim 1, as is described above, which comprises rounds of training a neural network and updating a data augmentation policy. Tian particularly teaches that “Proximal Policy Optimization” can be used to update the data augmentation policy (see e.g. “Search Pipeline” on page 6). However, Tian does not explicitly disclose that updating the data augmentation policy uses a gradient computed using a score-based method, as is required by claim 2. Schulman generally describes Proximal Policy Optimization (PPO) (see e.g. the Abstract). Schulman suggests that updating a policy using PPO comprises using a gradient computed using a score-based method (see e.g. the Abstract and section 5 “Algorithm,” which describes a loss function upon which a gradient is computed to update the policy. Computing a gradient based on the loss function is considered a “score-based method” like claimed.). It would have been obvious to one of ordinary skill in the art, having the teachings of Tian and Schulman before the effective filing date of the claimed invention, to modify the method taught by Tian such that PPO like taught by Schulman is applied to update the data augmentation policy in each round, whereby the policy is updated in part by computing a gradient using a score-based method. It would have been advantageous to one of ordinary skill to utilize such PPO because it has some of the benefits of other policy optimization algorithms (e.g. trust region policy optimization) but is relatively simpler to implement, as is taught by Schulman (see e.g. the Abstract). Accordingly, Tian and Schulman teach, to one of ordinary skill in the art, a method like that of claim 2. As per claim 4, it would have been obvious, as is described above, to modify the method taught by Tian such that PPO like taught by Schulman is applied to update the data augmentation policy in each round. Schulman teaches that PPO uses a divergence and entropy-based regularization to update the policy (see e.g. section 3 “Clipped Surrogate Objective” and section 5 “Algorithm:” Schulman teaches that the loss function used to update the policy comprises a term L t C L I P ( θ ) that penalizes too large a policy update. Schulman further discloses that the loss function comprises another term S that indicates entropy-based regularization is also used – see e.g. section 5 “Algorithm.”). Accordingly, the above-described combination of Tian and Schulman is further considered to teach a method like that of claim 4. Regarding claim 13, Tian teaches a method like that of claim 1, as is described above, which comprises rounds of training a neural network and updating a data augmentation policy. Tian particularly teaches that “Proximal Policy Optimization” can be used to update the data augmentation policy (see e.g. “Search Pipeline” on page 6). However, Tian does not explicitly disclose that updating the data augmentation policy uses a divergence or entropy-based regularization, wherein the regularization forces the current data augmentation policy to stay close to an anchor policy that comprises an updated data augmentation policy from a prior round, as is required by claim 13. As noted above, Schulman generally describes Proximal Policy Optimization (PPO) (see e.g. the Abstract). Schulman suggests that updating a policy using PPO comprises using a divergence and entropy-based regularization, wherein the regularization forces the current data augmentation policy to stay close to an anchor policy that comprises an updated data augmentation policy from a prior round (see e.g. section 3 “Clipped Surrogate Objective” and section 5 “Algorithm:” Schulman teaches that the loss function used to update the policy comprises a term L t C L I P ( θ ) that penalizes too large a policy update. The term is thus indicative of a divergence-based regularization that forces the current policy to stay close to an anchor policy that comprises an updated policy from a prior round. Schulman further discloses that the loss function comprises another term S that indicates entropy-based regularization is also used – see e.g. section 5 “Algorithm.”). It would have been obvious to one of ordinary skill in the art, having the teachings of Tian and Schulman before the effective filing date of the claimed invention, to modify the method taught by Tian such that PPO like taught by Schulman is applied to update the data augmentation policy in each round, whereby the policy is updated by using a divergence or entropy-based regularization and wherein the regularization forces the current policy to stay close to an anchor policy that comprises an updated policy from a prior round. It would have been advantageous to one of ordinary skill to utilize such PPO because it has some of the benefits of other policy optimization algorithms (e.g. trust region policy optimization) but is relatively simpler to implement, as is taught by Schulman (see e.g. the Abstract). Accordingly, Tian and Schulman teach, to one of ordinary skill in the art, a method like that of claim 13. Regarding claim 21, Tian teaches a method like that of claim 1, as is described above, which comprises iteratively training a data augmentation policy using bilevel optimization. However, Tian does not explicitly disclose that the training incorporates a regularization mechanism that is entropy-based, as is required by claim 21. As noted above, Schulman generally describes Proximal Policy Optimization (PPO) (see e.g. the Abstract). Schulman suggests that updating a policy using PPO comprises using an entropy-based regularization mechanism (see e.g. section 5 “Algorithm:” Schulman discloses that a loss function used to update a policy comprises a term S that indicates an entropy bonus. The use of such an entropy-bonus is considered an entropy-based regularization mechanism.). It would have been obvious to one of ordinary skill in the art, having the teachings of Tian and Schulman before the effective filing date of the claimed invention, to modify the method taught by Tian such that PPO like taught by Schulman is applied to update the data augmentation policy in each round (i.e. to train the data augmentation policy), whereby the updating incorporates a regularization mechanism and the regularization mechanism is entropy-based. It would have been advantageous to one of ordinary skill to utilize such PPO because it has some of the benefits of other policy optimization algorithms (e.g. trust region policy optimization) but is relatively simpler to implement, as is taught by Schulman (see e.g. the Abstract). Accordingly, Tian and Schulman teach, to one of ordinary skill in the art, a method like that of claim 21. Claims 14 and 16 are rejected under 35 U.S.C. 103 as being unpatentable over the article by Tian described above, and also over the article entitled, “Train faster, generalize better: Stability of stochastic gradient descent” by Hardt et al. (“Hardt”). Regarding claim 14, Tian teaches a method like that of claim 1, as is described above, which comprises rounds of training a neural network and updating a data augmentation policy. However, Tian does not explicitly disclose that training the neural network comprises computing a stochastic gradient of a loss over the neural network parameters, and that updating the data augmentation policy comprises computing a stochastic gradient of a loss over data augmentation parameters of the data augmentation policy, as is required by claim 14. Nevertheless, updating model parameters by computing a stochastic gradient of a loss over the parameters is well-known in the art. Hardt in particular teaches using stochastic gradient methods to optimize machine learning models, whereby such methods repeatedly compute the stochastic gradient of a loss function over model parameters on a single training example or a batch of examples, and update the model parameters accordingly (see e.g. section 1 “Introduction” on page 1, and section 3 “Stability of Stochastic Gradient Method” on pages 6-7). It would have been obvious to one of ordinary skill in the art, having the teachings of Tian and Hardt before the effective filing date of the claimed invention, to modify the method taught by Tian such that a stochastic gradient method like taught by Hardt is used to train the neural network parameters and to update the parameters of the data augmentation policy. It thus follows that training the neural network would comprise computing a stochastic gradient of a loss over the neural network parameters, and that updating the data augmentation policy would similarly comprise computing a stochastic gradient of a loss over data augmentation parameters of the data augmentation policy. It would have been advantageous to one of ordinary skill to utilize such a stochastic gradient method because it is “scalable, robust and performs well across many different domains,” as is taught by Hardt (see section 1 “Introduction” on page 1). Accordingly, Tian and Hardt are considered to teach, to one of ordinary skill in the art, a method like that of claim 14. Regarding claim 16, Tian teaches a method like that of claim 1, as is described above, which comprises rounds of training a neural network and updating a data augmentation policy. Tian further teaches that training the neural network approximates an optimal inner-level solution (see e.g. section 3.2 “Auto-Aug Formation,” section 3.3. “Our Proxy Task” and Algorithm 1). However, Tian does not explicitly disclose that training the neural network approximates the optimal inner-level solution using a sequence of differentiable optimization steps, as is required by claim 16. Like noted above, Hardt teaches using stochastic gradient methods to train and optimize machine learning models (see e.g. section 1 “Introduction” on page 1, and section 3 “Stability of Stochastic Gradient Method” on pages 6-7). Hardt further discloses that such stochastic gradient methods use a sequence of differentiable optimization steps (see e.g. section 1 “Introduction” on page 1, and section 3 “Stability of Stochastic Gradient Method” on pages 6-7: Hardt discloses that stochastic gradient methods comprise repeatedly computing the gradient of a loss function on a single training example or a batch of examples. Each computation of the gradient of a loss function is considered a differentiable optimization step.). It would have been obvious to one of ordinary skill in the art, having the teachings of Tian and Hardt before the effective filing date of the claimed invention, to modify the method taught by Tian such that a stochastic gradient method like taught by Hardt is used to train the neural network to approximate the optimal inner-level solution. It thus follows that training the neural network would comprise training the neural network to approximate the optimal inner-level solution using a sequence of differentiable optimization steps. It would have been advantageous to one of ordinary skill to utilize such a stochastic gradient method because it is “scalable, robust and performs well across many different domains,” as is taught by Hardt (see section 1 “Introduction” on page 1). Accordingly, Tian and Hardt are considered to teach, to one of ordinary skill in the art, a method like that of claim 16. Claim 23 is rejected under 35 U.S.C. 103 as being unpatentable over the article by Tian described above, and also over the article entitled, “Understanding the Impact of Entropy on Policy Optimization” by Ahmed et al. (“Ahmed”). Regarding claim 23, Tian teaches a method like that of claim 1, as is described above, which comprises iteratively training a data augmentation policy using bilevel optimization. However, Tian does not explicitly disclose that the training incorporates a regularization mechanism, wherein the regularization mechanism is based on an entropy formulation whose strength decays over training iterations, as is required by claim 23. Ahmed nevertheless generally teaches using entropy regularization to improve policy optimization in reinforcement learning (see e.g. the Abstract). Ahmed teaches that including entropy regularization and decaying it during optimization can reduce the proportion of sub-optimal policy solutions found by the optimization procedure (see e.g. section 3.1.3 “Why Does Using Entropy Regularization find Better Solutions?”). It would have been obvious to one of ordinary skill in the art, having the teachings of Tian and Ahmed before the effective filing date of the claimed invention, to modify the method taught by Tian such that the training of the data augmentation policy incorporates an entropy regularization mechanism like taught by Ahmed, wherein the regularization mechanism is based on an entropy formulation whose strength decays during optimization (i.e. over training iterations). It would have been advantageous to one of ordinary skill to utilize such an entropy regularization mechanism because it can reduce the proportion of sub-optimal policy solutions found by the optimization procedure, as is taught by Ahmed (see e.g. section 3.1.3 “Why Does Using Entropy Regularization find Better Solutions?”). Accordingly, Tian and Ahmed are considered to teach, to one of ordinary skill in the art, a method like that of claim 23. Claims 9, 18, 34, 35 and 37 are rejected under 35 U.S.C. 103 as being unpatentable over the article by Tian described above, and also over the article entitled, “Differentiable RandAugment: Learning Selecting Weights and Magnitude Distributions of Image Transformations” by Xiao et al. (“Xiao”). Regarding claim 9, Tian teaches a method like that of claim 1, as is described above, which comprises rounds of training a neural network and updating a data augmentation policy. However, Tian does not explicitly disclose that the data augmentation policy comprises a categorical distribution of data transformations and a continuous distribution of magnitudes for each data transformation, wherein the data augmentation policy generates a data augmentation by: (i) selecting K data transformations from the categorical distribution, where K is at least one; (ii) selecting values for magnitudes for each of the selected K data transformations; and (iii) composing the K data transformations to obtain the data augmentation, as is required by claim 9. Xiao nevertheless suggests learning a data augmentation policy that comprises a categorical distribution of data transformations (i.e. operations) and a continuous distribution (i.e. a normal distribution) of magnitudes for each data transformation, wherein the data augmentation policy generates a data augmentation by: (i) selecting K data transformations (i.e. a sequence of D operations) from the categorical distribution, where K is at least one; (ii) selecting values for magnitudes for each of the selected K data transformations; and (iii) composing the K data transformations to obtained the data augmentation (see e.g. section III.A “Reformulate Data Augmentation”). It would have been obvious to one of ordinary skill in the art, having the teachings of Tian and Xiao before the effective filing date of the claimed invention, to modify the method taught by Tian such that the data augmentation policy comprises a categorical distribution of data transformations and a continuous distribution of magnitudes for each data transformation like taught by Xiao, and wherein the data augmentation policy generates a data augmentation by: (i) selecting K data transformations from the categorical distribution, where K is at least one; (ii) selecting values for magnitudes for each of the selected K data transformations; and (iii) composing the K data transformations to obtain the data augmentation. It would have been advantageous to one of ordinary skill to utilize such a data augmentation policy because it can achieve relatively better performance on classification and object detection, as is suggested by Xiao (see e.g. section 1 “Introduction”). Accordingly, Tian and Xiao are considered to teach, to one of ordinary skill in the art, a method like that of claim 9. Regarding claim 18, Tian teaches a method like that of claim 1, as is described above, which comprises rounds of training a neural network and updating a data augmentation policy. However, Tian does not explicitly disclose that the data augmentation policy comprises a composite of multiple transformations, wherein the method learns a categorical probability distribution over the multiple transformations and a continuous distribution of a magnitude of each of the multiple transformations, as is required by claim 18. Xiao nevertheless suggests learning a data augmentation policy that comprises a composite of multiple transformations (i.e. operations), and comprises a categorical probability distribution over the multiple transformations and a continuous distribution (i.e. a normal distribution) of a magnitude of each of the multiple transformations (see e.g. section III.A “Reformulate Data Augmentation”). It would have been obvious to one of ordinary skill in the art, having the teachings of Tian and Xiao before the effective filing date of the claimed invention, to modify the method taught by Tian such that the data augmentation policy comprises a composite of multiple transformations like taught by Xiao, and wherein the method learns a categorical probability distribution over the multiple transformations and a continuous distribution of a magnitude of each of the multiple transformations. It would have been advantageous to one of ordinary skill to utilize such a data augmentation policy because it can achieve relatively better performance on classification and object detection, as is suggested by Xiao (see e.g. section 1 “Introduction”). Accordingly, Tian and Xiao are considered to teach, to one of ordinary skill in the art, a method like that of claim 18. Regarding claim 34, Tian generally teaches using “Augmentation-Wise Weight Sharing” (AWS) to form an evaluation process for augmentation polices, and which is used while automatically searching augmentation policies (see e.g. the Abstract). Like claimed, Tian particularly teaches: initializing a data augmentation policy (see e.g. section 3.3 “Our Proxy Task” on pages 4-5: Tian teaches partitioning the augmented training of neural network parameters into two parts: a first part in which a shared augmentation policy is applied to train a neural network and thereby obtain shared weights w s h a r e for the neural network; and a second part in which the neural network model is fine-tuned from the shared weights by a given data augmentation policy so that the policy can evaluated. Tian further discloses that the shared augmentation policy comprises a uniform sampling of augmentation transforms – see e.g. section 3.3 “Our Proxy Task” on page 5. The shared data augmentation policy is considered an initialized data augmentation policy.); initializing neural network parameters of a neural network with neural network parameters trained during a pretraining, wherein said pretraining uses data augmented by the initialized data augmentation policy (see e.g. section 3.3 “Our Proxy Task” on pages 4-5: as noted above, Tian teaches partitioning the augmented training of neural network parameters into two parts, wherein a first part applies a shared augmentation policy to train a neural network and thereby obtain shared weights w s h a r e for the neural network, and a second part that fine-tunes neural network model from the shared weights. The first part particularly trains the neural network on a task using a training dataset D t r that is augmented by the shared augmentation policy – see e.g. section 3.3 “Our Proxy Task” on pages 4-5 and the line “Obtain w s h a r e in Equ. 3;” in Algorithm 1 on page 6. This first part is considered to pretrain the neural network. The second part particularly comprises an iterative process in which, for each iteration: (i) the shared weights w s h a r e are loaded; (ii) the neural network is fined-tuned from the shared weights w s h a r e with the training dataset D t r modified by a current data augmentation policy; and (iii) the current data augmentation policy is updated based on the accuracy ACC of the fine-tuned neural network on a validation dataset D v a l – see e.g. section 3 “Method” on pages 3-6, and the following lines in Algorithm 1 on page 6: while T ≤ T m a x do Load w s h a r e ; Fine-tune w s h a r e to get w - θ * ; Use ACC ( w - θ * , D v a l ) to update θ ; end while Loading the shared weights w s h a r e during each iteration of the second part is considered initializing neural network parameters of a neural network with neural network parameters, i.e. with w s h a r e , that are trained during a pretraining, i.e. during the first part, wherein the pretraining uses data augmented by the initialized data augmentation policy, i.e. by the shared augmentation policy.); training a neural network on a task to update neural network parameters on a training dataset augmented by a current augmentation policy (see e.g. section 3 “Method” on pages 3-6, and Algorithm 1 on page 6: as noted above, Tian teaches partitioning the augmented training of neural network parameters into two parts, wherein the second part comprises an iterative process in which, for each iteration: (i) shared weights w s h a r e are loaded; (ii) the neural network is fined-tuned from the shared weights w s h a r e with a training dataset D t r modified by a current data augmentation policy; and (iii) the current data augmentation policy is updated based on the accuracy ACC of the fine-tuned neural network on a validation dataset D v a l . The fine-tuning in each iteration is considered training a neural network on a task to update neural network parameters on a training dataset augmented by a current augmentation policy.); and updating the data augmentation policy based on the updated neural network parameters using an evaluation dataset that is separate from the training dataset, without augmenting the evaluation dataset, to define the current data augmentation policy for the next step or the data augmentation policy on the last round (see e.g. section 3 “Method” on pages 3-6, and Algorithm 1 on page 6: as noted above, Tian describes an iterative process in which each iteration comprises updating the current data augmentation policy based on the accuracy ACC of the fine-tuned neural network on a validation dataset D v a l . The validation dataset D v a l is separate from the training dataset D t r used to fine-tune the neural network, and data augmentation is not applied to the validation dataset D v a l – see e.g. section 3 “Method” on pages 3-6. The updated data augmentation policy is then used in the next step, i.e. within a next iteration, or is returned if T > T m a x – see e.g. section 3 “Method” on pages 3-6, and Algorithm 1 on page 6. Accordingly, Tian teaches updating the data augmentation policy based on the updated neural network parameters, i.e. based on the fine-tuned neural network, using an evaluation dataset D v a l that is separate from the training dataset D t r , without augmenting the evaluation dataset, to define the current augmentation policy for the next step or the data augmentation policy on the last round, i.e. the last iteration.). Tian thus teaches a method similar to that of claim 34. However, Tian does not explicitly disclose that the data augmentation policy comprises a categorical distribution of data transformations, and a continuous distribution of magnitudes for each data transformation, as is required by claim 34. Xiao nevertheless suggests learning a data augmentation policy that comprises a categorical distribution of data transformations (i.e. operations) and a continuous distribution (i.e. a normal distribution) of magnitudes for each data transformation (see e.g. section III.A “Reformulate Data Augmentation”). It would have been obvious to one of ordinary skill in the art, having the teachings of Tian and Xiao before the effective filing date of the claimed invention, to modify the method taught by Tian such that the data augmentation policy comprises a categorical distribution of data transformations and a continuous distribution of magnitudes for each data transformation, as is taught by Xiao. It would have been advantageous to one of ordinary skill to utilize such a data augmentation policy because it can achieve relatively better performance on classification and object detection, as is suggested by Xiao (see e.g. section 1 “Introduction”). Accordingly, Tian and Xiao are considered to teach, to one of ordinary skill in the art, a computer-implemented method like that of claim 34, which is for learning a data augmentation policy represented by data augmentation parameters. As per claim 35, Tian further teaches that initializing the neural network parameters, training a neural network, and updating the data augmentation policy are each performed over each of one or more rounds (see e.g. section 3 “Method” on pages 3-6, and Algorithm 1 on page 6: as noted above, Tian teaches partitioning the augmented training of neural network parameters into two parts, wherein the second part comprises an iterative process in which, for each iteration: (i) shared weights w s h a r e are loaded; (ii) the neural network is fined-tuned from the shared weights w s h a r e with a training dataset D t r modified by a current data augmentation policy; and (iii) the current data augmentation policy is updated based on the accuracy ACC of the fine-tuned neural network on a validation dataset D v a l . Each iteration is considered a round, like claimed. Accordingly, (i) initializing the neural network parameters, i.e. loading the shared weights w s h a r e , (ii) training the neural network, i.e. fine-tuning the neural network from the shared weights w s h a r e , and (iii) updating the data augmentation policy, are performed over each of one or more rounds.). Tian further teaches that training a neural network and updating the data augmentation policy are performed over each of a plurality of steps within each round (see e.g. section 3.3 “Our Proxy Task” on page 5, and section 4.2 “Implementation Details” on page 6: Tian teaches that the fine-tuning in each iteration occurs over a plurality of epochs – see e.g. section 3.3 “Our Proxy Task” on page 5, and section 4.2 “Implementation Details” on page 6. Accordingly, the fine-tuning is understandably performed over a plurality of steps within each round/iteration. Tian also suggests that the data augmentation policy is updated over a plurality of steps within each iteration, as would be necessary to determine a validation loss and then determine updated policy parameters based on the validation loss – see e.g. section 3.3 “Our Proxy Task” on page 5-6.). Accordingly, the above-described combination of Tian and Xiao is further considered to teach a method like that of claim 35. As per claim 37, Tian further teaches training, over an additional plurality of steps within each round, the neural network on the task to update the neural network parameters on the training dataset without updating the data augmentation policy, the training dataset being augmented by the current data augmentation policy for the current step (see e.g. section 3 “Method” on pages 3-6, and Algorithm 1 on page 6: as noted above, Tian describes an iterative process in which each iteration comprises: (i) loading the shared weights w s h a r e ; (ii) fine-tuning the neural network from the shared weights w s h a r e with the training dataset D t r modified by a current data augmentation policy; and (iii) updating the current data augmentation policy based on the accuracy ACC of the fine-tuned neural network on a validation dataset D v a l . Tian discloses that the fine-tuning can occur over a number of steps, e.g. over a number of epochs – see e.g. section 3.3 “Our Proxy Task” on page 5, and section 4.2 “Implementation Details” on page 6. Tian thus teaches fine-tuning, over an additional plurality of steps, e.g. epochs, within each round, the neural network on the task to update the neural network parameters on the training dataset without updating the data augmentation policy, wherein the training dataset is augmented by the current data augmentation policy for the current step.). Accordingly, the above-described combination of Tian and Xiao further teaches a method like that of claim 37. Claims 11 and 36 are rejected under 35 U.S.C. 103 as being unpatentable over the combination of Tian and Xiao described above, and also over U.S. Patent Application Publication No. 2021/0097348 to Shlens et al. (“Shlens”). Regarding claim 11, Tian and Xiao teach a method like that of claim 9, as is described above, which comprises rounds of training a neural network and updating a data augmentation policy. In particular, as is described above, it would have been obvious to modify the method taught by Tian such that the data augmentation policy comprises a categorical distribution of data transformations and a continuous distribution of magnitudes for each data transformation like taught by Xiao. Xiao particularly teaches that the data transformations comprise elementary transformations s i in a set S (i.e. operators in a set of candidate operators) from which elementary transformations t are sampled, and wherein the magnitudes m of each selected elementary transformation t i in S in the data augmentation policy are sampled from a continuous distribution (i.e. a normal distribution) (see e.g. section III.A “Reformulate Data Augmentation”). Accordingly, Tian and Xiao further teach a method similar to that of claim 11, but do not explicitly disclose that the magnitudes m of each selected elementary transformation t i in S in the data augmentation policy are sampled from a smoothed uniform distribution between 0 , u i whose upper bound u i is learned during each round, as is required by claim 11. Similar to Tian and Xiao, Shlens teaches identifying a suitable data augmentation policy by iteratively training a neural network and updating a candidate data augmentation policy (see e.g. paragraphs 0034-0043). Shlens particularly teaches that the data augmentation policy can provide a categorical distribution of data transformations (i.e. transformation operations) and a continuous distribution (i.e. a range) of magnitudes for each data transformation, wherein the data transformations comprise elementary transformations s i in a set S from which elementary transformations t are sampled (see e.g. paragraphs 0053-0055, 0059, and 0063-0064). Moreover, Shlens further suggests that the magnitudes m of each selected elementary transformation in t i in S in the data augmentation policy can be sampled from a smoothed uniform distribution (i.e. are randomly sampled) between 0 , u i whose upper bound u i is learned during each iteration (see e.g. paragraphs 0043, 0055, 0059 and 0062). It would have been obvious to one of ordinary skill in the art, having the teachings of Tian, Xiao and Shlens before the effective filing date of the claimed invention, to modify the method taught by Tian and Xiao such that the magnitudes m of each selected elementary transformation t i in S in the data augmentation policy are alternatively sampled from a smoothed uniform distribution between 0 , u i whose upper bound u i is learned during each round, as is taught by Shlens. It would have been advantageous to one of ordinary skill to utilize such a magnitude selection mechanism, because it can provide a more effective data augmentation policy, as is suggested by Shlens (see e.g. paragraphs 0010, 0055 and 0061-0062). Accordingly, Tian, Xiao and Shlens are considered to teach, to one of ordinary skill in the art, a method like that of claim 11. Regarding claim 36, Tian and Xiao teach a method like that of claim 34, as is described above, which is for learning a data augmentation policy represented by data augmentation parameters, and wherein the data augmentation policy comprises a categorical distribution of data transformations and a continuous distribution of magnitudes for each transformation. Tian and Xiao, however, do not explicitly disclose that the continuous distribution of magnitudes for each data transformation comprises a smoothed uniform distribution parameterized by an upper bound, as is required by claim 36. Like noted above, Shlens teaches identifying a suitable data augmentation policy by iteratively training a neural network and updating a candidate data augmentation policy (see e.g. paragraphs 0034-0043). Shlens particularly teaches that the data augmentation policy can provide a categorical distribution of data transformations (i.e. transformation operations) and a continuous distribution (i.e. a range) of magnitudes for each data transformation. Moreover, Shlens suggests that the continuous distribution of magnitudes for each data transformation can comprise a smoothed uniform distribution parameterized by an upper bound (see e.g. paragraphs 0055, 0059 and 0062). It would have been obvious to one of ordinary skill in the art, having the teachings of Tian, Xiao and Shlens before the effective filing date of the claimed invention, to modify the method taught by Tian and Xiao such that continuous distribution of magnitudes for each data transformation comprises a smoothed uniform distribution parameterized by an upper bound, as is taught by Shlens. It would have been advantageous to one of ordinary skill to utilize such a smoothed uniform distribution, because it can provide a more effective data augmentation policy, as is suggested by Shlens (see e.g. paragraphs 0010, 0055 and 0061-0062). Accordingly, Tian, Xiao and Shlens are considered to teach, to one of ordinary skill in the art, a method like that of claim 36. Claim 26 is rejected under 35 U.S.C. 103 as being unpatentable over the article to Tian described above, and also over the article entitled, “Feature Extraction using Convolution Neural Networks (CNN) and Deep Learning” by Jogin et al. (“Jogin”). Regarding claim 26, Tian teaches a method like that of claim 1, as is described above, which comprises rounds of training a neural network and updating a data augmentation policy. Tian, however, does not explicitly disclose that the neural network comprises a convolutional neural network (CNN), as is required by claim 26. Convolutional neural networks are nevertheless well-known in the art. Jogin for example describes convolutional neural networks (see e.g. section I. “Introduction,” section II.D “Convolutional Neural networks” and section III.E “Convolutional Neural Network”). Jogin teaches that convolutional neural networks are an often-used architecture that perform well in computer vision tasks such as image classification (see e.g. the Abstract, section III.F “Result Analysis” and section IV. “Conclusion”). It would have been obvious to one of ordinary skill in the art, having the teachings of Tian and Jogin before the effective filing date of the claimed invention, to modify the method taught by Tian such that the neural network comprises a convolutional neural network like taught by Jogin. It would have been advantageous to one of ordinary skill to utilize such a convolutional neural network because it is an often-used architecture that performs well in computer vision tasks, as is taught by Jogin (see e.g. the Abstract, section III.F “Result Analysis” and section IV. “Conclusion”). Accordingly, Tian and Jogin are considered to teach, to one of ordinary skill in the art, a method like that of claim 26. Claims 29 and 30 are rejected under 35 U.S.C. 103 as being unpatentable over the article to Tian described above, and also over the article entitled, “Learning Data Augmentation Strategies for Object Detection” by Zoph et al. (“Zoph”). Regarding claim 29, Tian teaches a method like that of claim 27, as is described above, which comprises generating augmented data from a dataset using an augmentation policy trained on a neural network, and training the neural network or a different neural network on the task using the augmented data. Tian, however, does not explicitly teach evaluating the trained neural network on the task using a testing dataset that is separate from the dataset, without data augmentation on the testing dataset, as is required by claim 29. Similar to Tian, Zoph generally teaches training a data augmentation policy, including by training a neural network on a task to update the neural network parameters on a training dataset (see e.g. the Abstract and section 3 “Methods”). Zoph further teaches generating augmented data from a dataset using the trained data augmentation policy, and training the neural network or a different neural network on the task using the augmented data (see e.g. the Abstract, section 1 “Introduction,” section 4.3 “Exploiting learned augmentation policies achieves state-of-the-art object detection,” and section 4.4 “Learned augmentation policies transfer to other detection datasets:” Zoph teaches that the learned augmentation policy can be applied to other neural network architectures, which would entail generating augmented data from a dataset using the trained data augmentation policy, and training a different neural network on the task using the augmented data.). Moreover, Zoph suggests evaluating the trained neural network on the task using a testing dataset that is separate from the dataset used to train the different neural network, without data augmentation on the testing dataset (see e.g. section 1 “Introduction,” which teaches that the transformations of a data augmentation policy are only used during training and not during testing. Section 4.3 “Exploiting learned augmentation policies achieves state-of-the-art object detection” and section 4.4 “Learned augmentation policies transfer to other detection datasets” suggest that the testing dataset is separate from the training dataset, and that the testing dataset is used to evaluate the trained neural network on the task.). It would have been obvious to one of ordinary skill in the art, having the teachings of Tian and Zoph before the effective filing date of the claimed invention, to modify the method taught by Tian so as to evaluate the trained neural network on the task using a testing dataset that is separate from the dataset, without data augmentation on the testing dataset, as is taught by Zoph. It would have been advantageous to one of ordinary skill to utilize such a testing dataset, because it can provide an indication of the performance of the trained neural network relative to other neural networks, as is suggested by Zoph (see e.g. Section 4.3 “Exploiting learned augmentation policies achieves state-of-the-art object detection” and section 4.4 “Learned augmentation policies transfer to other detection datasets”). Accordingly, Tian and Zoph are considered to teach, to one of ordinary skill in the art, a method like that of claim 29. Regarding claim 30, Tian teaches a method like that of claim 27, as is described above, which comprises generating augmented data from a dataset using a trained data augmentation policy – the data augmentation policy being trained in part by training a neural network on a task using a training dataset – and training the neural network or a different neural network on a task using the augmented data. Tian further teaches that updating the data augmentation policy comprises training the data augmentation policy on a validation dataset that is separate from the training dataset starting from the trained neural network with the updated neural network parameters without data augmentation on the validation dataset, to update data augmentation policy parameters of the data augmentation policy (see e.g. section 3 “Method” on pages 3-6, and Algorithm 1 on page 6: as noted above, Tian teaches partitioning the augmented training of neural network parameters into two parts: a first part in which a shared augmentation policy is applied to train a neural network and thereby obtain shared weights w s h a r e for the neural network; and a second part in which the neural network model is fine-tuned from the shared weights by a given data augmentation policy so that the policy can evaluated. The second part particularly comprises an iterative process in which, for each iteration: (i) the shared weights w s h a r e are loaded; (ii) the neural network is fined-tuned from the shared weights w s h a r e with the training dataset D t r modified by a current data augmentation policy; and (iii) the current data augmentation policy is updated based on the accuracy ACC of the fine-tuned neural network on a validation dataset D v a l – see e.g. section 3 “Method” on pages 3-6, and Algorithm 1 on page 6. The validation dataset D v a l is separate from the training dataset D t r used to fine-tune the neural network, and data augmentation is not applied to the validation dataset D v a l – see e.g. section 3 “Method” on pages 3-6. Tian thus teaches that updating the data augmentation policy comprises training the data augmentation policy on a validation dataset D v a l that is separate from the training dataset D t r starting from the trained neural network with the updated neural network parameters, i.e. with shared weights w s h a r e , without data augmentation on the validation dataset, to update data augmentation policy parameters of the data augmentation policy.). Tian, however, does not explicitly disclose that the dataset (i.e. the dataset used to train the neural network or a different neural network) is from a different domain as the training dataset and the evaluation dataset (i.e. the datasets used to train the data augmentation policy), as is required by claim 30. Like noted above, Zoph generally teaches training a data augmentation policy, including by training a neural network on a task to update the neural network parameters on a training dataset (see e.g. the Abstract and section 3 “Methods”). Similar to Tian, Zoph further teaches generating augmented data from a dataset using the trained data augmentation policy, and training the neural network or a different neural network on the task using the augmented data (see e.g. the Abstract, section 1 “Introduction,” section 4.3 “Exploiting learned augmentation policies achieves state-of-the-art object detection,” and section 4.4 “Learned augmentation policies transfer to other detection datasets:” Zoph teaches that the learned augmentation policy can be applied to other neural network architectures, which would entail generating augmented data from a dataset using the trained data augmentation policy, and training a different neural network on the task using the augmented data.). Zoph particularly teaches that the dataset (i.e. the dataset used to train the different neural network) can be from a different domain as the training dataset and the evaluation dataset (i.e. the datasets used to train the data augmentation policy) (see e.g. the Abstract, section 4.3 “Exploiting learned augmentation policies achieves state-of-the-art object detection,” and section 4.4 “Learned augmentation policies transfer to other detection datasets:” Zoph teaches that the learned augmentation policy can be applied to other neural network architectures and datasets. In such instances, the dataset to which the learned data augmentation is applied would be from a different domain as the datasets, i.e. the training dataset and the validation/evaluation dataset, used to train the augmentation policy). It would have been obvious to one of ordinary skill in the art, having the teachings of Tian and Zoph before the effective filing date of the claimed invention, to modify the method taught by Tian such that the dataset is from a different domain as the training dataset and the evaluation dataset, as is taught by Zoph. It would have been advantageous to one of ordinary skill to utilize such a combination, because it can improve the performance of the neural network on the dataset, as is evident from Zoph (see e.g. section 4.4 “Learned augmentation policies transfer to other detection datasets”). Accordingly, Tian and Zoph are also considered to teach, to one of ordinary skill in the art, a method like that of claim 30. Claim 32 is rejected under 35 U.S.C. 103 as being unpatentable over the article to Tian described above, and also over the article entitled, “Deep convolutional neural network based medical image classification for disease diagnosis” by Yadav et al. (“Yadav”). Regarding claim 32, Tian teaches a method like that of claim 1, as is described above, which comprises rounds of training a neural network and updating a data augmentation policy. Tian teaches that the data augmentation policy can be applied to data that comprises image data (e.g. from the CIFAR-10 dataset), that the data augmentation can comprise transforming the image data using one or more image transformations, and that the task can comprise classifying a visual input (i.e. image classification) (see e.g. section 1 “Introduction,” section 3.2 “Auto-Aug Formulation,” section 3.4 “Augmentation Policy Space and Search Pipeline” and section 4.1 “Datasets and Comparison Methods”). Tian, however, does not explicitly disclose that the visual input comprises natural images, medical images, sketches, spectral images and/or infrared images, as is further required by claim 32. Classifying such types of visual input is nevertheless well-known in the art. Yadav for example generally teaches training and applying convolutional neural networks to medical images so as to perform classification thereof (see e.g. the “Introduction” on pages 1-2). Yadav further teaches that data augmentation can be applied to small medical image datasets used to train such convolutional neural networks, and that the data augmentation increases the performance of the image classification (see e.g. the “Introduction” on pages 1-2). It would have been obvious to one of ordinary skill in the art, having the teachings of Tian and Yadav before the effective filing date of the claimed invention, to apply the method taught by Tian to medical images like taught by Yadav, i.e. where the visual input comprises medical images. It would have been advantageous to one of ordinary skill to utilize medical images, because medical image classification plays an essential role in clinical treatment and teaching tasks, as is taught by Yadav (see e.g. the Abstract). Accordingly, Tian and Yadav are also considered to teach, to one of ordinary skill in the art, a method like that of claim 32. Conclusion The prior art made of record on form PTO-892 and not relied upon is considered pertinent to applicant’s disclosure. The applicant is required under 37 C.F.R. §1.111(C) to consider these references fully when responding to this action. In particular, the article by Yang et al. cited therein (“A survey of automated data augmentation algorithms for deep learning-based image classification tasks”) provides a survey of major works in the field of automated data augmentation. The U.S. Patent Application Publication to Vasudevan et al. cited therein describes methods for learning a data augmentation policy for training a machine learning model. The U.S. Patent Application Publication to Mounsaveng et al. cited therein describes a method and a system for joint data augmentation and classification learning, where an augmentation network learns to perform transformations and a classification network is trained. Any inquiry concerning this communication or earlier communications from the examiner should be directed to BLAINE T BASOM whose telephone number is (571)272-4044. The examiner can normally be reached Monday-Friday, 9:00 am - 5:30 pm, EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Matt Ell can be reached at (571)270-3264. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /BTB/ 8/6/2026 /MATTHEW ELL/Supervisory Patent Examiner, Art Unit 2141
Read full office action

Prosecution Timeline

Apr 12, 2024
Application Filed
Aug 12, 2026
Non-Final Rejection mailed — §102, §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12704937
METHODS, APPARATUSES, AND COMPUTER-READABLE MEDIA FOR ENHANCING DIGITAL PATHOLOGY PLATFORM
4y 1m to grant Granted Aug 11, 2026
Patent 12688528
Design Resources
5y 1m to grant Granted Jul 21, 2026
Patent 12669920
SYSTEM AND GRAPHICAL USER INTERFACE FOR GUIDED NEW SPACE CREATION FOR A CONTENT COLLABORATION SYSTEM
3y 9m to grant Granted Jun 30, 2026
Patent 12663907
DEVICES, METHODS, AND GRAPHICAL USER INTERFACES FOR GAZE-BASED NAVIGATION
5y 3m to grant Granted Jun 23, 2026
Patent 12632794
METHOD AND SYSTEM FOR CROSS-CHAIN CONSENSUS ORIENTED TO FEDERATED LEARNING
4y 5m to grant Granted May 19, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
43%
Grant Probability
64%
With Interview (+20.8%)
4y 6m (~2y 0m remaining)
Median Time to Grant
Low
PTA Risk
Based on 338 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month