Prosecution Insights
Last updated: October 02, 2026
Application No. 17/720,212

SYNTHETIC DATA GENERATION USING DEEP REINFORCEMENT LEARNING

Final Rejection §103
Filed
Apr 13, 2022
Examiner
BOSTWICK, SIDNEY VINCENT
Art Unit
2124
Tech Center
2100 — Computer Architecture & Software
Assignee
Microsoft Technology Licensing, LLC
OA Round
4 (Final)
51%
Grant Probability
Moderate
5-6
OA Rounds
0m
Est. Remaining
86%
With Interview

Examiner Intelligence

Grants 51% of resolved cases
51%
Career Allowance Rate
78 granted / 152 resolved
-3.7% vs TC avg
Strong +35% interview lift
Without
With
+35.1%
Interview Lift
resolved cases with interview
Typical timeline
4y 5m
Avg Prosecution
41 currently pending
Career history
216
Total Applications
across all art units

Statute-Specific Performance

§101
24.6%
-15.4% vs TC avg
§103
46.5%
+6.5% vs TC avg
§102
4.6%
-35.4% vs TC avg
§112
24.0%
-16.0% vs TC avg
Black line = Tech Center average estimate • Based on career data from 152 resolved cases

Office Action

§103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Remarks This Office Action is responsive to Applicants' Amendment filed on July 28, 2026, in which claims 1, 6, 10, 15, and 19 are currently amended. Claims 1-20 are currently pending. Response to Arguments Applicant’s arguments with respect to rejection of claims 1-20 under 35 U.S.C. 103 based on amendment have been considered and are persuasive. The argument is moot in view of a new ground of rejection set forth below. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claims 1-5, 8, 10-14, 16, and 19 are rejected under U.S.C. §103 as being unpatentable over the combination of Fedus (“MaskGAN: Better Text Generation via Filling in the ______”, 2018) and Tucker (“Inverse reinforcement learning for video games”, 2018). Regarding claim 1, Fedus teaches A system for deep reinforcement learning, the system comprising: a processor;([p. 1] "We overcome this by using Reinforcement Learning (RL) to train the generator") a first neural network implemented on the processor; ([p. 3] "Our generator consists of an encoding module and decoding module." [p. 4] "MaskGAN source code available at: https://github.com/tensorflow/models/tree/ master/research/maskgan" Generator interpreted as first neural network) a second neural network implemented on the processor, the second neural network being different from the first neural network;([p. 3] "The discriminator has an identical architecture to the generator1 except that the output is a scalar probability at each time point, rather than a distribution over the vocabulary size" Discriminator interpreted as second neural network) and a memory storing instructions that, when executed by the processor, cause the processor to:([p. 4] "MaskGAN source code available at: https://github.com/tensorflow/models/tree/ master/research/maskgan") generate, by the first neural network, a synthetic data based on an original data,([p. 2] "In this task, portions of a body of text are deleted or redacted. The goal of the model is to then infill the missing portions of text so that it is indistinguishable from the original data" [p. 3] "the decoder fills in the missing tokens auto-regressively, however, it is now conditioned on both the masked text m(x)" [p. 2] "Despite the entire sequence now being clearly synthetic as a result of the errant token, a discriminative model that produces a high loss signal to the outlier token, but not to the others, will likely yield a more informative error signal to the generator.") wherein the original data has a label with a first binary identifier and the synthetic data has a label with a second binary identifier, ([p. 3] "For a discrete sequence x =(x1,··· ,xT), a binary mask is generated (deterministically or stochastically) of the same length m=(m1,··· ,mT)where each mt ∈ {0,1}, selects which tokens will remain. The token at time t, xt is then replaced with a special mask token <m> if the mask is 0 and remains unchanged if the mask is 1." mt=1 designates a position whose original token remains, mt=0 designates a position whose original token is removed and subsequently populated by synthetic data) the first binary identifier and the second binary identifier applied to mask the identity of original data and the synthetic data from the second neural network, respectively;([p. 4] "We give the discriminator the true context, otherwise, this algorithm has a critical failure mode. For instance, without this context, if the discriminator is given the filled-in sequence the director guided the series, it will fail to reliably identify the director bigram as fake text, despite this bigram potentially never appearing in the training corpus (aside from an errant typo) [...] our discriminator Dφ computes the probability of each token ˜xt being real given the true context of the masked sequence m(x)" Fedus explicitly gives the discriminator/second neural network the masked text masking the identity of original data and synthetic data. Fedus explains that without masking the identity the discriminator will fail) provide the original data and the generated synthetic data to the second neural network, ([p. 3] "The discriminator is given the filled-in sequence from the generator" [p. 6] "we then pretrain the discriminator on the samples produced from the current generator and real training text" Fedus explicitly trains the second neural network/discriminator on both real training text/original data and generated synthetic data) wherein the first neural network and the second neural network are structured to continuously learn using deep reinforcement learning including a state, an action, a reward, and a next state, ([p. 4] "the logarithm of the discriminator estimates are regarded as the reward […] The critic estimates the value function, which is the discounted total return of the filled-in sequence Rt = T s=tγsrs, where γ is the discount factor at each position in the sequence [...] the action at is the token chosen by the generator at ≡ ˆxt" [p. 5] "the state st are the current tokens produced up to that point st ≡ ˆx1,··· , ˆxt−1. This approach is an actor-critic architecture where G determines the policy π(st) and the baseline bt is the critic") wherein the state is a value of an input feature in the original data, ([p. 2] "In this task, portions of a body of text are deleted or redacted. The goal of the model is to then infill the missing portions of text so that it is indistinguishable from the original data" [p. 4] "the state st are the current tokens produced up to that point st ≡ ˆx1,··· , ˆxt−1." In Fedus the state is explicitly values of missing tokens in the original data) the action [is continuous and] includes the synthetic data, ([p. 4] "the action at is the token chosen by the generator at ≡ ˆxt"") the reward is a measure of how unsuccessful the second neural network is in discriminating the original data from the synthetic data, ([p. 4] "our discriminator Dφ computes the probability of each token ˜xt being real given the true context of the masked sequence m(x) [...] the logarithm of the discriminator estimates are regarded as the reward […] The critic estimates the value function, which is the discounted total return of the filled-in sequence Rt = T s=tγsrs, where γ is the discount factor at each position in the sequence") and the next state is a next group of examples to generate a next iteration of the synthetic data based on the original data,([p. 4] "the action at is the token chosen by the generator at ≡ ˆxt" [p. 5] "the state st are the current tokens produced up to that point st ≡ ˆx1,··· , ˆxt−1. This approach is an actor-critic architecture where G determines the policy π(st) and the baseline bt is the critic" [p. 3] "As in standard language-modeling, the decoder fills in the missing tokens auto-regressively" In Fedus the action results in a next state which explicitly is a next group of examples (missing token values) which is then in turn used to generate the subsequent action (synthetic data based on the original data missing token) and this process is explicitly performed iteratively/auto-regressively) generate, by the second neural network a prediction identifying the original data and the generated synthetic data, and([p. 4] "our discriminator Dφ computes the probability of each token ˜xt being real given the true context of the masked sequence m(x)") based at least in part on the prediction incorrectly identifying the generated synthetic data, export the generated synthetic data.([p. 4] "our discriminator Dφ computes the probability of each token ˜xt being real given the true context of the masked sequence m(x) [...] the logarithm of the discriminator estimates are regarded as the reward […] The critic estimates the value function, which is the discounted total return of the filled-in sequence Rt = T s=tγsrs, where γ is the discount factor at each position in the sequence" [p. 6] "We present both conditional and unconditional samples generated on the PTB and IMDB data sets at word-level. MaskGAN refers to our GAN-trained variant and MaskMLE refers to our maximum likelihood trained variant. Additional samples are supplied in Appendix B." [p. 2] "the generation of synthetic training data"). However, Fedus does not explicitly teach the action is continuous. Tucker, in the same field of endeavor, teaches the action is continuous([p. 3 §3] "We build on adversarial IRL (AIRL), an algorithm achieving state-of-the-art performance on simulated robotics tasks (Fu et al., 2017). The reference implementation of AIRL assumes a continuous action space and uses a fully-connected policy and reward network [...] Adversarial IRL formulates the inverse reinforcement learning problem as a GAN (Goodfellow et al., 2014). We learn a reward function fθ(s,a) for taking action a in state s and a stochastic policy π(a | s). The policy is the generator, and is trained using forward RL on the reward function fθ(s,a). The discriminator is restricted to have the special form"). Fedus as well as Tucker are directed towards hybrid reinforcement learning generative adversarial neural network models. Therefore, Fedus as well as Tucker are analogous art in the same field of endeavor. It would have been obvious before the effective filing date of the claimed invention to combine the teachings of Fedus with the teachings of Tucker by using a continuous action space. One of ordinary skill in the art would realize that the action space must be either discrete or continuous and Tucker explicitly acknowledges continuous action space in reinforcement learning generative adversarial networks and provides as additional motivation for combination ([p. 1 §1] “Recent deep IRL algorithms have achieved good performance on a variety of continuous control tasks”). This motivation for combination also applies to the remaining claims which depend on this combination. Regarding claim 2, the combination of Fedus and Tucker teaches The system of claim 1, wherein the instructions further cause the processor to: based at least in part on the prediction correctly identifying the generated synthetic data, control the first neural network to execute a machine learning model to calculate a loss for the generated synthetic data and update parameters; and(Fedus [p. 5] "Finally, as in conventional GAN training, our discriminator will be updated according to the gradient" [p. 5] "the gradient to the generator associated with producing ˆxt will depend on all the discounted future rewards (s ≥ t) assigned by the discriminator" [p. 4] "the logarithm of the discriminator estimates are regarded as the reward" [p. 14] "We perform 3 gradient descent steps on the discriminator for every step on the generator and critic" See Eqn. 7 which shows gradient calculated by log loss formula explicitly using the generator G (first neural network)) using the updated parameters, control the first neural network to generate a second synthetic data.(Fedus [p. 5] "the gradient to the generator associated with producing ˆxt will depend on all the discounted future rewards (s ≥ t) assigned by the discriminator" [p. 4] "the logarithm of the discriminator estimates are regarded as the reward" [p. 14] "We perform 3 gradient descent steps on the discriminator for every step on the generator and critic" Fedus is explicit that the generator and discriminator are trained iteratively and directly on one another’s updates). Regarding claim 3, the combination of Fedus and Tucker teaches The system of claim 2, wherein the instructions further cause the first neural network to update the parameters by subtracting a gradient of the calculated loss and minimizing values in an opposite direction of the gradient(Fedus [p. 4] "We optimize the parameters of the generator, xt θ, by performing gradient ascent on EG(θ)[R] […] ∇θEG[Rt] = (Rt −bt)∇θlogGθ(ˆxt)"). Regarding claim 4, the combination of Fedus and Tucker teaches The system of claim 1, wherein the instructions further cause the processor to: based at least in part on the prediction incorrectly identifying the generated synthetic data, control the second neural network to execute a machine learning algorithm to calculate a loss for the second neural network and update parameters.(Fedus [p. 5] "Finally, as in conventional GAN training, our discriminator will be updated according to the gradient" [p. 5] "the gradient to the generator associated with producing ˆxt will depend on all the discounted future rewards (s ≥ t) assigned by the discriminator" [p. 4] "the logarithm of the discriminator estimates are regarded as the reward" [p. 14] "We perform 3 gradient descent steps on the discriminator for every step on the generator and critic" See Eqn. 7 which shows gradient calculated by log loss formula explicitly using the generator G (first neural network)). Regarding claim 5, the combination of Fedus and Tucker teaches The system of claim 4, wherein the instructions further cause the second neural network to update the parameters by subtracting a gradient of the calculated loss and minimizing values in an opposite direction of the gradient.(Fedus [p. 4] "We optimize the parameters of the generator, xt θ, by performing gradient ascent on EG(θ)[R] […] ∇θEG[Rt] = (Rt −bt)∇θlogGθ(ˆxt)"). Regarding claim 8, the combination of Fedus and Tucker teaches The system of claim 1, wherein the instructions further cause the second neural network to: generate the prediction by alternating affine and non-linear activation functions.(Tucker [p. 5] "Our pixel-class CNN outputs class logits zijk for each pixel (i,j), rather than making a direct prediction of the pixel value. The class label cij is sampled from softmax(zij)" CNN layer is affine and softmax is nonlinear). Regarding claim 10, Fedus teaches A computer-implemented method for deep reinforcement learning,([p. 1] "We overcome this by using Reinforcement Learning (RL) to train the generator") the method comprising: generating, by a first neural network implemented on a processor, synthetic data based on original data;([p. 2] "In this task, portions of a body of text are deleted or redacted. The goal of the model is to then infill the missing portions of text so that it is indistinguishable from the original data" [p. 3] "the decoder fills in the missing tokens auto-regressively, however, it is now conditioned on both the masked text m(x)" [p. 2] "Despite the entire sequence now being clearly synthetic as a result of the errant token, a discriminative model that produces a high loss signal to the outlier token, but not to the others, will likely yield a more informative error signal to the generator.") providing the original data and the generated synthetic data to a second neural network implemented on the processor;([p. 3] "For a discrete sequence x =(x1,··· ,xT), a binary mask is generated (deterministically or stochastically) of the same length m=(m1,··· ,mT)where each mt ∈ {0,1}, selects which tokens will remain. The token at time t, xt is then replaced with a special mask token <m> if the mask is 0 and remains unchanged if the mask is 1." mt=1 designates a position whose original token remains, mt=0 designates a position whose original token is removed and subsequently populated by synthetic data) wherein the first neural network and the second neural network are structured to continuously learn using deep reinforcement learning including a state, an action, a reward, and a next state,([p. 4] "the logarithm of the discriminator estimates are regarded as the reward […] The critic estimates the value function, which is the discounted total return of the filled-in sequence Rt = T s=tγsrs, where γ is the discount factor at each position in the sequence [...] the action at is the token chosen by the generator at ≡ ˆxt" [p. 5] "the state st are the current tokens produced up to that point st ≡ ˆx1,··· , ˆxt−1. This approach is an actor-critic architecture where G determines the policy π(st) and the baseline bt is the critic") wherein the state is a value of an input feature in the original data, ([p. 2] "In this task, portions of a body of text are deleted or redacted. The goal of the model is to then infill the missing portions of text so that it is indistinguishable from the original data" [p. 4] "the state st are the current tokens produced up to that point st ≡ ˆx1,··· , ˆxt−1." In Fedus the state is explicitly values of missing tokens in the original data) wherein the original data has a label with a first binary identifier and the synthetic data has a label with a second binary identifier, the first binary identifier and the second binary identifier applied to mask the identity of original data and the synthetic data from the second neural network, respectively;([p. 4] "We give the discriminator the true context, otherwise, this algorithm has a critical failure mode. For instance, without this context, if the discriminator is given the filled-in sequence the director guided the series, it will fail to reliably identify the director bigram as fake text, despite this bigram potentially never appearing in the training corpus (aside from an errant typo) [...] our discriminator Dφ computes the probability of each token ˜xt being real given the true context of the masked sequence m(x)" Fedus explicitly gives the discriminator/second neural network the masked text masking the identity of original data and synthetic data. Fedus explains that without masking the identity the discriminator will fail) the action [is continuous and] includes the synthetic data, ([p. 4] "the action at is the token chosen by the generator at ≡ ˆxt"") the reward is a measure of how unsuccessful the second neural network is in discriminating the original data from the synthetic data, ([p. 4] "our discriminator Dφ computes the probability of each token ˜xt being real given the true context of the masked sequence m(x) [...] the logarithm of the discriminator estimates are regarded as the reward […] The critic estimates the value function, which is the discounted total return of the filled-in sequence Rt = T s=tγsrs, where γ is the discount factor at each position in the sequence") and the next state is a next group of examples to generate a next iteration of the synthetic data based on the original data,([p. 4] "the action at is the token chosen by the generator at ≡ ˆxt" [p. 5] "the state st are the current tokens produced up to that point st ≡ ˆx1,··· , ˆxt−1. This approach is an actor-critic architecture where G determines the policy π(st) and the baseline bt is the critic" [p. 3] "As in standard language-modeling, the decoder fills in the missing tokens auto-regressively" In Fedus the action results in a next state which explicitly is a next group of examples (missing token values) which is then in turn used to generate the subsequent action (synthetic data based on the original data missing token) and this process is explicitly performed iteratively/auto-regressively) generating, by the second neural network, a prediction identifying the original data and the generated synthetic data; ([p. 4] "our discriminator Dφ computes the probability of each token ˜xt being real given the true context of the masked sequence m(x)") and based at least in part on the prediction incorrectly identifying the generated synthetic data, exporting the generated synthetic data ([p. 4] "our discriminator Dφ computes the probability of each token ˜xt being real given the true context of the masked sequence m(x) [...] the logarithm of the discriminator estimates are regarded as the reward […] The critic estimates the value function, which is the discounted total return of the filled-in sequence Rt = T s=tγsrs, where γ is the discount factor at each position in the sequence" [p. 6] "We present both conditional and unconditional samples generated on the PTB and IMDB data sets at word-level. MaskGAN refers to our GAN-trained variant and MaskMLE refers to our maximum likelihood trained variant. Additional samples are supplied in Appendix B." [p. 2] "the generation of synthetic training data"). However, Fedus does not explicitly teach the action is continuous and includes the synthetic data. Tucker, in the same field of endeavor, teaches the action is continuous and includes the synthetic data, ([p. 3 §3] "We build on adversarial IRL (AIRL), an algorithm achieving state-of-the-art performance on simulated robotics tasks (Fu et al., 2017). The reference implementation of AIRL assumes a continuous action space and uses a fully-connected policy and reward network [...] Adversarial IRL formulates the inverse reinforcement learning problem as a GAN (Goodfellow et al., 2014). We learn a reward function fθ(s,a) for taking action a in state s and a stochastic policy π(a | s). The policy is the generator, and is trained using forward RL on the reward function fθ(s,a). The discriminator is restricted to have the special form"). Fedus as well as Tucker are directed towards hybrid reinforcement learning generative adversarial neural network models. Therefore, Fedus as well as Tucker are analogous art in the same field of endeavor. It would have been obvious before the effective filing date of the claimed invention to combine the teachings of Fedus with the teachings of Tucker by using a continuous action space. One of ordinary skill in the art would realize that the action space must be either discrete or continuous and Tucker explicitly acknowledges continuous action space in reinforcement learning generative adversarial networks and provides as additional motivation for combination ([p. 1 §1] “Recent deep IRL algorithms have achieved good performance on a variety of continuous control tasks”). This motivation for combination also applies to the remaining claims which depend on this combination. Regarding claim 11, the combination of Fedus and Tucker teaches The computer-implemented method of claim 10, further comprising: based at least in part on the prediction correctly identifying the generated synthetic data, executing, by the first neural network, a machine learning model to calculate a loss for the generated synthetic data (Fedus [p. 5] "Finally, as in conventional GAN training, our discriminator will be updated according to the gradient" [p. 5] "the gradient to the generator associated with producing ˆxt will depend on all the discounted future rewards (s ≥ t) assigned by the discriminator" [p. 4] "the logarithm of the discriminator estimates are regarded as the reward" [p. 14] "We perform 3 gradient descent steps on the discriminator for every step on the generator and critic" See Eqn. 7 which shows gradient calculated by log loss formula explicitly using the generator G (first neural network)) and update parameters for the generated synthetic data; and generating, by the first neural network, a second synthetic data.(Fedus [p. 5] "the gradient to the generator associated with producing ˆxt will depend on all the discounted future rewards (s ≥ t) assigned by the discriminator" [p. 4] "the logarithm of the discriminator estimates are regarded as the reward" [p. 14] "We perform 3 gradient descent steps on the discriminator for every step on the generator and critic" Fedus is explicit that the generator and discriminator are trained iteratively and directly on one another’s updates). Regarding claim 12, the combination of Fedus and Tucker teaches The computer-implemented method of claim 11, further comprising: updating, by the first neural network, the parameters by subtracting a gradient of the calculated loss and minimizing values in an opposite direction of the gradient.(Fedus [p. 4] "We optimize the parameters of the generator, xt θ, by performing gradient ascent on EG(θ)[R] […] ∇θEG[Rt] = (Rt −bt)∇θlogGθ(ˆxt)"). Regarding claim 13, the combination of Fedus and Tucker teaches The computer-implemented method of claim 10, further comprising: based at least in part on the prediction incorrectly identifying the generated synthetic data, executing, by the second neural network, a machine learning algorithm to calculate a loss for the second neural network and update parameters for the second neural network.(Fedus [p. 5] "the gradient to the generator associated with producing ˆxt will depend on all the discounted future rewards (s ≥ t) assigned by the discriminator" [p. 4] "the logarithm of the discriminator estimates are regarded as the reward" [p. 14] "We perform 3 gradient descent steps on the discriminator for every step on the generator and critic" Fedus is explicit that the generator and discriminator are trained iteratively and directly on one another’s updates). Regarding claim 14, the combination of Fedus and Tucker teaches The computer-implemented method of claim 13, further comprising: updating, by the second neural network, the parameters by subtracting a gradient of the calculated loss and minimizing values in an opposite direction of the gradient.(Fedus [p. 4] "We optimize the parameters of the generator, xt θ, by performing gradient ascent on EG(θ)[R] […] ∇θEG[Rt] = (Rt −bt)∇θlogGθ(ˆxt)"). Regarding claim 16, the combination of Fedus and Tucker teaches The computer-implemented method of claim 10, wherein generating the prediction further comprises: alternating affine and non-linear activation functions.(Tucker [p. 5] "Our pixel-class CNN outputs class logits zijk for each pixel (i,j), rather than making a direct prediction of the pixel value. The class label cij is sampled from softmax(zij)" CNN layer is affine and softmax is nonlinear). Regarding claim 19, Fedus teaches One or more computer-storage memory devices embodied with executable operations that, when executed by a processor, cause the processor to:([p. 3] "Our generator consists of an encoding module and decoding module." [p. 4] "MaskGAN source code available at: https://github.com/tensorflow/models/tree/ master/research/maskgan" Generator interpreted as first neural network) receive an original data group; generate, by a first neural network, a first synthetic data group based on the original data group;([p. 2] "In this task, portions of a body of text are deleted or redacted. The goal of the model is to then infill the missing portions of text so that it is indistinguishable from the original data" [p. 3] "the decoder fills in the missing tokens auto-regressively, however, it is now conditioned on both the masked text m(x)" [p. 2] "Despite the entire sequence now being clearly synthetic as a result of the errant token, a discriminative model that produces a high loss signal to the outlier token, but not to the others, will likely yield a more informative error signal to the generator.") wherein the original data group has a label with a first binary identifier and the first synthetic data group has a label with a second binary identifier, ([p. 3] "For a discrete sequence x =(x1,··· ,xT), a binary mask is generated (deterministically or stochastically) of the same length m=(m1,··· ,mT)where each mt ∈ {0,1}, selects which tokens will remain. The token at time t, xt is then replaced with a special mask token <m> if the mask is 0 and remains unchanged if the mask is 1." mt=1 designates a position whose original token remains, mt=0 designates a position whose original token is removed and subsequently populated by synthetic data) the first binary identifier and the second binary identifier applied to mask the identity of original data and the synthetic data from the second neural network, respectively;([p. 4] "We give the discriminator the true context, otherwise, this algorithm has a critical failure mode. For instance, without this context, if the discriminator is given the filled-in sequence the director guided the series, it will fail to reliably identify the director bigram as fake text, despite this bigram potentially never appearing in the training corpus (aside from an errant typo) [...] our discriminator Dφ computes the probability of each token ˜xt being real given the true context of the masked sequence m(x)" Fedus explicitly gives the discriminator/second neural network the masked text masking the identity of original data and synthetic data. Fedus explains that without masking the identity the discriminator will fail) provide the original data group and the generated first synthetic data group to a second neural network different from the first neural network; ([p. 3] "The discriminator is given the filled-in sequence from the generator" [p. 6] "we then pretrain the discriminator on the samples produced from the current generator and real training text" Fedus explicitly trains the second neural network/discriminator on both real training text/original data and generated synthetic data) generate, by a second neural network, a first prediction identifying the original data group and the generated first synthetic data group;([p. 4] "our discriminator Dφ computes the probability of each token ˜xt being real given the true context of the masked sequence m(x)") based at least in part on the first prediction correctly identifying the generated synthetic data group: execute, by the first neural network, a first machine learning (ML) model to calculate a loss for the generated first synthetic data group, update parameters, ([p. 5] "Finally, as in conventional GAN training, our discriminator will be updated according to the gradient" [p. 5] "the gradient to the generator associated with producing ˆxt will depend on all the discounted future rewards (s ≥ t) assigned by the discriminator" [p. 4] "the logarithm of the discriminator estimates are regarded as the reward" [p. 14] "We perform 3 gradient descent steps on the discriminator for every step on the generator and critic" See Eqn. 7 which shows gradient calculated by log loss formula explicitly using the generator G (first neural network)) and, using the updated parameters, generate a second synthetic data group, wherein the generated second synthetic data group is a second iteration of the generated first synthetic data group based on the original data group, and([p. 4] "the action at is the token chosen by the generator at ≡ ˆxt" [p. 5] "the state st are the current tokens produced up to that point st ≡ ˆx1,··· , ˆxt−1. This approach is an actor-critic architecture where G determines the policy π(st) and the baseline bt is the critic" [p. 3] "As in standard language-modeling, the decoder fills in the missing tokens auto-regressively" In Fedus the action results in a next state which explicitly is a next group of examples (missing token values) which is then in turn used to generate the subsequent action (synthetic data based on the original data missing token) and this process is explicitly performed iteratively/auto-regressively) execute, by the second neural network, a second ML model to calculate a loss for the second neural network and update parameters;([p. 5] "Finally, as in conventional GAN training, our discriminator will be updated according to the gradient" [p. 5] "the gradient to the generator associated with producing ˆxt will depend on all the discounted future rewards (s ≥ t) assigned by the discriminator" [p. 4] "the logarithm of the discriminator estimates are regarded as the reward" [p. 14] "We perform 3 gradient descent steps on the discriminator for every step on the generator and critic" See Eqn. 7 which shows gradient calculated by log loss formula explicitly using the generator G (first neural network)) provide the original data group and the generated second synthetic data group to the second neural network;([p. 4] "the action at is the token chosen by the generator at ≡ ˆxt" [p. 5] "the state st are the current tokens produced up to that point st ≡ ˆx1,··· , ˆxt−1. This approach is an actor-critic architecture where G determines the policy π(st) and the baseline bt is the critic" [p. 3] "As in standard language-modeling, the decoder fills in the missing tokens auto-regressively" In Fedus the action results in a next state which explicitly is a next group of examples (missing token values) which is then in turn used to generate the subsequent action (synthetic data based on the original data missing token) and this process is explicitly performed iteratively/auto-regressively) wherein the first neural network and the second neural network are structured to continuously learn using deep reinforcement learning including a state, an action, a reward, and a next state,([p. 4] "the logarithm of the discriminator estimates are regarded as the reward […] The critic estimates the value function, which is the discounted total return of the filled-in sequence Rt = T s=tγsrs, where γ is the discount factor at each position in the sequence [...] the action at is the token chosen by the generator at ≡ ˆxt" [p. 5] "the state st are the current tokens produced up to that point st ≡ ˆx1,··· , ˆxt−1. This approach is an actor-critic architecture where G determines the policy π(st) and the baseline bt is the critic") wherein the state is a value of an input feature in the original data group, ([p. 2] "In this task, portions of a body of text are deleted or redacted. The goal of the model is to then infill the missing portions of text so that it is indistinguishable from the original data" [p. 4] "the state st are the current tokens produced up to that point st ≡ ˆx1,··· , ˆxt−1." In Fedus the state is explicitly values of missing tokens in the original data) the action [is continuous and] includes the first synthetic data group,([p. 4] "the action at is the token chosen by the generator at ≡ ˆxt"") the reward is a measure of how unsuccessful the second neural network is in discriminating the original data group from the first synthetic data group, ([p. 4] "our discriminator Dφ computes the probability of each token ˜xt being real given the true context of the masked sequence m(x) [...] the logarithm of the discriminator estimates are regarded as the reward […] The critic estimates the value function, which is the discounted total return of the filled-in sequence Rt = T s=tγsrs, where γ is the discount factor at each position in the sequence") and the next state is a next group of examples to generate the second synthetic data group based on the second iteration of the generated first synthetic data group based on the original data group;([p. 4] "the action at is the token chosen by the generator at ≡ ˆxt" [p. 5] "the state st are the current tokens produced up to that point st ≡ ˆx1,··· , ˆxt−1. This approach is an actor-critic architecture where G determines the policy π(st) and the baseline bt is the critic" [p. 3] "As in standard language-modeling, the decoder fills in the missing tokens auto-regressively" In Fedus the action results in a next state which explicitly is a next group of examples (missing token values) which is then in turn used to generate the subsequent action (synthetic data based on the original data missing token) and this process is explicitly performed iteratively/auto-regressively) generate, by the second neural network, a second prediction identifying the original data group and the generated second synthetic data group; and([p. 4] "our discriminator Dφ computes the probability of each token ˜xt being real given the true context of the masked sequence m(x)") based at least in part on the second prediction incorrectly identifying the generated second synthetic data group, export the generated second synthetic data group.([p. 4] "our discriminator Dφ computes the probability of each token ˜xt being real given the true context of the masked sequence m(x) [...] the logarithm of the discriminator estimates are regarded as the reward […] The critic estimates the value function, which is the discounted total return of the filled-in sequence Rt = T s=tγsrs, where γ is the discount factor at each position in the sequence" [p. 6] "We present both conditional and unconditional samples generated on the PTB and IMDB data sets at word-level. MaskGAN refers to our GAN-trained variant and MaskMLE refers to our maximum likelihood trained variant. Additional samples are supplied in Appendix B." [p. 2] "the generation of synthetic training data"). However, Fedus does not explicitly teach the action is continuous and includes the first synthetic data group,. Tucker, in the same field of endeavor, teaches the action is continuous and includes the first synthetic data group,([p. 3 §3] "We build on adversarial IRL (AIRL), an algorithm achieving state-of-the-art performance on simulated robotics tasks (Fu et al., 2017). The reference implementation of AIRL assumes a continuous action space and uses a fully-connected policy and reward network [...] Adversarial IRL formulates the inverse reinforcement learning problem as a GAN (Goodfellow et al., 2014). We learn a reward function fθ(s,a) for taking action a in state s and a stochastic policy π(a | s). The policy is the generator, and is trained using forward RL on the reward function fθ(s,a). The discriminator is restricted to have the special form"). Fedus as well as Tucker are directed towards hybrid reinforcement learning generative adversarial neural network models. Therefore, Fedus as well as Tucker are analogous art in the same field of endeavor. It would have been obvious before the effective filing date of the claimed invention to combine the teachings of Fedus with the teachings of Tucker by using a continuous action space. One of ordinary skill in the art would realize that the action space must be either discrete or continuous and Tucker explicitly acknowledges continuous action space in reinforcement learning generative adversarial networks and provides as additional motivation for combination ([p. 1 §1] “Recent deep IRL algorithms have achieved good performance on a variety of continuous control tasks”). This motivation for combination also applies to the remaining claims which depend on this combination. Claims 7 and 17 are rejected under U.S.C. §103 as being unpatentable over the combination of Fedus and Tucker and in further view of Luong (US20210089724A1). Regarding claim 7, the combination of Fedus, and Tucker teaches The system of claim 1. However, the combination of Fedus, and Tucker doesn't explicitly teach wherein the instructions further cause the first neural network to generate the synthetic data by: distorting feature values of the original data to introduce noise. Luong, in the same field of endeavor, teaches the instructions further cause the first neural network to generate the synthetic data by: distorting feature values of the original data to introduce noise. ([¶0050] "a “noised” example xnoised 22 can be created by replacing the masked-out tokens 20 a and 20 b with generator samples. The discriminator 12 can then be trained to predict which tokens in xnoised 22 do not match the original input x 18."). The combination of Fedus, and Tucker as well as Luong are directed towards reinforcement learning generative adversarial networks. Therefore, the combination of Fedus, and Tucker as well as Luong are analogous art in the same field of endeavor. It would have been obvious before the effective filing date of the claimed invention to combine the teachings of the combination of Fedus, and Tucker with the teachings of Luong by using a noise distribution over masked input data. Luong provides as additional motivation for combination ([¶0049] "For a given position t, the discriminator 12 predicts whether the token xt is “real,” i.e., that it comes from the data distribution rather than the generator distribution (e.g., a noise distribution)"). This motivation for combination also applies to the remaining claims which depend on this combination. Regarding claim 17, the combination of Fedus, and Tucker teaches The computer-implemented method of claim 10. However, the combination of Fedus, and Tucker doesn't explicitly teach, wherein generating the synthetic data further comprises: distorting feature values of the original data to introduce noise.. Luong, in the same field of endeavor, teaches generating the synthetic data further comprises: distorting feature values of the original data to introduce noise. ([¶0050] "a “noised” example xnoised 22 can be created by replacing the masked-out tokens 20 a and 20 b with generator samples. The discriminator 12 can then be trained to predict which tokens in xnoised 22 do not match the original input x 18."). The combination of Fedus, and Tucker as well as Luong are directed towards reinforcement learning generative adversarial networks. Therefore, the combination of Fedus, and Tucker as well as Luong are analogous art in the same field of endeavor. It would have been obvious before the effective filing date of the claimed invention to combine the teachings of the combination of Fedus, and Tucker with the teachings of Luong by using a noise distribution over masked input data. Luong provides as additional motivation for combination ([¶0049] "For a given position t, the discriminator 12 predicts whether the token xt is “real,” i.e., that it comes from the data distribution rather than the generator distribution (e.g., a noise distribution)"). This motivation for combination also applies to the remaining claims which depend on this combination. Claim 9 is rejected under U.S.C. §103 as being unpatentable over the combination of Fedus, Tucker, and Mathworks (“Monitor GAN Training Progress and Identify Common Failure Modes”, 2021). Regarding claim 9, the combination of Fedus, and Tucker teaches The system of claim 1, wherein the generated synthetic data is exported for training an external machine learning model based at least in part on the prediction incorrectly identifying the generated synthetic data a number of times(Fedus [p. 4] "our discriminator Dφ computes the probability of each token ˜xt being real given the true context of the masked sequence m(x) [...] the logarithm of the discriminator estimates are regarded as the reward […] The critic estimates the value function, which is the discounted total return of the filled-in sequence Rt = T s=tγsrs, where γ is the discount factor at each position in the sequence" [p. 6] "We present both conditional and unconditional samples generated on the PTB and IMDB data sets at word-level. MaskGAN refers to our GAN-trained variant and MaskMLE refers to our maximum likelihood trained variant. Additional samples are supplied in Appendix B." [p. 2] "the generation of synthetic training data"). However, the combination of Fedus, and Tucker doesn't explicitly teach a number of times that exceeds a threshold greater than one. Mathworks, in the same field of endeavor, teaches the generated synthetic data is exported for training an external machine learning model based at least in part on the prediction incorrectly identifying the generated synthetic data a number of times that exceeds a threshold greater than one.([p. 1] "Convergence Failure Convergence failure happens when the generator and discriminator do not reach a balance during training. Discriminator Dominates This scenario happens when the generator score reaches zero or near zero and the discriminator score reaches one or near one. This plot shows an example of the discriminator overpowering the generator. Notice that the generator score approaches zero and does not recover. In this case, the discriminator classifies most of the images correctly. In turn, the generator cannot produce any images that fool the discriminator and thus fails to learn." Mathworks Matlab GAN documentation explicitly shows that convergence fails if the discriminator does not incorrectly identify the generated synthetic data a number of times that exceeds a threshold greater than one. Note that Mathworks also acknowledges that in discriminator dominated convergence failure situations the discriminator still fails to correctly identify all of the generated synthetic data.). The combination of Fedus and Tucker as well as Mathworks are directed towards generative adversarial networks. Therefore, the combination of Fedus and Tucker as well as Zhang are analogous art in the same field of endeavor. It would have been obvious before the effective filing date of the claimed invention to combine the teachings of the combination of Fedus and Tucker with the teachings of Mathworks by training the GAN based on the discriminator being fooled more than once. Mathworks describes this as not only expected but describes any situation where this does not occur as common convergence failure. Claims 18 and 20 are rejected under U.S.C. §103 as being unpatentable over the combination of Fedus and Tucker and Huang (US11586911B2). Regarding claim 18, the combination of Fedus, and Tucker teaches The computer-implemented method of claim 10. However, the combination of Fedus, and Tucker doesn't explicitly teach, wherein the first neural network. the second neural network, and the original data are physically co-located or located within a same geographic region. Huang, in the same field of endeavor, teaches The computer-implemented method of claim 10, wherein the first neural network. the second neural network, and the original data are physically co-located or located within a same geographic region. ([Col. 8 l. 16-18] "System 10 includes GAN 28 that includes generator 30 and discriminator 32." Huang explicitly places the generator and discriminator together inside the same system 10). The combination of Fedus, and Tucker as well as Huang are directed towards reinforcement learning generative adversarial networks. Therefore, the combination of Fedus, and Tucker as well as Huang are analogous art in the same field of endeavor. It would have been obvious before the effective filing date of the claimed invention to combine the teachings of the combination of Fedus, and Tucker with the teachings of Huang by having the GAN generator and discriminator on the same system. Huang provides as additional motivation for combination ([Col. 11 l. 4-10] “the arrangements described herein train RL agent 36 such that RL agent 36 is able to provide better performance after several samples than other arrangements”). Regarding claim 20, the combination of Fedus, and Tucker teaches The one or more computer-storage memory devices of claim 19. However, the combination of Fedus, and Tucker doesn't explicitly teach, wherein the processor further: exports the generated second synthetic data group to a third ML model. Huang, in the same field of endeavor, teaches The one or more computer-storage memory devices of claim 19, wherein the processor further: exports the generated second synthetic data group to a third ML model. ([Col. 6 l. 65-67] "deep neural network (DNN) is added to learn this relation and enforce the data generated by the GAN" DNN interpreted as third ML model that generated second synthetic data group (experience data) is exported to). The combination of Fedus, and Tucker as well as Huang are directed towards reinforcement learning generative adversarial networks. Therefore, the combination of Fedus, and Tucker as well as Huang are analogous art in the same field of endeavor. It would have been obvious before the effective filing date of the claimed invention to combine the teachings of the combination of Fedus, and Tucker with the teachings of Huang by exporting the synthetic data to a third model. Huang provides as additional motivation for combination ([Col. 11 l. 4-10] “the arrangements described herein train RL agent 36 such that RL agent 36 is able to provide better performance after several samples than other arrangements”). Allowable Subject Matter Claims 6 and 15 are objected to as being dependent upon rejected base claims, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. Below are the closest cited references, each of which disclose various aspects of the claimed invention: Fedus (“MaskGAN: Better Text Generation via Filling in the ______”, 2018) Tucker (“Inverse reinforcement learning for video games”, 2018) Luong (US20210089724A1) Huang (US11586911B2) However, none of the prior art references of record, alone or in combination, disclose or suggest the combined features recited in the independent claims, including specifically (for claim 6): The system of claim 1, wherein the instructions further cause the processor to: randomly assign the first binary identifier and the second binary identifier to a label to each of the original data and the generated synthetic data, respectively; receive the prediction generated by the second neural network; compare the prediction generated by the second neural network to the first binary identifier and the second binary identifier labels randomly assigned to the original data and the generated synthetic data; based on the comparison, output a first label and a second label when the prediction is correct and incorrect, respectively; and append the synthetic data with the first or second label. While Fedus discloses reinforcement learning for a generative adversarial network utilizing a masked discriminator, Fedus does not disclose the order of operations in instant claim 6, and specifically does not disclose to “compare the prediction generated by the second neural network to the first binary identifier and the second binary identifier labels randomly assigned to the original data and the generated synthetic data; based on the comparison, output a first label and a second label when the prediction is correct and incorrect, respectively; and append the synthetic data with the first or second label.”. It would not have been obvious before the effective filing date of the claimed invention to combine Tucker, Luong, and/or Huang with Fedus to arrive at the claimed invention. Dependent claim 15 recites analogous limitations to that identified above and is allowable for the same reason. Conclusion THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any extension fee pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to SIDNEY VINCENT BOSTWICK whose telephone number is (571)272-4720. The examiner can normally be reached M-F 7:30am-5:00pm EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Miranda Huang can be reached on (571)270-7092. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /SIDNEY VINCENT BOSTWICK/Examiner, Art Unit 2124
Read full office action

Prosecution Timeline

Show 8 earlier events
Mar 02, 2026
Request for Continued Examination
Mar 11, 2026
Response after Non-Final Action
Apr 28, 2026
Non-Final Rejection mailed — §103
Jun 05, 2026
Interview Requested
Jun 16, 2026
Examiner Interview Summary
Jun 16, 2026
Applicant Interview (Telephonic)
Jul 28, 2026
Response Filed
Sep 23, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12699874
Leveraging Redundancy in Attention with Reuse Transformers
3y 10m to grant Granted Aug 04, 2026
Patent 12675673
NEURAL NETWORK PROCESSING DEVICE, METHOD, AND COMPUTER-READABLE RECORDING MEDIUM
3y 7m to grant Granted Jul 07, 2026
Patent 12645914
INSTRUCTION PRUNING FOR NEURAL NETWORKS
3y 6m to grant Granted Jun 02, 2026
Patent 12626139
SECRET SOFTMAX FUNCTION CALCULATION SYSTEM, SECRET SOFTMAX FUNCTION CALCULATION APPARATUS, SECRET SOFTMAX FUNCTION CALCULATION METHOD, SECRET NEURAL NETWORK CALCULATION SYSTEM, SECRET NEURAL NETWORK LEARNING SYSTEM, AND PROGRAM
4y 3m to grant Granted May 12, 2026
Patent 12619815
Magnitude Invariant Multimodal Agent for Efficient Image-Text Interface Automation
1y 6m to grant Granted May 05, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

5-6
Expected OA Rounds
51%
Grant Probability
86%
With Interview (+35.1%)
4y 5m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 152 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month