Prosecution Insights
Last updated: October 01, 2026
Application No. 18/534,206

SYNTHETIC TRAINING DATA FOR GENERATIVE MODELS

Final Rejection §103
Filed
Dec 08, 2023
Examiner
DUNAY, CHRISTOPHER E
Art Unit
Tech Center
Assignee
Google LLC
OA Round
2 (Final)
76%
Grant Probability
Favorable
3-4
OA Rounds
0m
Est. Remaining
90%
With Interview

Examiner Intelligence

Grants 76% — above average
76%
Career Allowance Rate
576 granted / 754 resolved
+16.4% vs TC avg
Moderate +14% lift
Without
With
+14.0%
Interview Lift
resolved cases with interview
Fast prosecutor
1y 10m
Avg Prosecution
21 currently pending
Career history
775
Total Applications
across all art units

Statute-Specific Performance

§101
0.4%
-39.6% vs TC avg
§103
52.6%
+12.6% vs TC avg
§102
21.9%
-18.1% vs TC avg
§112
21.4%
-18.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 754 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Information Disclosure Statement The information disclosure statements (IDS) submitted on 4/8/2025 and 5/19/2025 were filed and are being considered by the examiner. Examiner Comment The Hanze Dong et al reference appears to teach the applicant’s invention and any other claims are obvious known features in machine learning. This reference was found during PCT examination. The Examiner rejects the claims with this reference, as the claims were not amended from the PCT. The applicant must answer to this reference before advancing prosecution further. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1-19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Hanze Dong et al (RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment). In regard to claim 1, Hanze Dong et al disclose a method implemented by one or more processors, the method comprising: for each of a plurality of sets of input data: generating, using a machine-learned generative model, a plurality of generative outputs from a set of input data; determining, using a machine-learned reward model, a plurality of rewards from the plurality of generative outputs, each reward associated with one or more of the plurality of generative outputs; generating, for inclusion in a machine-learning training dataset, a training example comprising the respective input data, a positive training example and a negative training example, comprising: selecting, from the plurality of generative outputs, a first generative output as the positive training example in the machine-learning training dataset based on a reward associated with the first generative output; and selecting, from the plurality of generative outputs, a second generative output as the negative training example in the machine-learning training dataset based on a reward associated with the second generative output, wherein the reward associated with the first generative output and the reward associated with the second generative output indicate that the first generative output is preferred over the second generative output. In regard to the amendment filed 8/24/2026, Hanze Dong et al fail to disclose based on processing the plurality of generative outputs. However, this limitation is very broad. There is clear some processing of the generative outputs. In regard to claim 4, Hanze Dong et al disclose using the machine-learned generative model, the plurality of generative outputs from a respective set of input data comprises: generating a distribution over a set of potential generative outputs; and sampling the plurality of generative outputs from the distribution. In regard to claim 7, Hanze Dong et al disclose the reward model is a pointwise reward model and wherein determining, using the machine-learned reward model, the plurality of rewards from the plurality of generative outputs comprises:determining, using the reward model, a respective reward for each of the plurality of generative outputs; and generating, based on the rewards, a ranking of the plurality of generative outputs, wherein the first generative output is ranked higher in the ranking than the second generative output. In regard to claim 8, Hanze Dong et al disclose the first generative output is a highest ranked generative output from the plurality of generative outputs. In regard to claim 9, Hanze Dong et al disclose the second generative output is a lowest ranked generative output from the plurality of generative outputs. In regard to claim 10, Hanze Dong et al disclose the reward model is a pairwise reward model and wherein determining, using the machine-learned reward model, the plurality of rewards from the plurality of generative outputs comprises:determining a respective reward for each of a plurality of pairs of generative outputs, wherein the respective reward for a pair of generative outputs indicates a probability that a first generative output of the pair of generative outputs is preferred over a second generative output of the pair. In regard to claim 11, Hanze Dong et al disclose the first generative output corresponds to a first generative output of a pair of generative outputs with the highest reward; and the second generative output corresponds to a second generative output of the pair of generative outputs with the highest reward. In regard to claim 12, Hanze Dong et al disclose generating a training example comprises performing a two- way tournament between the plurality of pairs of generative outputs to determine the pair of generative outputs with the highest reward. In regard to claim 15, Hanze Dong et al disclose training the reward model or a further reward model based on the machine-learning training dataset. In regard to claim 16, Hanze Dong et al disclose training a generative machine-learning model using the reward model. In regard to claim 17, Hanze Dong et al disclose distilling a student reward model from the machine-learned reward model, the distilling comprising: training the student reward model based on the machine-learning training dataset. In regard to claim 2 and 3, Hanze Dong et al fail to disclose that the machine-learned generative model is an image generation model, and wherein the plurality of generative outputs comprises a plurality of images or that the machine-learned generative model is a large language model, wherein each respective set of input data comprises an input prompt, and wherein the plurality of generative outputs comprises a plurality of text sequences. However, images and text are obvious uses for the model described by Hanze Dong et al and it would have been obvious to one of ordinary skill in the art at the time of filing to use images or text in order to improve models using images or text. In regard to claim 5 and 6, Hanze Dong et al fail to disclose for each of one or more training examples: determining a confidence value that the first generative output is preferred over the second generative output; determining that the confidence value does not satisfy a threshold confidence value; and in response to determining that the confidence value does not satisfy the threshold confidence value, discarding the training example from the machine-learning training dataset or that for each of one or more training examples: determining a first likelihood value for that the first generative output of the training example and/or second likelihood value for the second generative output; determining that the first likelihood value and/or second likelihood value does not satisfy a threshold likelihood value; and in response to determining that the first likelihood value and/or second likelihood value does not satisfy the threshold likelihood value, discarding the training example from the machine-learning training dataset. However, filtering data is notoriously old and well-known, and it would have been obvious to one of ordinary skill in the art at the time of filing to filter the data in order to only use clear cases where there is a positive and negative case. In regard to claim 13, Hanze Dong et al fail to disclose combining the training examples with human labeled training examples to generate the machine-learning training dataset. However, it would have been obvious to one of ordinary skill in the art at the time of filing to combine human labeled data in order to improve the model. In regard to claim 14, Hanze Dong et al fail to disclose one or more of the plurality of sets of input data comprises a multimodal input comprising two or more of: one or more images; a text sequence; one or more audio samples; and/or one or more videos. However, it would have been obvious to one of ordinary skill in the art at the time of filing to use multimodal input data in order to use input data that is multimodal. In regard to claim 18, Hanze Dong et al fail to disclose the student reward model has a memory footprint below a threshold memory usage. However, it would have been obvious to one of ordinary skill in the art at the time of filing to use a memory limited distilled model in order to limit the size of the student model for deployment. In regard to claim 19, Hanze Dong et al disclose generate a machine-learning training dataset; and train a reward model, or a further reward model, based on the machine-learning training dataset; wherein in generating the machine-learning training dataset one or more of the processors are to: for each of a plurality of sets of input data: generate, using a machine-learned generative model, a plurality of generative outputs from a set of input data; determine, using the reward model, a plurality of rewards from the plurality of generative outputs, each reward associated with one or more of the plurality of generative outputs; generate, for inclusion in the machine-learning training dataset, a training example comprising the respective input data, a positive training example and a negative training example, wherein in generating the training example one or more of the processors are to: select, from the plurality of generative outputs, a first generative output as the positive training example in the machine-learning training dataset based on a reward associated with the first generative output; and select, from the plurality of generative outputs, a second generative output as the negative training example in the machine-learning training dataset based on a reward associated with the second generative output, wherein the reward associated with the first generative output and the reward associated with the second generative output indicate that the first generative output is preferred over the second generative output. Hanze Dong et al fail to explicitly disclose processors and memory. However, when the model is implemented, it would use processors and memory, and it would have been obvious to one of ordinary skill in the art at the time of filing to use processors and memory in order to implement the model on a computer. In regard to the amendment filed 8/24/2026, Hanze Dong et al fail to disclose based on processing the plurality of generative outputs. However, this limitation is very broad. There is clear some processing of the generative outputs. Response to Arguments Applicant's arguments filed 8/24/2026 have been fully considered but they are not persuasive. “…based on processing the plurality of generative outputs…” is a very broad limitation. There is inherently some “processing” of the outputs. Conclusion THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to CHRISTOPHER E DUNAY whose telephone number is (571)270-1222. The examiner can normally be reached 7:00 am - 6:00 pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, James (Jong-Suk) Lee can be reached at 571-272-7044. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /CHRISTOPHER E DUNAY/ Primary Examiner, Art Unit 2875
Read full office action

Prosecution Timeline

Dec 08, 2023
Application Filed
May 22, 2026
Non-Final Rejection mailed — §103
Aug 24, 2026
Response Filed
Sep 03, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12743061
PROMPT ENGINEERING FOR ARTIFICIAL INTELLIGENCE ASSISTED INDUSTRIAL AUTOMATION DEVICE CONFIGURATION
3y 2m to grant Granted Sep 22, 2026
Patent 12737438
ROBUST TRAJECTORY PREDICTIONS AGAINST ADVERSARIAL ATTACKS IN AUTONOMOUS MACHINES AND APPLICATIONS
3y 6m to grant Granted Sep 15, 2026
Patent 12736278
REFRIGERATOR AND HOME APPLIANCE
3y 1m to grant Granted Sep 15, 2026
Patent 12740199
ANTIOXIDANTS, BACKLIGHT MODULES AND MANUFACTURING METHOD THEREOF
2y 10m to grant Granted Sep 15, 2026
Patent 12736614
VEHICLE EXTERIOR COMPONENT
2y 6m to grant Granted Sep 15, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
76%
Grant Probability
90%
With Interview (+14.0%)
1y 10m (~0m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 754 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month