Prosecution Insights
Last updated: October 02, 2026
Application No. 18/468,046

SYSTEMS AND METHODS FOR GENERATING IMAGES USING DIFFUSION MODEL

Non-Final OA §103
Filed
Sep 15, 2023
Examiner
WU, YANNA
Art Unit
2615
Tech Center
2600 — Communications
Assignee
Woven By Toyota Inc.
OA Round
4 (Non-Final)
81%
Grant Probability
Favorable
4-5
OA Rounds
0m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 81% — above average
81%
Career Allowance Rate
369 granted / 456 resolved
+18.9% vs TC avg
Strong +34% interview lift
Without
With
+34.4%
Interview Lift
resolved cases with interview
Fast prosecutor
2y 2m
Avg Prosecution
23 currently pending
Career history
474
Total Applications
across all art units

Statute-Specific Performance

§101
9.8%
-30.2% vs TC avg
§103
69.7%
+29.7% vs TC avg
§102
6.8%
-33.2% vs TC avg
§112
8.1%
-31.9% vs TC avg
Black line = Tech Center average estimate • Based on career data from 456 resolved cases

Office Action

§103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . DETAILED ACTION This is in response to applicant’s amendment/response filed on 08/03/2026, which has been entered and made of record. Claim 1, 3-5, 14-15 are amended. Claims 1-5, 11-15 are pending in the application. Applicant is reminded to change the status of the withdrawn claims 6-10 and 16-20 to be cancelled. Response to Arguments Applicant arguments regarding claim rejections under 103 are considered, but are not persuasive. Applicant argues: PNG media_image1.png 148 654 media_image1.png Greyscale Examiner disagrees: Li teaches using a machine model to generate the synthetic object image as shown in FIG. 2 and also [0183]. Li in [0127] also teaches that the machine learning model can be any forms of machines learning models. Although it does not list all the possible machine learning models, e.g. a diffusion model, it would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to have chosen a diffusion model from a finite number of existing machine learning models to perform image generation with a reasonable expectation of success. Applicant argues: FIG. 4C 404 of Li does not discloses a location of road lane. Examiner disagrees: Li explicitly say FIG. 4C, 404 indicates the road markers in an image, which means the two line indicates the location of a road in the image. Furthermore, the Specification FIG. 2 also uses to lines to indicate the location of road in an image, which is consistent with the teachings of Li. The rest of arguments are moot in view of new ground of rejections. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1-5, 11-15 is/are rejected under 35 U.S.C. 103 as being unpatentable over Li et al. (US 2021/0261148 A1) in view of Rong et al. (US 2021/0383616 A1) and further in view of Xu et al. (US 2002/0175921 A1). Regarding claim 1, Li teaches: A method for generating a synthetic image, the method comprising: receiving an input image, wherein the receiving the input image comprises generating the input image based on a machine learning model; ([0053], “In an exemplary embodiment, the data reception module 118 of the action prediction application 106 may be configured to receive image data that may be associated with images captured of the surrounding environment of the ego vehicle 102 that may be provided by the vehicle camera system 110 of the ego vehicle 102.” [0054], “In one embodiment, the data reception module 118 may be configured to input the image frames 202 to the neural network 108 to be analyzed by I3D 204. The I3D 204 is configured to apply instance segmentation and semantic segmentation 210 to detect dynamic objects located within the driving scene and the driving scene characteristics of the driving scene.”) removing at least one pre-existing object from the input image; ([0057], “In an exemplary embodiment, the neural network 108 may complete image inpainting to electronically remove and replace each of pixels associated with each of the dynamic objects independently,”) inpainting a region where the at least one pre-existing object was removed; ([0057], “In an exemplary embodiment, the neural network 108 may complete image inpainting to electronically remove and replace each of pixels associated with each of the dynamic objects independently,”) estimating a position of another pre-existing object from the input image; ( PNG media_image2.png 460 472 media_image2.png Greyscale ) However, Li does not explicitly, but Rong teaches: generating a layout overlaid on the input image based on the estimated position; ([0178] , ”FIG. 9 depicts one possible implementation of the system's insertion location determination and object data selection steps. In this implementation, the insertion location determination step includes sampling locations 902 in the environment to determine where the insertion location 906 is going to be. The system can sample the locations and determine whether the locations are viable locations for an object to be placed. The location needs to meet the dynamics of the environment without leading to a collision. In this implementation, the system is aware of the movement of objects in the scene 902, and once a sampling location 906 is determined, the system determines whether the placement leads to a collision 904.” Layout is shown in FIG. 9 first graph.) wherein the layout indicates potential positions for a synthetic object to be added;([0178]-[0179], “FIG. 9 depicts one possible implementation of the system's insertion location determination and object data selection steps. In this implementation, the insertion location determination step includes sampling locations 902 in the environment to determine where the insertion location 906 is going to be. The system can sample the locations and determine whether the locations are viable locations for an object to be placed. The location needs to meet the dynamics of the environment without leading to a collision. In this implementation, the system is aware of the movement of objects in the scene 902, and once a sampling location 906 is determined, the system determines whether the placement leads to a collision 904. When an insertion location is finally determined to be a viable location for placement that does not lead to a collision, an object data set can be selected. The system may take data sets from an object bank 908 to process for selection.”) generating a synthetic object based on the layout. ([0181], “FIG. 10 depicts an example input and output of one implementation of the system. In this implementation, the input is an input video 1002, captured while a car is driving down the street. The output is an output simulated video 1006 that includes a new car 1008 in the input video 1002. In this embodiment, the output is photorealistic 1012, physically plausible 1014, and geometrically consistent 1016. The photorealistic, physically plausible, and geometrically consistent output may have been generated through the use of the method of FIG. 3 or another method or system disclosed herein.”) wherein the generating the synthetic object comprises generating, using a second diffusion model, the synthetic object in the generated input image based on the layout.([0183], “In this implementation, the augmented image 1110 is then processed by an image synthesis model 1114, or a refinement model, to generate a refined augmented image 1116 with corrected texture and lighting. In this implementation, the image synthesis 1114 included texture synthesis for the border between the inserted object and the environment in order to create a smooth transition and a more realistic look.” FIG. 2 step 212 points out the refine step is using a machine learning model based on data generated in the previous steps including insertion location of objects. [0127] teaches that the machine learning model can be any forms of machines learning models: “In addition or alternatively to the model(s) 1235 at the computing system 1400, the machine learning computing system 1200 can include one or more machine-learned models 1235. As examples, the machine-learned models 1235 can be or can otherwise include various machine-learned models such as, for example, neural networks (e.g., deep neural networks), decision trees, ensemble models, k-nearest neighbors models, Bayesian networks, or other types of models including linear models and/or non-linear models. Example neural networks include feed-forward neural networks, recurrent neural networks (e.g., long short-term memory recurrent neural networks), convolutional neural networks, or other forms of neural networks.” [0127] teaches that the machine learning model can be any forms of machines learning models. Although it does not list all the possible machine learning models, e.g. a diffusion model, it would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to have chosen a diffusion model from a finite number of existing machine learning models to perform image generation with a reasonable expectation of success.) Li teaches that a self-driving car system captures an environment, predicts driving actions, but does not teach adding a synthetic vehicle to the environment. Rong also teaches that a self-driving car system captures an environment and analyzes the environment and tests driving safety features by augmented to add another vehicle into the environment. It would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to have combined the teachings of Li with the specific teachings of Rong to effectively test the driving safety features for a self-driving system. However, Li in view of Rong does not, but Xu teaches: The machine learning model can be a first diffusion model ([0083], “It can be seen from FIG. 5d that the above-described disparity-driven non-linear diffusion model is particularly appropriate for both object segmentation and pattern recognition image processing. ”) Li in view of Rong teaches using a machine learning model to process input image and perform object segmentation for this image. Xu teaches a diffusion model can be used to perform object segmentation. It would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to have combined the teachings of Li in view of Rong with the specific teachings of Xu to use an appropriate machine learning model to perform the object segmentation for an image. (Xu [0083]) Regarding claim 2, Li in view of Rong teaches: The method according to claim 1, wherein the at least one pre-existing object is a foreground object, (Li, FIG. 4A, the object are cars.) wherein the another pre-existing object is a road structure object;(Li, FIG. 4C. ) and wherein the generated synthetic object is a vehicle. Regarding claim 3, Li in view of Rong teaches: The method according to claim 2, wherein estimating a position of the another pre-existing object further comprises: estimating a depth of the road structure object.(Li, [0064], “he necessity of defining spatial relation arises from that the interactions of two distant objects are usually scarce. To calculate this relation, the neural network 108 may unproject objects from the 2D image plane to the 3D space in the world frame:[x y z 1].sup.T=δ.sub.u, .sub.vP.sup.−1[u v 1].sup.T   (3) where [u v 1].sup.T and [x y z 1].sup.T are homogeneous representations in 2D and 3D coordinate systems, P is the camera intrinsic matrix, and δ.sub.u,v is the relative depth at (u, v) obtained by depth estimation.”) Regarding claim 4, Li in view of Rong teaches: The method according to claim 3, wherein the another pre-existing object is a road lane, (Li, FIG. 4C, 404) and wherein estimating the position of the another pre-existing object further comprises: estimating a location of the road lane, wherein the location of the road is used as horizontal bounds for the layout. (Li, [0064]-[0065] teaches estimating location and distance between different object. FIG. 4C shows one object on the road is road lane. FIG. 4C, 404 road markers indicate the road lane, which limits the horizontal bounds where a car can be placed. PNG media_image2.png 460 472 media_image2.png Greyscale ) Regarding claim 5, Li in view of Rong teaches: The method according to claim 4, wherein removing the at least one pre-existing object is performed by instance segmentation. (Li [0054], “The I3D 204 is configured to apply instance segmentation and semantic segmentation 210 to detect dynamic objects located within the driving scene and the driving scene characteristics of the driving scene.” After the objects are detected, the following paragraphs [0057] will remove the detected objects.) Regarding claim 11, Li in view of Rong teaches: An apparatus for generating a synthetic image, the apparatus comprising: at least one memory storing computer-executable instructions; and at least one processor configured to execute the computer-executable instructions to: (Li, [0005], “According to yet another aspect, non-transitory computer readable storage medium storing instructions that when executed by a computer, which includes a processor perform a method”) The rest of claim 11 recites similar limitations of claim 1, thus is rejected accordingly. claim 12-15 recites similar limitations of claim 2-5 respectively, thus are rejected accordingly. Conclusion Relevant reference: Voroninski et al. (US 12205197 B1) para 71: teaches using a diffusion model to generate a synthetic image. 7 Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to YANNA WU whose telephone number is (571)270-0725. The examiner can normally be reached Monday-Thursday 8:00-5:30 ET. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Alicia Harrington can be reached at 5712722330. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /YANNA WU/Primary Examiner, Art Unit 2615
Read full office action

Prosecution Timeline

Show 3 earlier events
Feb 02, 2026
Final Rejection mailed — §103
Apr 02, 2026
Response after Non-Final Action
Apr 30, 2026
Request for Continued Examination
May 04, 2026
Response after Non-Final Action
May 08, 2026
Non-Final Rejection mailed — §103
Aug 03, 2026
Response Filed
Aug 17, 2026
Final Rejection mailed — §103
Sep 16, 2026
Response after Non-Final Action

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12743561
Generating Technical Drawings From Building Information Models
2y 0m to grant Granted Sep 22, 2026
Patent 12724582
SYSTEMS, APPARATUSES, AND METHODS FOR REAL-TIME COLLABORATION WITH A GRAPHICAL RENDERING PROGRAM
3y 0m to grant Granted Sep 01, 2026
Patent 12725352
INFORMATION PROCESSING DEVICE AND INFORMATION PROCESSING METHOD
2y 1m to grant Granted Sep 01, 2026
Patent 12718475
SITE MODEL UPDATING METHOD AND SYSTEM
3y 2m to grant Granted Aug 25, 2026
Patent 12711690
SPECULATIVE EXECUTION OF HIT AND INTERSECTION SHADERS ON PROGRAMMABLE RAY TRACING ARCHITECTURES
1y 10m to grant Granted Aug 18, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

4-5
Expected OA Rounds
81%
Grant Probability
99%
With Interview (+34.4%)
2y 2m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 456 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month