Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
DETAILED ACTION
This is in response to applicant’s amendment/response filed on 08/03/2026, which has
been entered and made of record. Claim 1, 3-5, 14-15 are amended. Claims 1-5, 11-15 are pending in the application. Applicant is reminded to change the status of the withdrawn claims 6-10 and 16-20 to be cancelled.
Response to Arguments
Applicant arguments regarding claim rejections under 103 are considered, but are not persuasive.
Applicant argues:
PNG
media_image1.png
148
654
media_image1.png
Greyscale
Examiner disagrees: Li teaches using a machine model to generate the synthetic object image as shown in FIG. 2 and also [0183]. Li in [0127] also teaches that the machine learning model can be any forms of machines learning models. Although it does not list all the possible machine learning models, e.g. a diffusion model, it would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to have chosen a diffusion model from a finite number of existing machine learning models to perform image generation with a reasonable expectation of success.
Applicant argues: FIG. 4C 404 of Li does not discloses a location of road lane.
Examiner disagrees: Li explicitly say FIG. 4C, 404 indicates the road markers in an image, which means the two line indicates the location of a road in the image. Furthermore, the Specification FIG. 2 also uses to lines to indicate the location of road in an image, which is consistent with the teachings of Li.
The rest of arguments are moot in view of new ground of rejections.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-5, 11-15 is/are rejected under 35 U.S.C. 103 as being unpatentable over Li et al. (US 2021/0261148 A1) in view of Rong et al. (US 2021/0383616 A1) and further in view of Xu et al. (US 2002/0175921 A1).
Regarding claim 1, Li teaches:
A method for generating a synthetic image, the method comprising:
receiving an input image, wherein the receiving the input image comprises generating the input image based on a machine learning model; ([0053], “In an exemplary embodiment, the data reception module 118 of the action prediction application 106 may be configured to receive image data that may be associated with images captured of the surrounding environment of the ego vehicle 102 that may be provided by the vehicle camera system 110 of the ego vehicle 102.” [0054], “In one embodiment, the data reception module 118 may be configured to input the image frames 202 to the neural network 108 to be analyzed by I3D 204. The I3D 204 is configured to apply instance segmentation and semantic segmentation 210 to detect dynamic objects located within the driving scene and the driving scene characteristics of the driving scene.”)
removing at least one pre-existing object from the input image; ([0057], “In an exemplary embodiment, the neural network 108 may complete image inpainting to electronically remove and replace each of pixels associated with each of the dynamic objects independently,”)
inpainting a region where the at least one pre-existing object was removed; ([0057], “In an exemplary embodiment, the neural network 108 may complete image inpainting to electronically remove and replace each of pixels associated with each of the dynamic objects independently,”)
estimating a position of another pre-existing object from the input image; (
PNG
media_image2.png
460
472
media_image2.png
Greyscale
)
However, Li does not explicitly, but Rong teaches:
generating a layout overlaid on the input image based on the estimated position; ([0178] , ”FIG. 9 depicts one possible implementation of the system's insertion location determination and object data selection steps. In this implementation, the insertion location determination step includes sampling locations 902 in the environment to determine where the insertion location 906 is going to be. The system can sample the locations and determine whether the locations are viable locations for an object to be placed. The location needs to meet the dynamics of the environment without leading to a collision. In this implementation, the system is aware of the movement of objects in the scene 902, and once a sampling location 906 is determined, the system determines whether the placement leads to a collision 904.” Layout is shown in FIG. 9 first graph.)
wherein the layout indicates potential positions for a synthetic object to be added;([0178]-[0179], “FIG. 9 depicts one possible implementation of the system's insertion location determination and object data selection steps. In this implementation, the insertion location determination step includes sampling locations 902 in the environment to determine where the insertion location 906 is going to be. The system can sample the locations and determine whether the locations are viable locations for an object to be placed. The location needs to meet the dynamics of the environment without leading to a collision. In this implementation, the system is aware of the movement of objects in the scene 902, and once a sampling location 906 is determined, the system determines whether the placement leads to a collision 904. When an insertion location is finally determined to be a viable location for placement that does not lead to a collision, an object data set can be selected. The system may take data sets from an object bank 908 to process for selection.”)
generating a synthetic object based on the layout. ([0181], “FIG. 10 depicts an example input and output of one implementation of the system. In this implementation, the input is an input video 1002, captured while a car is driving down the street. The output is an output simulated video 1006 that includes a new car 1008 in the input video 1002. In this embodiment, the output is photorealistic 1012, physically plausible 1014, and geometrically consistent 1016. The photorealistic, physically plausible, and geometrically consistent output may have been generated through the use of the method of FIG. 3 or another method or system disclosed herein.”) wherein the generating the synthetic object comprises generating, using a second diffusion model, the synthetic object in the generated input image based on the layout.([0183], “In this implementation, the augmented image 1110 is then processed by an image synthesis model 1114, or a refinement model, to generate a refined augmented image 1116 with corrected texture and lighting. In this implementation, the image synthesis 1114 included texture synthesis for the border between the inserted object and the environment in order to create a smooth transition and a more realistic look.” FIG. 2 step 212 points out the refine step is using a machine learning model based on data generated in the previous steps including insertion location of objects. [0127] teaches that the machine learning model can be any forms of machines learning models: “In addition or alternatively to the model(s) 1235 at the computing system 1400, the machine learning computing system 1200 can include one or more machine-learned models 1235. As examples, the machine-learned models 1235 can be or can otherwise include various machine-learned models such as, for example, neural networks (e.g., deep neural networks), decision trees, ensemble models, k-nearest neighbors models, Bayesian networks, or other types of models including linear models and/or non-linear models. Example neural networks include feed-forward neural networks, recurrent neural networks (e.g., long short-term memory recurrent neural networks), convolutional neural networks, or other forms of neural networks.” [0127] teaches that the machine learning model can be any forms of machines learning models. Although it does not list all the possible machine learning models, e.g. a diffusion model, it would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to have chosen a diffusion model from a finite number of existing machine learning models to perform image generation with a reasonable expectation of success.)
Li teaches that a self-driving car system captures an environment, predicts driving actions, but does not teach adding a synthetic vehicle to the environment. Rong also teaches that a self-driving car system captures an environment and analyzes the environment and tests driving safety features by augmented to add another vehicle into the environment.
It would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to have combined the teachings of Li with the specific teachings of Rong to effectively test the driving safety features for a self-driving system.
However, Li in view of Rong does not, but Xu teaches:
The machine learning model can be a first diffusion model ([0083], “It can be seen from FIG. 5d that the above-described disparity-driven non-linear diffusion model is particularly appropriate for both object segmentation and pattern recognition image processing. ”)
Li in view of Rong teaches using a machine learning model to process input image and perform object segmentation for this image. Xu teaches a diffusion model can be used to perform object segmentation.
It would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to have combined the teachings of Li in view of Rong with the specific teachings of Xu to use an appropriate machine learning model to perform the object segmentation for an image. (Xu [0083])
Regarding claim 2, Li in view of Rong teaches:
The method according to claim 1, wherein the at least one pre-existing object is a foreground object, (Li, FIG. 4A, the object are cars.) wherein the another pre-existing object is a road structure object;(Li, FIG. 4C. ) and wherein the generated synthetic object is a vehicle.
Regarding claim 3, Li in view of Rong teaches:
The method according to claim 2, wherein estimating a position of the another pre-existing object further comprises: estimating a depth of the road structure object.(Li, [0064], “he necessity of defining spatial relation arises from that the interactions of two distant objects are usually scarce. To calculate this relation, the neural network 108 may unproject objects from the 2D image plane to the 3D space in the world frame:[x y z 1].sup.T=δ.sub.u, .sub.vP.sup.−1[u v 1].sup.T (3)
where [u v 1].sup.T and [x y z 1].sup.T are homogeneous representations in 2D and 3D coordinate systems, P is the camera intrinsic matrix, and δ.sub.u,v is the relative depth at (u, v) obtained by depth estimation.”)
Regarding claim 4, Li in view of Rong teaches:
The method according to claim 3, wherein the another pre-existing object is a road lane, (Li, FIG. 4C, 404) and wherein estimating the position of the another pre-existing object further comprises: estimating a location of the road lane, wherein the location of the road is used as horizontal bounds for the layout. (Li, [0064]-[0065] teaches estimating location and distance between different object. FIG. 4C shows one object on the road is road lane. FIG. 4C, 404 road markers indicate the road lane, which limits the horizontal bounds where a car can be placed.
PNG
media_image2.png
460
472
media_image2.png
Greyscale
)
Regarding claim 5, Li in view of Rong teaches:
The method according to claim 4, wherein removing the at least one pre-existing object is performed by instance segmentation. (Li [0054], “The I3D 204 is configured to apply instance segmentation and semantic segmentation 210 to detect dynamic objects located within the driving scene and the driving scene characteristics of the driving scene.” After the objects are detected, the following paragraphs [0057] will remove the detected objects.)
Regarding claim 11, Li in view of Rong teaches:
An apparatus for generating a synthetic image, the apparatus comprising: at least one memory storing computer-executable instructions; and at least one processor configured to execute the computer-executable instructions to: (Li, [0005], “According to yet another aspect, non-transitory computer readable storage medium storing instructions that when executed by a computer, which includes a processor perform a method”) The rest of claim 11 recites similar limitations of claim 1, thus is rejected accordingly.
claim 12-15 recites similar limitations of claim 2-5 respectively, thus are rejected accordingly.
Conclusion
Relevant reference: Voroninski et al. (US 12205197 B1) para 71: teaches using a diffusion model to generate a synthetic image. 7
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to YANNA WU whose telephone number is (571)270-0725. The examiner can normally be reached Monday-Thursday 8:00-5:30 ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Alicia Harrington can be reached at 5712722330. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/YANNA WU/Primary Examiner, Art Unit 2615