Prosecution Insights
Last updated: October 02, 2026
Application No. 18/518,430

NEURAL NETWORKS TO GENERATE OBJECTS WITHIN DIFFERENT IMAGES

Non-Final OA §103
Filed
Nov 22, 2023
Examiner
WANG, YUEHAN
Art Unit
2617
Tech Center
2600 — Communications
Assignee
NVIDIA Corporation
OA Round
3 (Non-Final)
82%
Grant Probability
Favorable
3-4
OA Rounds
0m
Est. Remaining
96%
With Interview

Examiner Intelligence

Grants 82% — above average
82%
Career Allowance Rate
416 granted / 504 resolved
+20.5% vs TC avg
Moderate +14% lift
Without
With
+13.6%
Interview Lift
resolved cases with interview
Typical timeline
2y 5m
Avg Prosecution
26 currently pending
Career history
544
Total Applications
across all art units

Statute-Specific Performance

§101
4.9%
-35.1% vs TC avg
§103
72.4%
+32.4% vs TC avg
§102
6.9%
-33.1% vs TC avg
§112
6.2%
-33.8% vs TC avg
Black line = Tech Center average estimate • Based on career data from 504 resolved cases

Office Action

§103
DETAILED ACTION Response to Amendment Applicant’s amendments filed on 06 March 2026 have been entered. Claims 1-2, 5, 7-9, and 14-16 have been amended. Claims 1-20 are still pending in this application, with claims 1, 8 and 15 being independent. Continued Examination Under 37 CFR 1.114 A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 06 March 2026 has been entered. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1, 2, 4-9, 11-16 and 18-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Couleaud et al. (US 20250078347 A1), referred herein as Couleaud in view of Park et al. (US 12524937 B2), referred herein as Park. Regarding Claim 1, Couleaud in view of Park teaches one or more processors, comprising: circuitry to use one or more neural networks to: (Couleaud Abst: Systems and methods are described for inputting text input to a trained machine learning model; [0002] Text-to-image models are types of neural networks that generate images based on a text prompt; FIG. 9:911: control circuitry): receive one or more prompts that indicate an object and content related to the object (Couleaud [0039] As a non-limiting illustrative example, as shown in FIG. 2, text prompt 102 may correspond to text prompt 202 of FIG. 2 of “a picture of a young woman holding a cat in front of a Victorian era building, 1870, high quality, soft focus, f/18, 60 mm, in the style of Auguste Renoir.” In some embodiments, text prompt 102 may be provided as an input to a first trained machine learning model (e.g., text-to-image model 108 of FIG. 1)); and in response to receiving the one or more prompts, generate at least a first image and a second image depicting the object (Couleaud [0040] Based on input prompt 102, model 108 may be configured to output image 110 (e.g., representing an interpretation of text input 202, as determined by model 108), which may correspond to first image 210 of FIG. 2; [0042] As shown in FIG. 1, the image processing system may perform segmentation 114 of image 110 to obtain images 116, 118, and 120 (e.g., different identified portions, segments, and/or objects of image 210 of FIG. 2), which may respectively correspond to images 216, 218, and 220. Images 216, 218, and 220 may each comprise a depiction of a respective object of the plurality of objects 211, 213 and 215 of image 210), wherein the first image is to depict the object within a first background corresponding to the content and the second image is to depict the object subject within a second background corresponding to the content Couleaud [0056] textual prompts 128, 130, and 132 (and/or textual prompts 127, 129 and 131) may be input to trained machine learning model 108, and based on such input, trained machine learning model 108 may be configured to generate and output a second plurality of images 134, 136, and 138 (which may correspond to images 234, 236, and 238, respectively, of FIG. 2)). Couleaud does not disclose the second background is different from the first background. However, Park disclosed systems and methods for text-based image generation, which is an analogous art. Park teaches and different from the first background (Park col 8, ln 38-46: generates an image based on the text prompt, and displays the image to the user via user interface 300 (e.g., in first image display 320 and second image display 325). As shown in FIG. 3, the image generation apparatus generates two sets of images at two different resolutions in response to the text prompt, where the sets of images jointly depict similar looking vases of flowers on a white tabletop, but vary in angle of depiction, lighting, background detail, etc). It would have been obvious for a person of ordinary skill in the art before the effective filing date of the claimed invention to have modified Couleaud to incorporate the teachings of Park, and applying the two sets of images in response to the text prompt with varying in angle of depiction, lighting, background details into systems and methods for generating multi-layer images. Doing so would allow a user to produce an image without having to use an original image as an input, and therefore makes image generation easier for a layperson and also more readily automated. Regarding Claim 2, Couleaud in view of Park teaches the one or more processors of claim 1, and further teaches wherein the object within the first image and the second image include a same feature (Couleaud [0061] images 234, 236 and 238 may correspond to respective layers of a multi-layer image, and may be composited together to form a multi-layer image, e.g., composite image 240, 242, 244, or 246, which may comprise suitable transparency for such a multi-layer image. As shown in FIG. 2, the image processing system may output to a user a plurality of different alternative for the composite image. For example, composite image 244 may be the most similar to image 210 (e.g., in terms of the woman's center and prominent position with the cat)). Regarding Claim 4, Couleaud in view of Park teaches the one or more processors of claim 1, and further teaches wherein the one or more neural networks include a diffusion neural network (Park col 3. Ln 17-21: an image-generating machine learning model is a diffusion model, which generates an output by removing noise from an input. Diffusion models can produce higher-quality output images than GANs, but at the cost of a reduced processing speed). Regarding Claim 5, Couleaud in view of Park teaches the one or more processors of claim 1, and further teaches wherein the one or more prompts include one or more input text prompts, wherein the one or more input text prompts are generated by one or more users (Couleaud [0002] Text-to-image models are types of neural networks that generate images based on a text prompt, such as a sentence or a paragraph describing the desired image to be generated. Such models have recently been incorporated in new, popular tools for users to generate images based on a text input). Regarding Claim 6, Couleaud in view of Park teaches the one or more processors of claim 1, and further teaches wherein the one or more neural networks receive only text prompts as inputs (Couleaud [0005] large language models can be used to extract object and layer information from a prompt, for use as a guide in a text-to-image generation; [0072] Zero-Shot Text-to-Image Generation). Regarding Claim 7, Couleaud in view of Park teaches the one or more processors of claim 1, and further teaches wherein the object is a same object subject and in different poses within the first image and the second image (Couleaud [0061] composite image 244 may be the most similar to image 210 (e.g., in terms of the woman's center and prominent position with the cat) and other variations (e.g., in a position of, or presence of, the woman, cat and building and/or other objects) may be provided via the other composite images 240, 242, 244, or 246. As another example, composite image 240, 242, 244, or 246 may be generated based on receiving user input, e.g., to edit or move around objects and/or layers corresponding to images 234, 236, and 238). Regarding Claim 8, 9 and 11-14, Couleaud in view of Park teaches a system comprising: one or more processors to use one or more neural networks to (Couleaud Abst: Systems and methods are described for inputting text input to a trained machine learning model; [0002] Text-to-image models are types of neural networks that generate images based on a text prompt; FIG. 9:911: control circuitry): The metes and bounds of the rest of the limitations of the claims substantially correspond to the limitations set forth in claims 1, 2 and 4-7; thus they are rejected on similar grounds and rationale as their corresponding limitations. Regarding Claim 15, 16 and 18-20, Couleaud in view of Park teaches a method comprising (Couleaud Abst: Systems and methods are described for inputting text input to a trained machine learning model; [0002] Text-to-image models are types of neural networks that generate images based on a text prompt; FIG. 9:911: control circuitry): The metes and bounds of the rest of the limitations of the claims substantially correspond to the limitations set forth in claims 1, 2 and 4-6; thus they are rejected on similar grounds and rationale as their corresponding limitations. Claim(s) 3, 10 and 17 is/are rejected under 35 U.S.C. 103 as being unpatentable over Couleaud et al. (US 20250078347 A1), referred herein as Couleaud in view of Park et al. (US 12524937 B2), referred herein as Park and Sharma (US 20250095230 A1), referred herein as Sharma. Regarding Claim 3, Couleaud in view of Park teaches the one or more processors of claim 1, but does not teach the claimed limitations herein. However, Sharma disclosed systems and methods for generating digital images utilizing a diffusion neural network to preserve color harmony and image composition from a sample digital image while modifying image content, which is an analogous art. Sharma teaches wherein the one or more neural networks include one or more layers that denoise images to generate the first image and the second image (Sharma [0038] the image content modification system 102 utilizes the architecture of the diffusion neural network 206 (e.g., a plurality of denoising layers that remove noise or recreate a digital image) to generate a digital image (e.g., the generated digital image 208) from the noise map/inversion; [0041] the image context modification system 102 generates digital images using a diffusion neural network to not only preserve color harmony and image composition of a sample digital image, but also to reflect content indicated by a text prompt). It would have been obvious for a person of ordinary skill in the art before the effective filing date of the claimed invention to have modified Couleaud to incorporate the teachings of Sharma, and applying the denoising layers into systems and methods for generating multi-layer images. Doing so would be able to generate more accurate search results that include digital images which not only reflect content indicated by a search query but that also portray color harmonies and image compositions of a generated digital image. Regarding Claim 10, Couleaud in view of Park teaches the system of claim 8. The metes and bounds of the claim substantially correspond to the limitations set forth in claim 3; thus they are rejected on similar grounds and rationale as their corresponding limitations. Regarding Claim 17, Couleaud in view of Park teaches the method of claim 15. The metes and bounds of the claim substantially correspond to the limitations set forth in claim 3; thus they are rejected on similar grounds and rationale as their corresponding limitations. Response to Arguments Applicant's arguments filed on 06 March 2026, with respect to the 103 rejection have been fully considered but are moot in view of the new grounds of rejection. Examiner notes that independent claims 1, 8 and 15 have been amended to include new limitation. Examiner finds these limitations to be unpatentable as can be found in above detail action. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to Samantha (Yuehan) Wang whose telephone number is (571)270-5011. The examiner can normally be reached Monday-Friday, 8am-5pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, King Poon can be reached on (571)272-7440. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /Samantha (YUEHAN) WANG/ Primary Examiner Art Unit 2617
Read full office action

Prosecution Timeline

Show 9 earlier events
Jan 29, 2026
Examiner Interview Summary
Mar 06, 2026
Response after Non-Final Action
Mar 12, 2026
Request for Continued Examination
Mar 15, 2026
Response after Non-Final Action
Jul 14, 2026
Non-Final Rejection mailed — §103
Sep 14, 2026
Interview Requested
Sep 16, 2026
Examiner Interview Summary
Sep 16, 2026
Examiner Interview (Telephonic)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12731302
RAPID RENDERING AND/OR REALISTIC VISUALIZATION OF APPAREL DESIGN DRAFT FILES THROUGH APPLICATION OF ONE OR MORE GENERATIVE ARTIFICIAL NEURAL NETWORKS
2y 6m to grant Granted Sep 08, 2026
Patent 12718478
VIRTUAL ENVIRONMENT GUIDING METHOD AND SYSTEM
2y 3m to grant Granted Aug 25, 2026
Patent 12700146
SYSTEM FOR CONTEXTUAL DIMINISHED REALITY FOR METAVERSE IMMERSIONS
2y 3m to grant Granted Aug 04, 2026
Patent 12700192
AUGMENTED REALITY SYSTEM
2y 3m to grant Granted Aug 04, 2026
Patent 12694620
DEPTH RENDERING FROM NEURAL RADIANCE FIELDS FOR 3D MODELING
2y 2m to grant Granted Jul 28, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
82%
Grant Probability
96%
With Interview (+13.6%)
2y 5m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 504 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month