Prosecution Insights
Last updated: October 01, 2026
Application No. 18/952,023

UTILIZING A GENERATIVE NEURAL NETWORK TO INTERACTIVELY CREATE AND MODIFY DIGITAL IMAGES BASED ON NATURAL LANGUAGE FEEDBACK

Non-Final OA §102§103
Filed
Nov 19, 2024
Priority
Jan 14, 2022 — continuation of 12/148,119
Examiner
AUGUSTIN, MARCELLUS
Art Unit
Tech Center
Assignee
Adobe Inc.
OA Round
1 (Non-Final)
82%
Grant Probability
Favorable
1-2
OA Rounds
8m
Est. Remaining
98%
With Interview

Examiner Intelligence

Grants 82% — above average
82%
Career Allowance Rate
711 granted / 869 resolved
+21.8% vs TC avg
Strong +16% interview lift
Without
With
+16.0%
Interview Lift
resolved cases with interview
Typical timeline
2y 7m
Avg Prosecution
23 currently pending
Career history
885
Total Applications
across all art units

Statute-Specific Performance

§101
10.6%
-29.4% vs TC avg
§103
50.9%
+10.9% vs TC avg
§102
20.8%
-19.2% vs TC avg
§112
12.1%
-27.9% vs TC avg
Black line = Tech Center average estimate • Based on career data from 869 resolved cases

Office Action

§102 §103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to filed Amendments/Remarks and Continuation Request The request for a continuation is acknowledged. For a continuation to be applied as per the MPEP, the application discloses and claims only subject matter disclosed in prior Applications, and names the inventor or at least one joint inventor named in the prior application. Accordingly, this application may constitute a continuation or divisional. Should applicant desire to claim the benefit of the filing date of the prior application, attention is directed to 35 U.S.C. 120, 37 CFR 1.78, and MPEP § 211 et seq. Applicant’s filed IDS of 11/27/2024 have been entered and considered. New claims 2-20 have been added. Claims 1-20 are currently pending. Please refer to the action below. Examiner Notes The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. However, the claimed subject matter, not the specification, is the measure of the invention. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-3, 5-6, and 16-18 is/are rejected under 35 U.S.C. 103 as obvious over Liu et al. (US 2022/0108417, A1) in view Garcia et al. (EP 3754548, A1). Regarding claim 1, Liu teaches a method (at least Figs. 1-3 teaches a method of receiving a natural language prompt indicating one or more image modifications for a digital image and configured further in Figs. 1B-3 further adapted generating, utilizing a generative neural network of further para. 0066-0067, a modified digital image with the one or more image modifications) comprising: receiving a natural language prompt indicating one or more image modifications for a digital image (receiving speech and/or text input 152 of Figs. 1-3 comprising said natural language prompt illustrating “add a boat” indicating one or more image modifications for a digital image 114); generating, utilizing a text encoder, a textual feature vector from the natural language prompt (utilizing a text encoder of at least para. 0065-0067 corresponding to user inputs of at least para. 0051 for generating encoded textual feature vector from the natural language prompt); and generating, utilizing a generative neural network, a modified digital image with the one or more image modifications in at least para. 0065-0067 for generating a modified digital image 218 of Figs. 1B-3 with the one or more image modifications). Liu is silent regarding the above lined-out items such as specifically teaching said generating modified digital image with the one or more image modifications by conditioning the generative neural network on the textual feature vector. Garcia teaches at least in Figure 5 and the disclosure a text encoder including at least a text encoder configured to encode textual feature vector from the natural language prompt and to feed said encoded text feature to a conditional generative model 605 (such as the generative model 131 and 522) trained and conditioned to combine the encoded text features and noise data such as “textual samples may be combined with noise and may be used by the generative model to generate the images. The training of the generator and discriminator of the GAN may then be performed in an adversarial way as with training the GAN, but the generator is conditioned with the textual samples. This may enable a better convergence of the training” to output a digital image by conditioning the generative neural network on the textual feature vector corresponding to the user instructed input commands, said output image as understood in the art may obviously comprise one of a modified image based on understood user inputs depicting in a case a requested modification input to an already generated image. It would have been obvious to one of ordinary skill in the art at the time the invention was made to combine the teachings of Liu in view of Garcia to include wherein said generating modified digital image with the one or more image modifications by conditioning the generative neural network on the textual feature vector, as discussed above, as Liu in view of Garcia are in the same field of endeavor including methods and systems for utilizing a conditioning generative neural network for generating output images corresponding to user inputted text/speech commands, Garcia’s combination architecture of conditioning the generative neural network on the textual feature vector comprises combining noise with the textual feature vector further complements the method of Liu of utilizing said generative neural network for generating the modified digital image with the one or more image modifications, in a sense that said method of Liu when combined to the combination architecture of Garcia, it enables the systems and methods of Liu to become advantageously enabled to output the said generated modified digital image with the one or more image modifications based on the cited model conditioning methods of Garcia by combining noise and the textual feature vector in such a way to provide an optimized and accurate output image which may obviously be further modified according to another or subsequent user input modification instructions to generate the claimed modified image of Liu based on the said model conditioning which may be further realized according to further known methods to yield predictable results since known work in one field of endeavor may prompt variations of it for use in either the same field or a different one based on design incentives or other market forces if the variations are predictable to one of ordinary skill in the art as said combination is thus the adaptation of an old idea or invention using newer technology that is either commonly available and understood in the art thereby a variation on already known art (See MPEP 2143, KSR Exemplary Rationale F). Regarding claim 2 (according to claim 1), Liu is silent regarding wherein conditioning the generative neural network on the textual feature vector comprises combining noise with the textual feature vector. Garcia further teaches at least in Figure 5 and the disclosure a conditional generative model 605 conditioned to combine and synthesized encoded text features and noise data such as cited “textual samples may be combined with noise and may be used by the generative model to generate the images. The training of the generator and discriminator of the GAN may then be performed in an adversarial way as with training the GAN, but the generator is conditioned with the textual samples. This may enable a better convergence of the training” to output a generated generating output digital image, utilizing the generative neural network, from a combination of the noise and the textual feature vector. It would have been obvious to one of ordinary skill in the art at the time the invention was made to combine the teachings of Liu in view of Garcia to include wherein said conditioning the generative neural network on the textual feature vector comprises combining noise with the textual feature vector, as discussed above, as Liu in view of Garcia are in the same field of endeavor including methods and systems for utilizing a conditioning generative neural network for generating output images corresponding to user inputted text/speech commands, Garcia’s combination architecture of conditioning the generative neural network on the textual feature vector comprises combining noise with the textual feature vector further complements the method of Liu of utilizing said generative neural network for generating the modified digital image with the one or more image modifications, in a sense that said method of Liu when combined to the combination architecture of Garcia, it enables the systems and methods of Liu to become advantageously enabled to output the said generated modified digital image with the one or more image modifications based on the cited model conditioning methods of Garcia by combining noise and the textual feature vector in such a way to provide an optimized and accurate output image which may obviously be further modified according to another or subsequent user input modification instructions to generate the claimed modified image of Liu based on the said model conditioning which may be further realized according to further known methods to yield predictable results since known work in one field of endeavor may prompt variations of it for use in either the same field or a different one based on design incentives or other market forces if the variations are predictable to one of ordinary skill in the art as said combination is thus the adaptation of an old idea or invention using newer technology that is either commonly available and understood in the art thereby a variation on already known art (See MPEP 2143, KSR Exemplary Rationale F). Regarding claim 3 (according to claim 2), Liu is silent regarding wherein generating the modified digital image comprises synthesizing, utilizing the generative neural network, the one or more image modifications from a combination of the noise and the textual feature vector. Garcia further teaches at least in Figure 5 and the disclosure a conditional generative model 605 conditioned to combine and synthesized encoded text features and noise data such as cited “textual samples may be combined with noise and may be used by the generative model to generate the images. The training of the generator and discriminator of the GAN may then be performed in an adversarial way as with training the GAN, but the generator is conditioned with the textual samples. This may enable a better convergence of the training” to output a generated generating output digital image, utilizing the generative neural network, from a combination of the noise and the textual feature vector. It would have been obvious to one of ordinary skill in the art at the time the invention was made to combine the teachings of Liu in view of Garcia to include wherein said generating modified digital image comprises synthesizing, utilizing the generative neural network, the one or more image modifications from a combination of the noise and the textual feature vector, as discussed above, as Liu in view of Garcia are in the same field of endeavor including methods and systems for utilizing a conditioning generative neural network for generating output images corresponding to user inputted text/speech commands, Garcia’s combination architecture of conditioning the generative neural network on the textual feature vector comprises combining noise with the textual feature vector further complements the method of Liu of utilizing said generative neural network for generating the modified digital image with the one or more image modifications, in a sense that said method of Liu when combined to the combination architecture of Garcia, it enables the systems and methods of Liu to become advantageously enabled to output the said generated modified digital image with the one or more image modifications based on the cited model conditioning methods of Garcia by combining noise and the textual feature vector in such a way to provide an optimized and accurate output image which may obviously be further modified according to another or subsequent user input modification instructions to generate the claimed modified image of Liu based on the said model conditioning which may be further realized according to further known methods to yield predictable results since known work in one field of endeavor may prompt variations of it for use in either the same field or a different one based on design incentives or other market forces if the variations are predictable to one of ordinary skill in the art as said combination is thus the adaptation of an old idea or invention using newer technology that is either commonly available and understood in the art thereby a variation on already known art (See MPEP 2143, KSR Exemplary Rationale F). Regarding claim 5 (according to claim 1), Liu further illustrates wherein generating the modified digital image comprises replacing, within a graphical user interface on a client device, the digital image with the modified digital image in response to receiving the natural language prompt (at least Fig. 1B illustrates images 156/218 comprising the generated modified digital image comprises replacing, within a graphical user interface on a client device, the digital image with the modified digital image in response to receiving the natural language prompt 152). Regarding claim 6 (according to claim 1), Liu further illustrates wherein receiving the natural language prompt comprises receiving a textual input (a case in at least Figs. 1-3 where the system is configured for receiving said natural language prompt comprises receiving a textual input). Regarding claim 16, Liu teaches a non-transitory computer-readable medium comprising instructions that, when executed by at least one processor, cause a computing device to perform operations (computing device of at least para. 0122 or that of para. 0312 comprises said computer-readable medium comprising instructions and the at least one processor), comprising: receiving a natural language prompt indicating one or more image modifications for a digital image (para. 0048-0051 and Figs. 1-3 further teaches a case of receiving second modification natural language prompt indicating adding of elements as the one or more image modifications to make to the digital image); and generating, utilizing a text encoder, a textual feature vector from the natural language prompt (the prompt of at least para. 0048-0051 may be processed utilizing a text encoder of at least para. 0065-0067 corresponding to user inputs of at least para. 0051 for generating encoded textual feature vector from the natural language prompt); and generating, utilizing a generative neural network, a modified digital image with the one or more image modifications Liu is silent regarding the above lined-out items such as specifically teaching said generating modified digital image with the one or more image modifications by conditioning the generative neural network on the textual feature vector. Garcia teaches at least in Figure 5 and the disclosure a text encoder including at least a text encoder configured to encode textual feature vector from the natural language prompt and to feed said encoded text feature to a conditional generative model 605 (such as the generative model 131 and 522) trained and conditioned to combine the encoded text features and noise data such as “textual samples may be combined with noise and may be used by the generative model to generate the images. The training of the generator and discriminator of the GAN may then be performed in an adversarial way as with training the GAN, but the generator is conditioned with the textual samples. This may enable a better convergence of the training” to output a digital image by conditioning the generative neural network on the textual feature vector corresponding to the user instructed input commands, said output image as understood in the art may obviously comprise one of a modified image based on understood user inputs depicting in a case a requested modification input to an already generated image. It would have been obvious to one of ordinary skill in the art at the time the invention was made to combine the teachings of Liu in view of Garcia to include wherein said generating modified digital image with the one or more image modifications by conditioning the generative neural network on the textual feature vector, as discussed above, as Liu in view of Garcia are in the same field of endeavor including methods and systems for utilizing a conditioning generative neural network for generating output images corresponding to user inputted text/speech commands, Garcia’s combination architecture of conditioning the generative neural network on the textual feature vector comprises combining noise with the textual feature vector further complements the method of Liu of utilizing said generative neural network for generating the modified digital image with the one or more image modifications, in a sense that said method of Liu when combined to the combination architecture of Garcia, it enables the systems and methods of Liu to become advantageously enabled to output the said generated modified digital image with the one or more image modifications based on the cited model conditioning methods of Garcia by combining noise and the textual feature vector in such a way to provide an optimized and accurate output image which may obviously be further modified according to another or subsequent user input modification instructions to generate the claimed modified image of Liu based on the said model conditioning which may be further realized according to further known methods to yield predictable results since known work in one field of endeavor may prompt variations of it for use in either the same field or a different one based on design incentives or other market forces if the variations are predictable to one of ordinary skill in the art as said combination is thus the adaptation of an old idea or invention using newer technology that is either commonly available and understood in the art thereby a variation on already known art (See MPEP 2143, KSR Exemplary Rationale F). Regarding claim 17 (according to claim 16), Liu is silent regarding wherein conditioning the generative neural network on the textual feature vector comprises combining a noise vector with the textual feature vector. Garcia further teaches at least in Figure 5 and the disclosure a conditional generative model 605 conditioned to combine and synthesized encoded text features and noise data such as cited “textual samples may be combined with noise and may be used by the generative model to generate the images. The training of the generator and discriminator of the GAN may then be performed in an adversarial way as with training the GAN, but the generator is conditioned with the textual samples. This may enable a better convergence of the training” to output a generated generating output digital image, utilizing the generative neural network, from a combination of the noise and the textual feature vector. It would have been obvious to one of ordinary skill in the art at the time the invention was made to combine the teachings of Liu in view of Garcia to include wherein said conditioning the generative neural network on the textual feature vector comprises combining noise with the textual feature vector, as discussed above, as Liu in view of Garcia are in the same field of endeavor including methods and systems for utilizing a conditioning generative neural network for generating output images corresponding to user inputted text/speech commands, Garcia’s combination architecture of conditioning the generative neural network on the textual feature vector comprises combining noise with the textual feature vector further complements the method of Liu of utilizing said generative neural network for generating the modified digital image with the one or more image modifications, in a sense that said method of Liu when combined to the combination architecture of Garcia, it enables the systems and methods of Liu to become advantageously enabled to output the said generated modified digital image with the one or more image modifications based on the cited model conditioning methods of Garcia by combining noise and the textual feature vector in such a way to provide an optimized and accurate output image which may obviously be further modified according to another or subsequent user input modification instructions to generate the claimed modified image of Liu based on the said model conditioning which may be further realized according to further known methods to yield predictable results since known work in one field of endeavor may prompt variations of it for use in either the same field or a different one based on design incentives or other market forces if the variations are predictable to one of ordinary skill in the art as said combination is thus the adaptation of an old idea or invention using newer technology that is either commonly available and understood in the art thereby a variation on already known art (See MPEP 2143, KSR Exemplary Rationale F). Regarding claim 18 (according to claim 17), Liu is silent regarding wherein generating the modified digital image comprises synthesizing, utilizing the generative neural network, the one or more image modifications by transforming the noise vector into the one or more image modifications in a manner informed by the textual feature vector. Garcia further teaches at least in Figure 5 and the disclosure a conditional generative model 605 conditioned to combine and synthesized encoded text features and noise data such as cited “textual samples may be combined with noise and may be used by the generative model to generate the images. The training of the generator and discriminator of the GAN may then be performed in an adversarial way as with training the GAN, but the generator is conditioned with the textual samples. This may enable a better convergence of the training” to output a generated generating output digital image, utilizing the generative neural network, from a combination of the noise and the textual feature vector. It would have been obvious to one of ordinary skill in the art at the time the invention was made to combine the teachings of Liu in view of Garcia to include wherein said generating modified digital image comprises synthesizing, utilizing the generative neural network, the one or more image modifications from a combination of the noise and the textual feature vector, as discussed above, as Liu in view of Garcia are in the same field of endeavor including methods and systems for utilizing a conditioning generative neural network for generating output images corresponding to user inputted text/speech commands, Garcia’s combination architecture of conditioning the generative neural network on the textual feature vector comprises combining noise with the textual feature vector further complements the method of Liu of utilizing said generative neural network for generating the modified digital image with the one or more image modifications, in a sense that said method of Liu when combined to the combination architecture of Garcia, it enables the systems and methods of Liu to become advantageously enabled to output the said generated modified digital image with the one or more image modifications based on the cited model conditioning methods of Garcia by combining noise and the textual feature vector in such a way to provide an optimized and accurate output image which may obviously be further modified according to another or subsequent user input modification instructions to generate the claimed modified image of Liu based on the said model conditioning which may be further realized according to further known methods to yield predictable results since known work in one field of endeavor may prompt variations of it for use in either the same field or a different one based on design incentives or other market forces if the variations are predictable to one of ordinary skill in the art as said combination is thus the adaptation of an old idea or invention using newer technology that is either commonly available and understood in the art thereby a variation on already known art (See MPEP 2143, KSR Exemplary Rationale F). Claim Rejections - 35 USC § 102 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. (a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. Claims 7-11 is/are rejected under 35 U.S.C. 102 (a)(1) as being unpatentable over Liu et al. (US 2022/0108417, A1). Regarding claim 7, Liu teaches at least in the Abstract a system comprising: a memory component (para. 0122); and one or more processing devices coupled to the memory component (para. 0122), the one or more processing devices to perform operations comprising: receiving a first natural language prompt indicating one or more image elements for a digital image (receiving further in Figs. 1-3 and para. 0048-0051 at least input 102 as one of a plurality of natural language prompt indicating one or more image elements for generating a digital image, the user may be in a case received text data of further para. 0050-0051); generating the digital image with the one or more image elements by injecting textual information from the first natural language prompt into a generative neural network as part of a first text-to-image generation process (generated image 114 of at least Figs. 1-3 by injecting in a case textual information of para. 0050-0051 from the first natural language prompt into a generative neural network of at least para. 0054-0055 and 0065-0066 as part of a first text-to-image generation process); receiving a second natural language prompt indicating one or more image modifications to make to the digital image (para. 0048-0051 and Figs. 1-3 further teaches a case of receiving second modification natural language prompt indicating adding of elements as the one or more image modifications to make to the digital image); and generating a modified digital image with the one or more image modifications by injecting textual information from the second natural language prompt into the generative neural network as part of a second text-to-image generation process (Figs. 1-3 and 0048-0051 further indicate the generating modified digital image with the one or more image modifications by injecting textual information from the second natural language prompt into the said generative neural network of at least para. 0054-0055 and 0065-0066 as part of a second text-to-image generation process). Regarding claim 8 (according to claim 7), Liu further teaches wherein generating the modified digital image comprises replacing, within a graphical user interface on a client device, the digital image with the modified digital image in response to receiving the second natural language prompt (at least Figs. 1-3 further illustrates images 218/156 further generating the modified digital image further comprises replacing, within a graphical user interface on a client device, the digital image 114 with the modified digital image 156 in response to receiving the second natural language prompt). Regarding claim 9 (according to claim 7), Liu further teaches wherein the modified digital image comprises the one or more image elements and the one or more image modifications (at least Figs. 1-3 further illustrates images 218/156 further comprises the modified digital image comprises the one or more image elements and the one or more image modifications). Regarding claim 10 (according to claim 7), Liu further teaches wherein generating the modified digital image with the one or more image modifications comprises generating one or more additional image elements (at least Figs. 1-3 further illustrates images 218/156 further comprises the modified digital image with the one or more image modifications comprises generating one or more additional image elements). Regarding claim 11 (according to claim 7), Liu further teaches wherein generating the modified digital image with the one or more image modifications comprises modifying at least one of the one or more image elements (at least Figs. 1-3 further illustrates images 218/156 further comprises the modified digital image with the one or more image modifications comprises modifying at least one of the one or more image elements). Claims 4, 12-15, and 19-20 is/are rejected under 35 U.S.C. 103 as obvious over Liu in view Garcia, and further in view of Zhang et al et al. (US 2023/0081171, A1). Regarding claim 4 (according to claim 1), Liu in view of Garcia are silent regarding wherein generating, utilizing the text encoder, the textual feature vector from the natural language prompt comprises utilizing a contrastive language image pre-training model to encode the natural language prompt into the textual feature vector. Zhang teaches at least in para. 0037-0038 and 0047-0072 a machine learning based text-to-image synthesis system such as a Cross-Modal Contrastive Generative Adversarial Network (XMC-GAN) comprising a text encoder 140 to generate textual feature vector from the natural language prompt comprises utilizing the pre-trained Cross-Modal Contrastive Generative Adversarial Network (XMC-GAN) to encode the natural language prompt into the textual feature vector. It would have been obvious to one of ordinary skill in the art at the time the invention was made to combine the teachings of Liu in view of Garcia, and further in view of Zhang to include wherein said generating, utilizing the text encoder, the textual feature vector from the natural language prompt comprises utilizing a contrastive language image pre-training model to encode the natural language prompt into the textual feature vector, as discussed above, as Liu in view of Garcia and further in view of Zhang are in the same field of endeavor including methods and systems for utilizing a conditioning generative neural network for generating output images corresponding to user inputted text commands, Zhang’s combination architecture of conditioning the generative neural network on the textual feature vector and utilizing a contrastive language image pre-training model to encode the natural language prompt into the textual feature vector further complements the methods of Liu in view of Garcia for generating in a case the modified digital image with the one or more image modifications, in a sense that said method of Liu in view of Garcia when combined to the combination architecture of Zhang, it further enables the systems and methods of Liu in view of Garcia to become of processing and transforming noise vectors from the input data which may be in a case concatenated with the generated text feature elements to output a generated modified digital image as intended by the user inputs with the one or more image modifications according to further known methods to yield predictable results since known work in one field of endeavor may prompt variations of it for use in either the same field or a different one based on design incentives or other market forces if the variations are predictable to one of ordinary skill in the art as said combination is thus the adaptation of an old idea or invention using newer technology that is either commonly available and understood in the art thereby a variation on already known art (See MPEP 2143, KSR Exemplary Rationale F). Regarding claim 12 (according to claim 7), Liu further teaches wherein generating the modified digital image with the one or more image modifications by injecting the textual information from the second natural language prompt into the generative neural network as part of the second text-to-image generation process comprises synthesizing, utilizing the generative neural network, the one or more image modifications However, Liu in view of Garcia are silent regarding wherein synthesizing, utilizing the generative neural network, the one or more image modifications by transforming a noise vector into the one or more image modifications in a manner informed by the second natural language prompt. Zhang further teaches the pre-trained model of at least para 0072 and the disclosure may comprise a transformer configured for concatenating, utilizing the generative neural network, the one or more image elements to as cited “transform a random noise vector and conditioning data derived from the textual description directly into the output image rendition at a target resolution” indicating said into the one or more image elements in a manner informed by the first natural language prompt. It would have been obvious to one of ordinary skill in the art at the time the invention was made to combine the teachings of Liu in view of Garcia, and further in view of Zhang to include wherein said synthesizing, utilizing the generative neural network, the one or more image modifications by transforming a noise vector into the one or more image modifications in a manner informed by the second natural language prompt, as discussed above, as Liu in view of Garcia and further in view of Zhang are in the same field of endeavor including methods and systems for utilizing a conditioning generative neural network for generating output images corresponding to user inputted text commands, Zhang’s combination architecture of conditioning the generative neural network on the textual feature vector and utilizing a transformer for transforming a noise vector into the one or more image elements in a manner informed by the natural language prompt to encode the natural language prompt into the textual feature vector further complements the methods of Liu in view of Garcia for generating in a case the modified digital image with the one or more image modifications, in a sense that said method of Liu in view of Garcia when combined to the combination architecture of Zhang, it further enables the systems and methods of Liu in view of Garcia to process and transform noise vectors from the input data which may be in a case concatenated with the generated text feature elements to output obviously a generated modified digital image as intended by the user inputs with the one or more image modifications according to further known methods to yield predictable results since known work in one field of endeavor may prompt variations of it for use in either the same field or a different one based on design incentives or other market forces if the variations are predictable to one of ordinary skill in the art as said combination is thus the adaptation of an old idea or invention using newer technology that is either commonly available and understood in the art thereby a variation on already known art (See MPEP 2143, KSR Exemplary Rationale F). Regarding claim 13 (according to claim 7), Liu further teaches wherein generating the digital image with the one or more image elements by injecting the textual information from the first natural language prompt into the generative neural network as part of the first text-to-image generation process comprises synthesizing, utilizing the generative neural network, the one or more image elements However, Liu in view of Garcia are silent regarding wherein synthesizing, utilizing the generative neural network, the one or more image elements by transforming a noise vector into the one or more image elements in a manner informed by the first natural language prompt. Zhang further teaches the pre-trained model of at least para 0072 and the disclosure may comprise a transformer configured for concatenating, utilizing the generative neural network, the one or more image elements to as cited “transform a random noise vector and conditioning data derived from the textual description directly into the output image rendition at a target resolution” indicating said into the one or more image elements in a manner informed by the first natural language prompt. It would have been obvious to one of ordinary skill in the art at the time the invention was made to combine the teachings of Liu in view of Garcia, and further in view of Zhang to include wherein said synthesizing, utilizing the generative neural network, the one or more image elements by transforming a noise vector into the one or more image elements in a manner informed by the first natural language prompt, as discussed above, as Liu in view of Garcia and further in view of Zhang are in the same field of endeavor including methods and systems for utilizing a conditioning generative neural network for generating output images corresponding to user inputted text commands, Zhang’s combination architecture of conditioning the generative neural network on the textual feature vector and utilizing a transformer for transforming a noise vector into the one or more image elements in a manner informed by the natural language prompt to encode the natural language prompt into the textual feature vector further complements the methods of Liu in view of Garcia for generating in a case the modified digital image with the one or more image modifications, in a sense that said method of Liu in view of Garcia when combined to the combination architecture of Zhang, it further enables the systems and methods of Liu in view of Garcia to process and transform noise vectors from the input data which may be in a case concatenated with the generated text feature elements to output obviously a generated modified digital image as intended by the user inputs with the one or more image modifications according to further known methods to yield predictable results since known work in one field of endeavor may prompt variations of it for use in either the same field or a different one based on design incentives or other market forces if the variations are predictable to one of ordinary skill in the art as said combination is thus the adaptation of an old idea or invention using newer technology that is either commonly available and understood in the art thereby a variation on already known art (See MPEP 2143, KSR Exemplary Rationale F). Regarding claim 14 (according to claim 13), Liu in view of Garcia are silent regarding wherein the operations further comprise generating, utilizing a text encoder, a textual feature vector from the first natural language prompt utilized contrastive language image pre-training model. Zhang teaches at least in para. 0037-0038 and 0047-0072 a machine learning based text-to-image synthesis system such as a Cross-Modal Contrastive Generative Adversarial Network (XMC-GAN) comprising a text encoder 140 to generate textual feature vector from the natural language prompt comprises utilizing the pre-trained Cross-Modal Contrastive Generative Adversarial Network (XMC-GAN) to encode the natural language prompt into the textual feature vector. It would have been obvious to one of ordinary skill in the art at the time the invention was made to combine the teachings of Liu in view of Garcia, and further in view of Zhang to include wherein said operations further comprise generating, utilizing a text encoder, a textual feature vector from the first natural language prompt utilized contrastive language image pre-training model, as discussed above, as Liu in view of Garcia and further in view of Zhang are in the same field of endeavor including methods and systems for utilizing a conditioning generative neural network for generating output images corresponding to user inputted text commands, Zhang’s combination architecture of conditioning the generative neural network on the textual feature vector and utilizing a contrastive language image pre-training model to encode the natural language prompt into the textual feature vector further complements the methods of Liu in view of Garcia for generating in a case the modified digital image with the one or more image modifications, in a sense that said method of Liu in view of Garcia when combined to the combination architecture of Zhang, it further enables the systems and methods of Liu in view of Garcia to become of processing and transforming noise vectors from the input data which may be in a case concatenated with the generated text feature elements to output a generated modified digital image as intended by the user inputs with the one or more image modifications according to further known methods to yield predictable results since known work in one field of endeavor may prompt variations of it for use in either the same field or a different one based on design incentives or other market forces if the variations are predictable to one of ordinary skill in the art as said combination is thus the adaptation of an old idea or invention using newer technology that is either commonly available and understood in the art thereby a variation on already known art (See MPEP 2143, KSR Exemplary Rationale F). Regarding claim 15 (according to claim 14), Liu is silent regarding wherein generating said digital image with the one or more image elements by injecting the textual information from the first natural language prompt into the generative neural network as part of the first text-to-image generation process comprises combining the noise vector with the textual feature vector. Garcia further teaches at least in Figure 5 and the disclosure a conditional generative model 605 conditioned to generating said digital image with the one or more image elements by injecting the textual information from the first natural language prompt into the generative neural network as part of the first text-to-image generation process comprises cited “textual samples may be combined with noise and may be used by the generative model to generate the images. The training of the generator and discriminator of the GAN may then be performed in an adversarial way as with training the GAN, but the generator is conditioned with the textual samples. This may enable a better convergence of the training” to output a generated generating output digital image, utilizing the generative neural network, from a combination of the noise and the textual feature vector. It would have been obvious to one of ordinary skill in the art at the time the invention was made to combine the teachings of Liu in view of Garcia to include wherein said generating said digital image with the one or more image elements by injecting the textual information from the first natural language prompt into the generative neural network as part of the first text-to-image generation process comprises combining the noise vector with the textual feature vector, as discussed above, as Liu in view of Garcia are in the same field of endeavor including methods and systems for utilizing a conditioning generative neural network for generating output images corresponding to user inputted text/speech commands, Garcia’s combination architecture of conditioning the generative neural network on the textual feature vector comprises combining noise with the textual feature vector further complements the method of Liu of utilizing said generative neural network for generating the modified digital image with the one or more image modifications, in a sense that said method of Liu when combined to the combination architecture of Garcia, it enables the systems and methods of Liu to become advantageously enabled to output the said generated modified digital image with the one or more image modifications based on the cited model conditioning methods of Garcia by combining noise and the textual feature vector in such a way to provide an optimized and accurate output image which may obviously be further modified according to another or subsequent user input modification instructions to generate the claimed modified image of Liu based on the said model conditioning which may be further realized according to further known methods to yield predictable results since known work in one field of endeavor may prompt variations of it for use in either the same field or a different one based on design incentives or other market forces if the variations are predictable to one of ordinary skill in the art as said combination is thus the adaptation of an old idea or invention using newer technology that is either commonly available and understood in the art thereby a variation on already known art (See MPEP 2143, KSR Exemplary Rationale F). Regarding claim 19 (according to claim 18), Liu in view of Garcia are silent regarding wherein synthesizing, utilizing the generative neural network, the one or more image modifications by transforming the noise vector comprises utilizing a transformer to iteratively transform the noise vector. Zhang further teaches the pre-trained model of at least para 0072 and the disclosure may comprise a transformer configured for concatenating, utilizing the generative neural network, the one or more image elements to as cited “transform a random noise vector and conditioning data derived from the textual description directly into the output image rendition at a target resolution” indicating said into the one or more image elements in a manner informed by the first natural language prompt. It would have been obvious to one of ordinary skill in the art at the time the invention was made to combine the teachings of Liu in view of Garcia, and further in view of Zhang to include wherein said synthesizing, utilizing the generative neural network, the one or more image modifications by transforming the noise vector comprises utilizing a transformer to iteratively transform the noise vector, as discussed above, as Liu in view of Garcia and further in view of Zhang are in the same field of endeavor including methods and systems for utilizing a conditioning generative neural network for generating output images corresponding to user inputted text commands, Zhang’s combination architecture of conditioning the generative neural network on the textual feature vector and utilizing a transformer for transforming a noise vector into the one or more image elements in a manner informed by the natural language prompt to encode the natural language prompt into the textual feature vector further complements the methods of Liu in view of Garcia for generating in a case the modified digital image with the one or more image modifications, in a sense that said method of Liu in view of Garcia when combined to the combination architecture of Zhang, it further enables the systems and methods of Liu in view of Garcia to process and transform noise vectors from the input data which may be in a case concatenated with the generated text feature elements to output obviously a generated modified digital image as intended by the user inputs with the one or more image modifications according to further known methods to yield predictable results since known work in one field of endeavor may prompt variations of it for use in either the same field or a different one based on design incentives or other market forces if the variations are predictable to one of ordinary skill in the art as said combination is thus the adaptation of an old idea or invention using newer technology that is either commonly available and understood in the art thereby a variation on already known art (See MPEP 2143, KSR Exemplary Rationale F). Regarding claim 20 (according to claim 16), Liu in view of Garcia are silent regarding wherein generating, utilizing the text encoder, the textual feature vector from the natural language prompt comprises utilizing a contrastive language image pre-training model to encode the natural language prompt into the textual feature vector. Zhang teaches at least in para. 0037-0038 and 0047-0072 a machine learning based text-to-image synthesis system such as a Cross-Modal Contrastive Generative Adversarial Network (XMC-GAN) comprising a text encoder 140 to generate textual feature vector from the natural language prompt comprises utilizing the pre-trained Cross-Modal Contrastive Generative Adversarial Network (XMC-GAN) to encode the natural language prompt into the textual feature vector. It would have been obvious to one of ordinary skill in the art at the time the invention was made to combine the teachings of Liu in view of Garcia, and further in view of Zhang to include wherein said generating, utilizing the text encoder, the textual feature vector from the natural language prompt comprises utilizing a contrastive language image pre-training model to encode the natural language prompt into the textual feature vector, as discussed above, as Liu in view of Garcia and further in view of Zhang are in the same field of endeavor including methods and systems for utilizing a conditioning generative neural network for generating output images corresponding to user inputted text commands, Zhang’s combination architecture of conditioning the generative neural network on the textual feature vector and utilizing a contrastive language image pre-training model to encode the natural language prompt into the textual feature vector further complements the methods of Liu in view of Garcia for generating in a case the modified digital image with the one or more image modifications, in a sense that said method of Liu in view of Garcia when combined to the combination architecture of Zhang, it further enables the systems and methods of Liu in view of Garcia to become of processing and transforming noise vectors from the input data which may be in a case concatenated with the generated text feature elements to output a generated modified digital image as intended by the user inputs with the one or more image modifications according to further known methods to yield predictable results since known work in one field of endeavor may prompt variations of it for use in either the same field or a different one based on design incentives or other market forces if the variations are predictable to one of ordinary skill in the art as said combination is thus the adaptation of an old idea or invention using newer technology that is either commonly available and understood in the art thereby a variation on already known art (See MPEP 2143, KSR Exemplary Rationale F). Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to MARCELLUS AUGUSTIN whose telephone number is (571)270-3384. The examiner can normally be reached 9 AM- 5 PM. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, BENNY TIEU can be reached at 571-272-7490. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /MARCELLUS J AUGUSTIN/Primary Examiner, Art Unit 2682 08/07/2026
Read full office action

Prosecution Timeline

Nov 19, 2024
Application Filed
Aug 12, 2026
Non-Final Rejection mailed — §102, §103
Sep 10, 2026
Interview Requested

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12749176
FOURIER TRANSFORM BASED MACHINE LEARNING FOR DEFECT EXAMINATION OF SEMICONDUCTOR SPECIMENS
2y 9m to grant Granted Sep 29, 2026
Patent 12743027
Methods And Systems For Model-less, Scatterometry Based Measurements Of Semiconductor Structures
3y 5m to grant Granted Sep 22, 2026
Patent 12738007
INFORMATION PROCESSING APPARATUS, INFORMATION PROCESSING METHOD, AND RECORDING MEDIUM
2y 8m to grant Granted Sep 15, 2026
Patent 12731379
CALIBRATING OUTPUT FROM AN IMAGE CLASSIFIER
3y 11m to grant Granted Sep 08, 2026
Patent 12725384
METHOD AND APPARATUS FOR PROVIDING FEEDBACK TO A USER INPUT
3y 6m to grant Granted Sep 01, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
82%
Grant Probability
98%
With Interview (+16.0%)
2y 7m (~8m remaining)
Median Time to Grant
Low
PTA Risk
Based on 869 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month