Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Continued Examination Under 37 CFR 1.114
A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 08/03/2026 has been entered.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claim(s) 16-17 and 19-20 is/are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Zhang (Zhang, Lvmin, Anyi Rao, and Maneesh Agrawala. "Adding conditional control to text-to-image diffusion models." Proceedings of the IEEE/CVF international conference on computer vision. 2023.)
Regarding claim 16, teaches a non-transitory computer-readable medium storing executable instructions that, when executed by a processing device, cause the processing device to perform operations (see §1 “We also show that in some tasks like depth-to-image, training ControlNets on a personal computer (one Nvidia RTX 3090TI) can achieve competitive results to commercial models trained on large computation clusters with terabytes of GPU memory and thousands of GPU hours.”), comprising:
receiving a text prompt and an image prompt for generating a digital image (Figure 3: Prompt → Text Encoder → Each layer of stable diffusion, i.e. a text prompt, and Condition → … → zero convolution → a layer of SD Decoder, i.e. an image prompt, and Output, i.e. a digital image, see Figure 4);
conditioning a first upsampling layer of a neural network with an image vector representation of the image prompt, the first upsampling layer operating at a first resolution (Figure 3: Any of zero convolution → a block of SD Decoder, which is at a certain resolution, e.g. 8x8);
conditioning a third upsampling layer of the neural network with a text vector representation of the text prompt without the image vector representation of the image prompt (Figure 3: Any of the other SD Decoder blocks with a Text Encoder input. The Text Encoder input does not include the image vector representation input into the first upsampling layer); and
generating, utilizing the neural network, the digital image from the image vector representation and the text vector representation (Figure 4).
To clarify, “third” is not interpreted to indicate that there exists an unclaimed second upsampling layer. Under the broadest reasonable interpretation, it merely indicates a unique upsampling layer. See MPEP 2111.03, I, “The court also emphasized that reference to "first," "second," and "third" blades in the claim was not used to show a serial or numerical limitation but instead was used to distinguish or identify the various members of the group.”
Regarding claim 17, Zhang teaches all of the limitations of claim 16, wherein:
conditioning the first upsampling layer of the neural network comprises conditioning a high-resolution upsampling layer of the neural network with the image vector representation of the image prompt (Figure 3, any of the 16x16, 32x32, or 64x64 decoder blocks, each of which has three layers); and
conditioning the third upsampling layer of the neural network comprises conditioning a low-resolution upsampling layer of the neural network with the text vector representation of the text prompt, wherein the high-resolution upsampling layer has a higher resolution than the low-resolution upsampling layer (Figure 3, Decoder block 8x8).
Regarding claim 19, Zhang teaches all of the limitations of claim 16, wherein generating, utilizing the neural network, the digital image from the image vector representation and the text vector representation comprises:
generating a first noise representation utilizing a first neural network of a first denoising iteration of a diffusion neural network; and generating a second noise representation utilizing a second neural network of a second denoising iteration of the diffusion neural network (Figure 3, Encoder and Decoder).
Regarding claim 20, Zhang teaches all of the limitations of claim 16, wherein the operations further comprise
conditioning a plurality of downsampling layers of the neural network with the text vector representation of the text prompt (Figure 3: Encoder layers).
Allowable Subject Matter
Claims 1-15 are allowed.
Claim 18 is objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Regarding claims 1 and 9, the prior art does not anticipate or render obvious the passing of the image vector representation (or the first vector representation) to separate upsampling layers where the upsampling layers are at different resolutions. The closest prior art is represented by Zhang and discloses passing image vector representations to multiple upsampling layers but the image representation is unique to each resolution, i.e. the same representation is not passed to upsampling layers of differing resolution.
Regarding claim 18, the prior art does not anticipate or render obvious the style weight controller.
Response to Arguments
Applicant’s remarks filed 08/03/2026 have been fully considered.
Applicant argues that Zhang fails to disclose the amended limitations in the response filed 08/03/2026. However, claim 16 does not recite the second upsampling layer at the second resolution and therefore the argument directed thereto is moot.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SCHYLER S SANKS whose telephone number is (571)272-6125. The examiner can normally be reached 06:30 - 15:30 Central Time, M-F.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Michael Huntley can be reached at (303) 297-4307. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/SCHYLER S SANKS/Primary Examiner, Art Unit 2129