DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 1-20 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claims 1, 8 and 15 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being incomplete for omitting essential elements, such omission amounting to a gap between the elements. See MPEP § 2172.01. The omitted elements are: user input of an image via a displayed user interface provided by an image processing apparatus as disclosed by Applicant’s Specification (Para 31, 37, 45); providing an input prompt via a user interface or a database as disclosed by Applicant’s Specification (Para 95); an image processing apparatus to generate a synthetic image as disclosed by Applicant’s Specification (Para 32); and a ground truth data and ground truth image for determining a score as disclosed by Applicant’s Specification (Para 25, 27; Fig. 4).
Claims 2-7, 9-14 and 16-20 are rejected based on dependency from a rejected base claim.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claim(s) 1, 3-5, 7, 8, 10-12, 14-15, 18 and 20 is/are rejected under 35 U.S.C. 102(a)(2) as being anticipated by Agoston Weisz et al., US 2025/0329084 A1.
Independent claim 1, discloses a method for image processing, comprising:
obtaining a prompt indicating an image element (Fig. 4 “454”);
generating a first score function using a base image generation model and a second score function using an auxiliary image generation model, wherein the first score function and the second score function are based on the prompt (Fig. 4 “456”; The cross-attention operation can therefore be considered to update or modify the encoding 204 of the input image by attending to the input prompt 203.- Para 46; The self-attention operation can enable the VLM to focus on the most relevant parts of the encoding 204 of the input image in accordance with the instructions in the input prompt 203 – Para 47);
combining the first score function and the second score function to obtain a combined score function (i.e. the VLM can include a combination of cross-attention and self-attention operations – Para 48), wherein the combined score function includes positive guidance from the first score function (i.e. The cross-attention operation can therefore be considered to update or modify the encoding 204 of the input image by attending to the input prompt 203. – Para 46) and negative guidance from the second score function (i.e. The self-attention operation can enable the VLM to focus on the most relevant parts of the encoding 204 of the input image in accordance with the instructions in the input prompt 203 – Para 47); and
generating a synthetic image that depicts the image element based on the combined score function (i.e. The generated output image 206 is a modified version of the input image 202, modified according to the instructions in the input prompt 203- Para 52).
Claim 3, Weisz discloses the method of claim 1, further comprising: generating a training image using the base image generation model (i.e. the image generation model 1545 is configured to generate the training image1550 based on the guidance embedding 1540 – Para 184); and training the auxiliary image generation model using the training image (i.e. an encoding of the input prompt 203 can be concatenated and provided as input to the VLM (e.g. auxiliary image generation model). The VLM can include one or more Transformer blocks with a self-attention operation – Para 47).
Claim 4, Weisz discloses the method of claim 1, further comprising: training another image generation model using the synthetic image as training data (i.e. the image generation engine 162 can interface with an external image generation system 180 to generate an image based upon the modified encoding 205 – Para 51).
Claim 5, Weisz discloses the method of claim 1, wherein combining the first score function and the second score function comprises: identifying a weight parameter, wherein the first score function and the second score function are combined based on the weight parameter (i.e. The plurality of VLMs can be a copy of a single VLM but with different parameters, - Para 59; the parameters of the VLM undergoing fine-tuning are adjusted using a reinforcement learning update rule – Para 63; The system can determine to continue fine-tuning the VLM until one or more conditions are satisfied – Para 64).
Claim 7, Weisz discloses the method of claim 1, wherein: the second score function represents an unnatural image artifact (i.e. self-attention operation can enable the VLM to focus on the most relevant parts of the encoding 204 of the input image – Para 47).
Independent claim 8, the claim is similar in scope to claim 1. Therefore, similar rationale as applied in the rejection of claim 1 applies herein.
Claims 10-12, 14, 18 and 20, the corresponding rationale as applied in the rejection of claims 1, 3-5 and 7 apply herein.
Independent claim 15, the claim is similar in scope to claim 1. Therefore, similar rationale as applied in the rejection of claim 1 applies herein.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claim(s) 2, 6, 9, 13,16-17 and 19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Agoston Weisz et al., US 2025/0329084 A1 as applied to claims 1, 8 and 15 above, and further in view of Zecheng He et al., US 2026/0101081 A1.
Claim 2, Weisz teaches the method of claim 1 includes encoding an input image in a learned latent space (Para 43); and applying differing parameters to the VLM (e.g. generative model – Para 8, 63).
He discloses the method of claim 1, wherein generating the synthetic image comprises: obtaining a noise map (i.e. noise input may be in a latent space – Para 53, 149); and denoising the noise map based on the combined score function (i.e. the image generation model performs a denoising process (e.g., a reverse diffusion process) to denoise the noise input – Para 54; guidance feature 1070 can be combined with the noisy feature 1035 using a cross-attention block within the reverse diffusion process 1040 – Para 141), which Weisz fails to disclose.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention at the time the invention was made to combine He’s known method comprising obtaining a noise map; and denoising the noise map based on the combined score function with the method of Weisz because each discloses a generative model for updating an input image to provide variations of a synthetic image (He, Para 128).
One of ordinary skill would have been motivated to combine He’s known method comprising obtaining a noise map; and denoising the noise map based on the combined score function with the method of Weisz and would have recognized that the results of the combination were predictable.
Claim 6, Weisz teaches the method of claim 1 that uses a multimodal model (Para 4) trained on multimodal data (Para 68); obtaining a prompt (Fig. 4 “454”) and encoding an input prompt (Para 49) indicating an image element based on an embedding from vocabulary (Para 50), wherein the first score function and the second score function are generated based on the prompt embedding (i.e. Fig. 4 “456”; The cross-attention operation can therefore be considered to update or modify the encoding 204 of the input image by attending to the input prompt 203.- Para 46; The self-attention operation can enable the VLM to focus on the most relevant parts of the encoding 204 of the input image in accordance with the instructions in the input prompt 203 – Para 47).
He discloses the method of claim 1, further comprising: encoding the prompt to obtain a prompt embedding (i.e. generating a multimodal embedding based on the input image and the text prompt – abstract; generating, using a multimodal encoder, a multimodal embedding – Para 4; encode the input prompts to generate a multimodal embedding – Para 42), which Weisz fails to disclose.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention at the time the invention was made to combine He’s known method comprising encoding the prompt to obtain a prompt embedding with the method of Weisz because each encodes prompts and provides an embedding that guides the image generation process to generate a synthetic image (He, Para 42).
One of ordinary skill would have been motivated to combine He’s known method comprising encoding the prompt to obtain a prompt embedding with the method of Weisz and would have recognized that the results of the combination were predictable.
Claims 9, 13 and 19, the corresponding rationale as applied in the rejection of claims 2 and 6 apply herein.
Claim 16, Weisz discloses the system of claim 15, wherein the base image generation model is a generative model (Para 70).
Weisz fails to disclose wherein: the base image generation model comprises a diffusion model, which He discloses (i.e. diffusion model 1000 is an example of, or includes aspects of, the image generation model – Para 135; Diffusion models are a class of generative neural networks that can be trained to generate new data with features similar to features found in training data- Para 136).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention at the time the invention was made to combine He’s known method wherein a base image generation model comprises a diffusion model with the method of Weisz because generative models are used in generating images, where a diffusion model is exemplary of a generative model. Thus, the combination yields predictable results.
Claim 17, Weisz discloses the system of claim 15, wherein the base image generation model is a generative model (Para 70).
Weisz fails to disclose wherein: the auxiliary image generation model comprises a diffusion model, which He discloses (i.e. diffusion model 1000 is an example of, or includes aspects of, the image generation model – Para 135; Diffusion models are a class of generative neural networks that can be trained to generate new data with features similar to features found in training data- Para 136).
Similar rationale as applied in the rejection of claim 16 applies herein.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to CHANTE HARRISON whose telephone number is (571)272-7659. The examiner can normally be reached Monday - Friday 8:00 am to 5:00 pm EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Alicia Harrington can be reached at 571-272-2330. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/CHANTE E HARRISON/Primary Examiner, Art Unit 2615