DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claim(s) 1, 5-6, 8 – 9, 13-14, 16-17, and 20 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Kumari et al. (made of reference in ids: US 20240185588 A1).
Regarding claim 1, Kumari teaches A method for a domain-specific attribute-adapter (Para. 03 – 06, 21-25, 53-55: teaches a method for fine tuning a pre trained text to image diffusion model using a limited collection images representing a new concept or sematic domain by updating model parameters. This selective adaptation functions as an adapter for incorporating new domain specific knowledge into the diffusion model), the method comprising: learning domain-specific attributes from a collection of domain-specific images (Para 21-25, 50-55, 63-68 and 76-80: teaches receiving training images representing a new concept and fine tuning the diffusion model using those images. The learned concept corresponds to attributes extracted from the domain specific image collection); encoding a latent space of a pre-trained text-to-image diffusion model according to the learned domain-specific attributes from the collection of domain-specific images ( Para 29, 56-62, 86-87 and 90-92: teaches a latent diffusion architecture including an encoder that map images into a latent space, a text transformer generating text-condition vectors, and updates to attention layer projection matrices during fine-tuning ); decoding the latent space in response to a received text prompt and one or more conditions ( Para. 43 -46, 56-62, 68-72: teaches generating text-condition vectors from an input text prompt, performing reverse diffusion conditioned on the text condition vector, and decoding the latent representation using a decoder to synthesize an image); and inferring a series of images based on decoding the latent space in response to the received text prompt and the one or more conditions (Para. 62-72 and 73-85: teaches the reverse diffusion process iteratively denoises latent vectors and generates synthesized images conditioned on the text prompt and learned concept).
Regarding claim 5, Kumari teaches the method of claim 1, in which inferring further comprises disconnecting an image prompt during the inferring ( Para.43 during interface, the diffusion model receives only an input text and generates text features. Para 60-62: the reverse diffusion process receives the latent/noisy image vector and the text condition vector to generate the output image. 63-68: during inference, the user simply provides a text prompt and the model generates a new image. ).
Regarding claim 6, Kumari teaches the method of claim 1, in which decoding comprises modeling and providing domain-specific attribute conditions C using a decoder (Kumari, Para 59-62: teach decoder 355 decoding the synthesized latent vector into the final synthesized image. Para 59-62: also teaches disclose decoder 355 reconstructing the synthesized image from the latent representation. Para. 3-4, 20, 26-30, 55-77 104-110: teach a continuous control model that generates attribute embeddings. Which are provided to the image generation model).
Regarding claim 8, Kumari teaches the method of claim 1, further comprising displaying, through a user interface, the series of images ( Para 32, 38, 63-68 and interface 215: teaches displaying generated synthetic images through user interface 215 after image generation).
Regarding claim 9, Kumari teaches A non-transitory computer-readable medium having program code recorded thereon for a domain-specific attribute-adapter (Para. 03 – 06, 21-25, 53-55: teaches a method for fine tuning a pre trained text to image diffusion model using a limited collection images representing a new concept or sematic domain by updating model parameters. This selective adaptation functions as an adapter for incorporating new domain specific knowledge into the diffusion model), the program code being executed by a processor and comprising: program code to learn domain-specific attributes from a collection of domain-specific images (Para 21-25, 50-55, 63-68 and 76-80: teaches receiving training images representing a new concept and fine tuning the diffusion model using those images. The learned concept corresponds to attributes extracted from the domain specific image collection); encoding a latent space of a pre-trained text-to-image diffusion model according to the learned domain-specific attributes from the collection of domain-specific images ( Para 29, 56-62, 86-87 and 90-92: teaches a latent diffusion architecture including an encoder that map images into a latent space, a text transformer generating text-condition vectors, and updates to attention layer projection matrices during fine-tuning ); decoding the latent space in response to a received text prompt and one or more conditions ( Para. 43 -46, 56-62, 68-72: teaches generating text-condition vectors from an input text prompt, performing reverse diffusion conditioned on the text condition vector, and decoding the latent representation using a decoder to synthesize an image); and inferring a series of images based on decoding the latent space in response to the received text prompt and the one or more conditions (Para. 62-72 and 73-85: teaches the reverse diffusion process iteratively denoises latent vectors and generates synthesized images conditioned on the text prompt and learned concept).
Regarding claim 13, it falls under the same rejection as claim 5 it is similar in scope dependent upon same references.
Regarding claim 14, it falls under the same rejection as claim 6 it is similar in scope dependent upon same references.
Regarding claim 16, it falls under the same rejection as claim 8 it is similar in scope dependent upon same references.
Regarding claim 17, Kumari teaches A system for a domain-specific attribute-adapter, the system comprising: a domain-specific attributes learning model to learn domain-specific attributes from a collection of domain-specific images; a latent space encoding model to encode a latent space of a pre-trained text-to-image diffusion model according to the learned domain-specific attributes from the collection of domain-specific images (Para 21-25, 50-55, 63-68 and 76-80: teaches receiving training images representing a new concept and fine tuning the diffusion model using those images. The learned concept corresponds to attributes extracted from the domain specific image collection); encoding a latent space of a pre-trained text-to-image diffusion model according to the learned domain-specific attributes from the collection of domain-specific images ( Para 29, 56-62, 86-87 and 90-92: teaches a latent diffusion architecture including an encoder that map images into a latent space, a text transformer generating text-condition vectors, and updates to attention layer projection matrices during fine-tuning ); decoding the latent space in response to a received text prompt and one or more conditions ( Para. 43 -46, 56-62, 68-72: teaches generating text-condition vectors from an input text prompt, performing reverse diffusion conditioned on the text condition vector, and decoding the latent representation using a decoder to synthesize an image); and inferring a series of images based on decoding the latent space in response to the received text prompt and the one or more conditions (Para. 62-72 and 73-85: teaches the reverse diffusion process iteratively denoises latent vectors and generates synthesized images conditioned on the text prompt and learned concept).
Regarding claim 20, it falls under the same rejection as claim 8 it is similar in scope dependent upon same references.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention
was made.
Claim(s) 2, 4, 7, 10, 12, 15, 18 are rejected under 35 U.S.C. 103 as being unpatentable over Kumari et al. (made of reference in ids: US 20240185588 A1) in view of Cheng (US 20250259340 A1) .
Regarding claim 2, Kumari teaches the method of claim 1, in which inferring the series of images comprises controlling, by an image prompt (IP) adapted text-to-image (T2I) ( Para. 27-29, 43, 60, 65-68 and 74: teaches generating synthetic images from text prompts/text conditions), but fails to teach generation of the series of images using domain-specific continuous attribute conditions C and on inference prediction from a conditional latent space Z without an image prompt X'.
Cheng Teaches the generation of the series of images using domain-specific continuous attribute conditions C and on inference prediction from a conditional latent space Z without an image prompt X' (Para. 71-77: teaches text embedding space, attribute embedding space, vector representations, combining text and attribute embeddings, and feeding them into a diffusion model. Para. 49-53, 67-77 and 108-110 teaches image generation is performed from the text prompt and attribute embedding. There is no requirement for an input image prompt/ reference during interface. It would have been obvious to incorporate the continuous attribute control techniques of Cheng into the latent diffusion text to image framework of Kumari in order to provide finer, continuous control over generated image attributes.
Regarding claim 4, Kumari in view of Cheng teaches the method of claim 1, in which encoding comprises separately performing image content embedding of an image prompt from a text embedding of the received text prompt ( Cheng, 3-4: describes generating a text embedding and a separate attribute embedding, then combining them. Para.29: continuous control model generates an attribute embedding. Para 71-77: separately generates a text embedding from the text prompt and an attribute embedding which is combined later for image generation).
Regarding claim 7, The method of claim 1, in which encoding comprises conditioning the latent space of the pre-trained text-to-image diffusion model on particular attributes for a specific domain ( Kumari, Para 20, 29, 57-61: teach a latent diffusion model operating in latent space conditioned on text condition vectors ), in which the learned domain-specific attributes comprise a pose, angle, point-of-view (POV), and/or a size of an in-domain object (Cheng, Para. 20-30, 33-39, 67-77 and 10-14-110: teach conditioning generation using continuous attribute embeddings for a particular object/domain).
Regarding claim 10, it falls under the same rejection as claim 2 it is similar in scope dependent upon same references.
Regarding claim 12, it falls under the same rejection as claim 4 it is similar in scope dependent upon same references.
Regarding claim 15, it falls under the same rejection as claim 7 it is similar in scope dependent upon same references.
Regarding claim 18, it falls under the same rejection as claim 2 it is similar in scope dependent upon same references.
Claim(s) 3, 11, 19 is rejected under 35 U.S.C. 103 as being unpatentable over Kumari et al. (made of reference in ids: US 20240185588 A1) in view of Fortkort (US-20250363304-A1) .
Regarding claim 3, Kumari teaches The method of claim 1, in which encoding comprises generating the latent space, but fails to teach using a conditional variational autoencoder (CVAE).
Fortkort teaches using a conditional variational autoencoder (CVAE) ( Para.377: teaches using a conditional variational autoencoder to encode data into a continuous latent space. It would have been obvious to implement the latent-space encoder of Kumari using the CVAE architecture taught by Fortkort because CVAE are a known technique for encoding data into a conditional latent space).
Regarding claim 11, it falls under the same rejection as claim 3 it is similar in scope dependent upon same references.
Regarding claim 19, it falls under the same rejection as claim 3 it is similar in scope dependent upon same references.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Hinz (US 20240320873 A1): discloses training and using text-to-image diffusion models with jointly trained text encoders, latent embeddings, style representations and conditioning mechanisms for controllable image generation.
Murez et al (US 11620527 B2): discloses domain adaptation using joint latent space with domain-agnostic feature learning, encoders/decoders, and reconstruction of domain specific representations.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to LATRELL ANTHONY CREARY whose telephone number is (703)756-1219. The examiner can normally be reached Mon - Fri 7:30am - 4:30pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Xiao WU can be reached on (571) 272-7761. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/LATRELL ANTHONY CREARY/Examiner, Art Unit 2613
/XIAO M WU/Supervisory Patent Examiner, Art Unit 2613