DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Amendment
Applicant’s amendments filed on 12 May 2026 have been entered. Claims 1, 4, 9, 15, and 18 have been amended. Claim 3 have been canceled. Claims 1, 2 and 4-20 are still pending in this application, with claims 1, 9 and 15 being independent.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1, 2, 5-15 and 18-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Liu et al. (US 12346995 B2), referred herein as Liu in view of Zeng et al. (US 20250166237 A1), referred herein as Zeng and Puri et al. (US 20220101047 A1), referred herein as Puri.
Regarding Claim 1, Liu in view of Zeng and Puri teaches a method comprising:
obtaining a reference image an input prompt describing an image element (Liu col 7, ln 10-13: The user input text 220 may include style and/or scene words that describe visual features that the user desires to have represented in the synthesized image 130; col 10, ln 25-29: A first set of synthesized images 502 is generated based at least on the input image 500 and a set of word embeddings that specify a portrait of the first user as a warrior with a mountain background; FIG. 2: 124: input image, 220: user input; FIG. 5: 502: a portrait of user ` AS a warrior, mountain background);
However, Zeng teaches
generating an object mask Zeng [0076] the extracted foreground masks are representations of each subject's pose separated from any background image);
Puri teaches
generating an object mask indicating a location of the object from the reference image (Puri [0043] The region identifier 204 may further generate a segmentation mask 222 that includes a segment 212B that corresponds to the object based on identifying the region 212A. The region identifier 204 may further detect a location of the object to define an area 230 of the source image 220 and/or segmentation mask 222. The pre-processor 206 may process at least a portion of the segment 212B in the area 230 of the segmentation mask 222 to produce an object mask 232).
The prior art further teaches
generating, using an image generation model, image features representing the object based on the reference image and the object mask (Liu Abst: The diffusion model is configured to receive the input feature vector and generate a synthesized image of the user based at least on the input feature vector; col 4, ln 19-30: A trained machine learning diffusion model 128 is configured to receive the image 124 of the user and generate a synthesized image 130 of the user based at least on the image 124 of the user captured via the camera 110. The synthesized image 130 includes a character having the same or similar visual features as the user (e.g., the same eye shape, eye color, nose shape, mouth shape, cheek bone shape, complexion etc.); col 9, ln 55-57: the video stream 416 includes the face of the user cropped or masked and overlaid on the synthesized image 414); and
generating, using the image generation model, a synthetic image (Liu ln 19-30: the synthesized image 130 includes additional stylized visual features. For example, the character may have different clothes, assume different body poses, and/or may be placed in a different scene) depicting the object at the location indicated by the object mask within the background scene (Puri [0027] The object mask data may be used to de-emphasize or otherwise modify the background of the source image; [0071] The method 600, at block B604, includes determining image data representative of the object based at least on the region. For example, the image data determiner 208 may determine image data representative of the object based at least on the region 212A of the object; [0072] The method 600, at block B606, includes generating a second image that includes the object having a second background using the image data) described by the input prompt based on the input prompt, the object mask, and the image features from the reference image (Zeng [0115] The generated images have as the foreground image input 601 with the subject “fox” 611 in different poses… The background of images 603, 604, 605, and 606 are scenes described by prompt 602. image 603 is input 601 with a background scene of a forest in spring).
Zeng discloses methods for using neural networks for generating multiple related images, which is analogous to the present patent application.
It would have been obvious for a person of ordinary skill in the art before the effective filing date of the claimed invention to have modified Liu to incorporate the teachings of Zeng, and apply the generated images from a text prompt and an input image into methods for generating a synthesized image of a user with a trained machine learning diffusion model.
Doing so would improve neural networks that generate images as well as ways to improve training of these neural networks.
Puri discloses a segmentation mask that may be generated and used to generate an object image that includes image data representing the object., which is analogous to the present patent application.
It would have been obvious for a person of ordinary skill in the art before the effective filing date of the claimed invention to have modified Liu to incorporate the teachings of Puri, and apply the object mask data into methods for generating a synthesized image of a user with a trained machine learning diffusion model.
Doing so would improve inferencing performed by a neural network trained using object mask data.
Regarding Claim 2, Liu in view of Zeng and Puri teaches the method of claim 1, and further teaches wherein: the reference image depicts the object in a first scene and the synthetic image depicts the object in a second scene described by the input prompt (Liu col 4, ln 26-29: the synthesized image 130 includes additional stylized visual features. For example, the character may have different clothes, assume different body poses, and/or may be placed in a different scene).
Regarding Claim 5, Liu in view of Zeng and Puri teaches the method of claim 1, and further teaches wherein generating the image features comprises: generating a plurality of layer-specific image features at a plurality of layers of the image generation model, respectively (Liu col 6, ln 1-14: the plurality of pre-trained layers are from a pre-trained Contrastive Language-Image Pre-training (CLIP) ViT. In some implementations, the image encoder 200 further includes a plurality of fine-tuned layers 210 (e.g., 8 fine-tuned layers) that are re-trained specifically to extract visual features of the user from the image 124 of the user. In some implementations, the plurality of fine-tuned layers are re-trained to extract visual features of a face of the user. In other implementations, the plurality of fine-tuned layers are re-trained to extract visual features of a body of the user. In some implementations, the image encoder 200 includes a fully connected layer 212 configured to generate the set of embeddings 206 based at least on the visual features of the face of the user extracted by the plurality of fine-tuned layers 210).
Regarding Claim 6, Liu in view of Zeng and Puri teaches the method of claim 1, and further teaches wherein: the image features are generated based on a plurality of reference images (Liu col 10, ln 18-21: FIGS. 5 and 6 show how different input images of different users produce different synthesized images even though the same word embeddings are used to generate the different synthesized images).
Regarding Claim 7, Liu in view of Zeng and Puri teaches the method of claim 1, and further teaches wherein: the input prompt comprises a nonce token corresponding to the object (Liu col 6, ln30-36: the user identifier is implanted in different word embeddings or combined with different sentences by the text encoder 202, and the diffusion model 204 synthesizes a character (i.e., a synthetically generated image of a person) having the visual features of the user in different contexts based at least on the user identifier and the word embeddings and/or sentences; Fig. 5: a portrait of user 1 as …).
Regarding Claim 8, Liu in view of Zeng and Puri teaches the method of claim 1, and further teaches wherein: the image generation model is fine-tuned to generate images depicting the object based on the reference image (Liu col 2, ln 38-41: these training images require specific visual characteristics in order for the conventional diffusion model to be fine-tuned to accurately learn visual features of the user).
Regarding Claim 9, Liu in view of Zeng and Puri teaches a method of training an image generation model, comprising (Liu col 2, ln 19-21: One type of conventional generative diffusion model is pre-trained on a large amount of image and corresponding text data to generate image content based on text inputs; col 4, ln 19-22: A trained machine learning diffusion model 128 is configured to receive the image 124 of the user and generate a synthesized image 130 of the user based at least on the image 124 of the user captured via the camera 110):
obtaining a training set (Liu col 7, ln 56-59: the set of training images 300 includes a plurality of synthesized training images 302 that are generated based at least on an initial training image of the set of training images 300), including
The prior art further teaches
training, using the training set and the synthetic image, the image generation model to generate synthetic images that preserve visual details of the object (Liu col 8, ln 41-53: The set of training feature vectors 312 are fed to the diffusion model 204 and the diffusion model 204 generates a set of predicted synthesized images 314 based at least on the set of set of training feature vectors 312. The diffusion model 204 compares the set of predicted synthesized images 314 to the original set of training images 300 and the amount of noise 304. The diffusion model 204 learns the differences between each of the training images 300 and the predicted synthesized images 314 by making iterative adjustments based at least on minimizing a loss function between each training image and each predicted synthesized image to train the diffusion model 204).
The rest claimed limitations of the claim substantially correspond to the limitations set forth in claim 1; thus they are rejected on similar grounds and rationale as their corresponding limitations.
Regarding Claim 10, Liu in view of Zeng and Puri teaches the method of 9, and further teaches
wherein: the image generation model is pre-trained in a first training phase Liu Fig. 2: 208: pre-trained layer, 210: fine-runed layers, 204: diffusion model; col 2, ln 19-28: One type of conventional generative diffusion model is pre-trained on a large amount of image and corresponding text data to generate image content based on text inputs. For this conventional diffusion model to generate synthesized photorealistic images of a particular user, it needs to be fine-tuned with multiple (10 or more) training images of the user. Once this conventional diffusion model is fine tuned for a particular user, it can generate various synthesized images of a synthetic person who has similar visual features as the user).
Zeng further teaches without receiving the image features at an attention layer (Zeng [0069] neural network 102 includes a convolution layer 103, a self-attention layer 104, and a cross-attention layer 105. In at least one embodiment, neural network 102 includes a diffusion model that is pre-trained model and learned to reverse a diffusion process).
Regarding Claim 11, Liu in view of Zeng and Puri teaches the method of 10, and further teaches wherein: each layer of the image generation model is updated during the second training phase (Liu col 8, ln 54-56: different elements of the machine learning diffusion model 128 can be trained together in some phases and separately in other phases; col 8, ln 66 – col 9, ln 4: Note that the image encoder 200 is pretrained based at least on the set of training images 300 during the training phase. Then, during the inference phase, the image encoder 200 is fine-tuned on the input image 124 of the user. The other elements of the machine learning diffusion model 128 are pretrained during the training phase; Zeng [0066] specific layers of a neural network are trained to identify what features of a subject are common in different images and how different text prompts correspond to image features; [0067] a generated feature map, which includes identified features and weight values, is then input into another layer that identifies image features that correspond to text prompts. In at least one embodiment, a layer receives a feature map (for each input image) and text prompts corresponding to each feature map).
Regarding Claim 12, Liu in view of Zeng and Puri teaches the method of 10, and further teaches wherein: the attention layer receives a different number of input tokens during the first training phase and the second training phase (Zeng [0084] receives one or more inputs 402 and generates one or more outputs images 408 where each image of output images 407 includes a shared image object (e.g., one of the subjects from subject set 203 {x.sub.n, x.sub.n+1, x.sub.n+2, . . . N} where each of x.sub.n, x.sub.n+1, x.sub.n+2, N are different subjects among the list of subjects) displayed over one or more different backgrounds; [0085] The feature image maps generated by convolution layer 403 are received by one or more concatenation layers 404 that concatenate each of the feature image maps into a single feature map, where said single feature map provided to one or more self-attention layers 405. In at least one embodiment, self-attention layers 405, performed by one or more processors, compare each element of said single feature map to determine which elements are more alike each other element).
Regarding Claim 13, Liu in view of Zeng and Puri teaches the method of 9, and further teaches wherein training the image generation model comprises:
computing a diffusion loss (Zeng [0100] input images 502, 503 can be better adapted to personalized image generation by adding an extra mask and input image channel to the diffusion model training; [0102] the training loss is: (formular in [00002])); and
updating parameters of the image generation model based on the diffusion loss (Zeng [0069] a diffusion model includes learned parameters to predict and subtract noise at each step; during model training 4114, by having reset or replaced output or loss layer(s) of initial model 4504, [0659] parameters may be updated and re-tuned for a new data set based on loss calculations associated with accuracy of output or loss layer(s) at generating predictions on new, customer dataset 4506).
Regarding Claim 14, Liu in view of Zeng and Puri teaches the method of 9, and further teaches wherein: the image generation model is trained to receive layer-specific image features for the object at a plurality of different layers (Liu col 5, ln 65 - col 6, ln 8: the image encoder 200 includes a plurality of pre-trained layers 208… the image encoder 200 further includes a plurality of fine-tuned layers 210 (e.g., 8 fine-tuned layers) that are re-trained specifically to extract visual features of the user from the image 124 of the user. In some implementations, the plurality of fine-tuned layers are re-trained to extract visual features of a face of the user).
Regarding Claim 15, Liu in view of Zeng and Puri teaches an apparatus, comprising; a memory component; and a processing device coupled to the memory component, the processing device configured to perform operations (Liu FIG. 1; Claim 1: one or more processors configured to execute instructions stored in memory).
The metes and bounds of the claim substantially correspond to the limitations set forth in claim 1; thus they are rejected on similar grounds and rationale as their corresponding limitations.
Regarding Claim 18, Liu in view of Zeng and Puri teaches the apparatus of claim 15, and further teaches further comprising: a mask generation network configured to generate an object mask that indicates a location of the object in the reference image (Puri [0043] The region identifier 204 may further generate a segmentation mask 222 that includes a segment 212B that corresponds to the object based on identifying the region 212A. The region identifier 204 may further detect a location of the object to define an area 230 of the source image 220 and/or segmentation mask 222. The pre-processor 206 may process at least a portion of the segment 212B in the area 230 of the segmentation mask 222 to produce an object mask 232).
Regarding Claim 19, Liu in view of Zeng and Puri teaches the apparatus of claim 15, and further teaches further comprising: a text encoder configured to encode the input prompt (Karpman col 6, ln 30-32: the user identifier is implanted in different word embeddings or combined with different sentences by the text encoder 202).
Regarding Claim 20, Liu in view of Zeng and Puri teaches the apparatus of claim 15, and further teaches wherein: the attention layer comprises a self-attention layer (Zeng [0069] neural network 102 includes a convolution layer 103, a self-attention layer 104, and a cross-attention layer 105).
Claim(s) 4 is/are rejected under 35 U.S.C. 103 as being unpatentable over Liu et al. (US 12346995 B2), referred herein as Liu in view of Zeng et al. (US 20250166237 A1), referred herein as Zeng, Puri et al. (US 20220101047 A1), referred herein as Puri and Cohen et al. (US 20190196698 A1), referred herein as Cohen.
Regarding Claim 4, Liu in view of Zeng and Puri teaches the method of claim 1. However Cohen teaches wherein: the input prompt describes the location of the object in the reference image, wherein generating the image features comprises determining the location of the object based on the input prompt, and wherein the image features are generated based on the location of the object (Cohen [0096] Continuing with the example directed user conversation 204 in FIG. 2, the user responds “A cloudy sky”, indicating to the device that the boring sky should be replaced by a cloudy sky. In response, the device generates harmonized image 206, which includes a cloudy sky. Harmonized image 206 includes is natural looking and lacks artifacts of editing often present in composite images, such as halos).
Cohen discloses systems and techniques for directing a user conversation to obtain an editing query, and removing and replacing objects in an image based on the editing query, which is analogous to the present patent application.
It would have been obvious for a person of ordinary skill in the art before the effective filing date of the claimed invention to have modified Liu to incorporate the teachings of Puri, and apply the conversation module into methods for generating a synthesized image of a user with a trained machine learning diffusion model.
Doing so, images can be efficiently provided to a user that satisfy an editing query and at the same time instruct the user on the use of the editing application while using the user's actual data.
Claim(s) 16 and 17 is/are rejected under 35 U.S.C. 103 as being unpatentable over Liu et al. (US 12346995 B2), referred herein as Liu in view of Zeng et al. (US 20250166237 A1), referred herein as Zeng, Puri et al. (US 20220101047 A1), referred herein as Puri, and Karpman et al. (US 11995803 B1), referred herein as Karpman.
Regarding Claim 16, Liu in view of Zeng and Puri teaches the apparatus of claim 15. However, Karpman teaches wherein: the image generation model comprises a diffusion U-Net (Karpman col 5, ln 20-21: The base image diffusion model 120 can include a U-net architecture (e.g., Efficient U-Net)).
Karpman discloses a method receives a text prompt and executes a text encoder on the text prompt to generate an embedding, which is analogous to the present patent application.
It would have been obvious for a person of ordinary skill in the art before the effective filing date of the claimed invention to have modified Liu to incorporate the teachings of Karpman, and apply the Efficient U-Net into methods for generating a synthesized image of a user with a trained machine learning diffusion model.
Doing so would provide a new and useful systems and methods for training and deploying text-to-image generative models.
Regarding Claim 17, Liu in view of Zeng, Puri and Karpman teaches the apparatus of claim 16, and further teaches wherein: the image generation model receives the image features at a plurality of attention layers corresponding to a plurality of decoder layers of the diffusion U-Net (Karpman col 5, ln 20-28: The base image diffusion model 120 can include a U-net architecture (e.g., Efficient U-Net) defined from residual and multi-head attention blocks that enable the base image diffusion model 120 to progressively denoise (e.g., infill, generate, augment) image data according to cross-attention inputs based on the text prompt. The base image diffusion model 120 can therefore: receive one or more text embeddings from the set of pre-trained text encoders 118).
Response to Arguments
Applicant’s arguments, see page 1, filed on 18 May 2026, with respect to claim objections have been fully considered and are persuasive. The objections of 18 November 2025 has been withdrawn.
Applicant’s arguments, filed on 05 May 2026 with respect to103 rejection on claims 1, 24 and 25 has been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Samantha (Yuehan) Wang whose telephone number is (571)270-5011. The examiner can normally be reached Monday-Friday, 8am-5pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, King Poon can be reached at (571)272-7440. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/Samantha (YUEHAN) WANG/
Primary Examiner
Art Unit 2617