Prosecution Insights
Last updated: October 01, 2026
Application No. 18/947,989

DIFFUSION MODEL MULTI-PERSON IMAGE GENERATION

Final Rejection §103
Filed
Nov 14, 2024
Priority
Nov 17, 2023 — provisional 63/600,451
Examiner
NGUYEN, DAVID VAN
Art Unit
Tech Center
Assignee
Snap Inc.
OA Round
2 (Final)
88%
Grant Probability
Favorable
3-4
OA Rounds
6m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 88% — above average
88%
Career Allowance Rate
7 granted / 8 resolved
+27.5% vs TC avg
Moderate +15% lift
Without
With
+14.6%
Interview Lift
resolved cases with interview
Typical timeline
2y 5m
Avg Prosecution
15 currently pending
Career history
24
Total Applications
across all art units

Statute-Specific Performance

§101
3.8%
-36.2% vs TC avg
§103
85.0%
+45.0% vs TC avg
§102
11.3%
-28.7% vs TC avg
Black line = Tech Center average estimate • Based on career data from 8 resolved cases

Office Action

§103
CTNF 18/947,989 CTNF 101398 DETAILED ACTION Notice of Pre-AIA or AIA Status 07-03-aia AIA 15-10-aia The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA. Claim Rejections - 35 USC § 103 07-06 AIA 15-10-15 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. 07-20-aia AIA The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. 07-21-aia AIA Claim (s) 1-3, 5-6, and 19-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Davis (US 20240331211 A1) and Zhang et al (“ControlCom: Controllable Image Composition using Diffusion Model”), hereinafter Davis and Zhang respectively . Regarding claim 19, Davis teaches a system comprising: at least one processor; and at least one memory component having instructions stored thereon that, “The processor 190 may also communicate with memory 180. The memory 180 may contain computer program instructions (grouped as modules or units in some embodiments) that the processor 190 executes in order to implement one or more aspects of the present disclosure.” – Par 50, Lines 1-5 when executed by the at least one processor, cause the at least one processor to perform operations comprising: accessing first and second artificial personalized images generated by first and second generative machine learning models, “At block 410, the synthesized human generation system 120 utilizes the selected ML models 170 and generates one or more images of one or more synthesized bodies wearing the garment. The ML models 170 may utilize the classifications from 404 and the segmented garment from 406 to generate the image(s) of synthesized bodies wearing the garment.” – Par 46, Lines 1-6 NOTE: Davis teaches machine learning models that are trained on different classifications and used to generate different artificial human bodies wearing clothes. These models would be able to create a first and second artificial image. It is also disclosed that these models are generative machine learning models: “For example, the synthesized human generation system 120 may implement one or more of the ML models 170 as a generative adversarial network (GAN).” – Par 26, Lines 3-5 the first generative machine learning model trained to generate the first artificial personalized image comprising a depiction of a first person, the second generative machine learning model trained to generate the second artificial personalized image comprising a depiction of a second person; “plurality of body generation machine learning models, wherein a first model of the plurality of body generation machine learning models is trained using training images depicting clothed humans of a first classification, wherein the second model of the plurality of body generation machine learning models is trained using training images depicting clothed humans of a second classification.” – Claim 10 NOTE: Davis discloses that a first and second classification are used to train a first and second generative machine learning model to create images of artificial bodies wearing different styles of clothes. This corresponds functionally to having a first model generating a first person and a second model generating a second person since one of ordinary skill in the art can implement the two models trained on different classifications to generate two artificial humans wearing different clothes. A source image of a person or mannequin is used as input to the generative model as a basis for the synthesis of the first and second person, see Abstract. Davis does not teach generating a foreground image that combines the depiction of the first person in the first artificial personalized image with the depiction of the second person in the second artificial personalized image; accessing background information; and generating a new artificial image comprising the foreground image on a background having visual attributes that correspond to the background information. However, Zhang teaches generating a foreground image that combines the depiction of the first person in the first artificial personalized image with the depiction of the second person in the second artificial personalized image; “generative image composition aims to synthesize an image Ic that composites the foreground object into the background, so that the region within the bounding box depicts the object as similar to the foreground object and fits harmoniously, while the other regions remain as same as possible to the background Ib.” – Section 3.1 Par 1, and Fig. 1 PNG media_image1.png 291 520 media_image1.png Greyscale Zhang et al, Fig. 1 NOTE: Zhang teaches a pair of images (foreground and background images) which can be composited together by placing the foreground image into a background image using a diffusion model. After the combination, the method for combining a foreground image to a region in the background image as taught by Zhang can be used by the body generation models as taught by Davis. This would allow the first model to generate an image of a person and set it as a background image so that the second model which generates a second person can be set as the foreground image. Then, the diffusion model of Zhang combines the two images created by Davis to generate a new image with both artificial persons. accessing background information; “Recently, generative image composition [46,52] targets at solving all issues in one unified model, which can greatly simplify the composition pipeline. These methods are generally built on pretrained diffusion model [41], due to its un precedented power in synthesizing realistic images. Specifically, they take in a foreground image and a background image with a user-specified bounding box to produce a realistic composite image, in which a pretrained image encoder [39] extracts the foreground embedding and the diffusion model incorporates this conditional foreground embedding into diffusion process.” – Section 1, Par 2 NOTE: Zhang teaches that the diffusion model takes in a foreground image and a background image as input to combine the images into one. Providing a background image can be interpreted as background information pertaining to the style or attributes the user may want. and generating a new artificial image comprising the foreground image on a background having visual attributes that correspond to the background information. “generative image composition aims to synthesize an image Ic that composites the foreground object into the background, so that the region within the bounding box depicts the object as similar to the foreground object and fits harmoniously, while the other regions remain as same as possible to the background Ib.” – Section 3.1 Par 1, and Fig. 1 NOTE: Zhang teaches the combination of the foreground and background images using the diffusion model. The generation of this new image would naturally use the information of the background image to create a new image of the foreground in front of the background with its corresponding visual attributes. After the combination, the method for combining a foreground and background images as taught by Zhang can be added to Davis’s body generation models to repeat the steps of Zhang to take the newly generated image of the two persons as input with a new desired background image input and combine the two images to make a new combined image with the first and second generated persons and the desired background. It would have been obvious to one of ordinary skill in the art before the effective filing date of the present invention to modify Davis by incorporating the teachings of Zhang to generate a foreground image comprising the generated images of a person from two models and create a new artificial image by combining the foreground image with a background image. One would be motivated to make this combination to give a user the ability depict multiple artificial objects (persons) in an artificial scene. This would have the predictable result of producing a high quality synthesized image that combines artificially created objects into one image. Regarding claim 1, the claim recites similar limitations to claim 19. Therefore, method claim 1 corresponds to the system disclosed in claim 19 and is rejected for the same reasons of obviousness as used above. Regarding claim 20, the claim recites similar limitations to claim 19. Therefore, non-transitory computer readable storage medium claim 20 corresponds to the system disclosed in claim 19 and is rejected for the same reasons of obviousness as used above. Regarding claim 2, Davis in view of Zhang teach the method of claim 1. Davis does not teach accessing a prompt comprising the background information; and processing the foreground image and the prompt by a third generative machine learning model to generate the new artificial image. However, Zhang teaches accessing a prompt comprising the background information; “Recently, generative image composition [46,52] targets at solving all issues in one unified model, which can greatly simplify the composition pipeline. These methods are generally built on pretrained diffusion model [41], due to its un precedented power in synthesizing realistic images. Specifically, they take in a foreground image and a background image with a user-specified bounding box to produce a realistic composite image, in which a pretrained image encoder [39] extracts the foreground embedding and the diffusion model incorporates this conditional foreground embedding into diffusion process.” – Section 1, Par 2 NOTE: The background image that is used as input to the diffusion model can be interpreted as a prompt (specifically an image prompt) that is used with the input foreground image to create a combined image. and processing the foreground image and the prompt by a third generative machine learning model to generate the new artificial image. “generative image composition aims to synthesize an image Ic that composites the foreground object into the background, so that the region within the bounding box depicts the object as similar to the foreground object and fits harmoniously, while the other regions remain as same as possible to the background Ib.” – Section 3.1 Par 1, and Fig. 1 NOTE: the foreground and background images used as input to the diffusion model are then processed to create a new artificial image that combines the foreground subject with the background. After the combination, the diffusion model for combining the foreground and background images as taught by Zhang can be interpreted as the third generative machine learning model that can combine the two generated persons created by the first and second generative machine learning models taught by Davis. It would have been obvious to one of ordinary skill in the art before the effective filing date of the present invention to modify Davis by incorporating the teachings of Zhang to obtain a prompt comprising background information and processing it with a foreground image to generate a new artificial image. One would be motivated to make this combination because it is well known in the art that a user prompt will guide the generation process of the new image. Processing this prompt with the foreground will allow the generative model to create a high quality image of the foreground subject with the desired background details and visuals. Regarding claim 3, Davis in view of Zhang teach the method of claim 2, Davis does not teach wherein the prompt comprises an image depicting the background information. However, Zhang teaches wherein the prompt comprises an image depicting the background information. “Image composition targets at synthesizing a realistic composite image from a pair of foreground and background images” – Abstract, and Fig 1 NOTE: The diffusion model in Fig 1 is shown to take a foreground and background image. The images can be interpreted as prompts used to guide the generation of a new artificial image. This would imply that the background image used as inputs can be understood as an image prompt depicting background information to be combined with the foreground image. It would have been obvious to one of ordinary skill in the art before the effective filing date of the present invention to modify Davis by incorporating the teachings of Zhang to have the background information prompt comprise an image. One would be motivated to make this combination in order for the user to input an image that describes in better detail the style and visual attributes. This would allow the user to create a high quality composite image with the desired background, Regarding claim 5, Davis in view of Zhang teaches the method of claim 2. Davis teaches wherein the first generative machine learning model and the second generative machine learning model comprises a respective diffusion machine learning model. “Additionally or alternatively, the synthesized human generation system 120 may implement one or more of the ML models 170 as a diffusion model” – Par 27, Lines 1-3 Davis does not teach the third generative machine learning model comprising a respective diffusion machine learning model . However, Zhang teaches the third generative machine learning model comprising a respective diffusion machine learning model To address the first issue, we propose a controllable image composition method named ControlCom based on conditional diffusion model” – Introduction Par 4, Lines 3-4 NOTE: After the combination, the third diffusion model configured to create a combined image of a foreground and background as taught by Zhang can be added to the Davis’s system for generating bodies using multiple diffusion models. This would allow the two images created by the two diffusion models respectively to then be combined by Zhang’s diffusion model to create the composite image of the two generated persons together in one image. It would have been obvious to one of ordinary skill in the art before the effective filing date of the present invention to modify Davis by incorporating the teachings of Zhang to have the first, second, and third generative models comprise a respective diffusion machine learning model. One would be motivated to make this combination since it is well-known in the art for diffusion models to be powerful image generation tools and stable during the training process. Regarding claim 6, Davis in view of Zhang teaches the method of claim 1. Davis further teaches wherein the first and second generative machine learning models each comprises a respective diffusion machine learning model. “Additionally or alternatively, the synthesized human generation system 120 may implement one or more of the ML models 170 as a diffusion model” – Par 27, Lines 1-3 07-21-aia AIA Claim (s) 4 is/are rejected under 35 U.S.C. 103 as being unpatentable over Davis, Zhang and Wang et al (“Imagen Editor and EditBench: Advancing and Evaluating Text-Guided Image Inpainting”), hereinafter Wang . Regarding claim 4, Davis in view of Zhang teach the method of claim 2. Davis does not teach wherein the prompt comprises a textual description of the background information. However, Wang teaches wherein the prompt comprises a textual description of the background information. “We focus on text-guided image inpainting, where a user provides an image, a masked area, and a text prompt and the model fills the masked area, consistent with both the prompt and the image context (Fig. 1)” – Introduction Par 1 and Fig. 1 NOTE: Wang discloses a test guided image inpainting that takes a user’s image, desired masked area, and text prompt to fill the masked area with the text. This functionally corresponds to a prompt comprising a textual description of the background image since the mask can be placed over the background. After the combination, the diffusion model to perform inpainting of an image by filling a determined mask based on a text prompt as taught by Wang can be added to Zhang’s diffusion model to allow text prompts regarding background information to be combined with the foreground image. This modification is then added to Davis’s system for generating artificial image of a first and second person with generative machine learning models which can be combined into one image using Zhang’s diffusion model. This combined image of the two generated persons can then be used as input again in Zhang’s diffusion model to generate a new image that combines the image of the two generated persons with a background based off a text prompt. It would have been obvious to one of ordinary skill in the art before the effective filing date of the present invention to modify Davis by incorporating the teachings of Zhang to have the prompt regarding background information be a text description. One would be motivated to make this combination to allow for intuitive use of the diffusion model to easily allow the user to generate new images based on their desired visuals . 07-21-aia AIA Claim (s) 8-10 is/are rejected under 35 U.S.C. 103 as being unpatentable over Davis, Zhang, Epstein et al US 20230260175 A1), hereinafter Epstein . Regarding claim 8, Davis in view of Zhang teaches the method of claim 1. Davis does not teach generating a first segmentation of the first person depicted in the first artificial personalized image. However, Epstein teaches generating a first segmentation of the first person depicted in the first artificial personalized image. “To generate the background digital images and the foreground digital images, the digital image collaging system 102 generates background digital image masks and foreground digital image masks, respectively.”- Par 87, Lines 12-15 NOTE: Epstein discloses that a mask for the background and foreground are generated for an image. This would imply that the foreground and background are segmented. After the combination, the system for segmentation of the foreground and background as taught by Epstein can be applied to the generated images of a person from one generative model as taught by Davis. This would allow the foreground object (person) to be segmented from the background. It would have been obvious to one of ordinary skill in the art before the effective filing date of the present invention to modify Davis by incorporating the teachings of Epstein to generate a first segmentation of the first person depicted in the artificial personalized image. One would be motivated to make this combination since it is well known in the art for accurately labelling objects or regions within an image. This would have the predictable result of a cleaner composite image when combining multiple images as it would prevent the use of unnecessary pixels that do not make up the object of interest. Regarding claim 9, Davis in view of Zhang, and Epstein teaches the method of claim 8. Davis does not teach extracting a first region of the first artificial personalized image based on the first segmentation, the first region comprising pixels that fall within the first segmentation; extracting a second region of the second artificial personalized image based on the first segmentation, the second region of the second artificial personalized image excluding pixels that fall within the first segmentation; and generating the foreground image by combining the first region and the second region. However, Epstein further teaches extracting a first region of the first artificial personalized image based on the first segmentation, the first region comprising pixels that fall within the first segmentation; “Along these lines, a “foreground digital image” refers to a digital image depicting pixels belonging to a foreground region or a foreground layer of a digital image” – Par 33, Lines 15-17 NOTE: Epstein discloses pixels belonging to the foreground region. The foreground region can be interpreted as the segmented portion of the foreground containing the object (a building as shown in Fig. 7). After the combination, the system for extracting a first region based on the pixels that fall within the first segmentation can be used by Davis’s system for generating the first person with the first generative model. This would allow the extraction of a first region comprising the pixels that make up the person in the image. extracting a second region of the second artificial personalized image based on the first segmentation, the second region of the second artificial personalized image excluding pixels that fall within the first segmentation; “Relatedly, a “background digital image” refers to a digital image including or depicting pixels belonging to a background region or a background layer of a digital image.” – Par 33, Lines 12-14 NOTE: Epstein discloses the use of two images of buildings that extract a foreground and background mask. Naturally, the background region would encompass pixels that fall outside the foreground region. After the combination, the system for extracting a region that lies outside the first segmentation as taught by Epstein can be used by the system for generating a first and second person using a first and second generative model as taught by Davis. This would allow Davis’s system to take the first generated image and extract the foreground object by segmenting the person and then extract the background region from the second image which would encompass the pixels outside of the foreground segment. and generating the foreground image by combining the first region and the second region. “Additionally, the digital image collaging system 102 generates the collage digital images by collaging (e.g., alpha compositing) the background digital images and the foreground digital images together.” – Par 87, Lines 13-16. And Fig. 7 NOTE: Epstein discloses an alpha compositing technique that will combine the first image’s foreground region with a second image’s background region. This will create a composite image where the foreground object of the first image is placed in front of the second image’s background, see Fig. 7. PNG media_image2.png 332 543 media_image2.png Greyscale Epstein et al, Fig. 7 It would have been obvious to one of ordinary skill in the art before the effective filing date of the present invention to modify Davis by incorporating the teachings of Epstein to extract a first region of the first generated image based on the first segmentation, extract a second region from the second generated image based on the first segmentation, and combine the first and second region. One would be motivated to make this combination as it would result in a high quality composite image that separates the regions in order to cleanly blend both visual attributes. Regarding claim 10, Davis in view of Zhang and Epstein teach the method of claim 9. Davis does not teach generating a second segmentation of the second person depicted in the second artificial personalized image; and combining the first and second segmentations to generate a combined segmentation. However, Epstein further teaches generating a second segmentation of the second person depicted in the second artificial personalized image; “To generate the background digital images and the foreground digital images, the digital image collaging system 102 generates background digital image masks and foreground digital image masks, respectively.”- Par 87, Lines 12-15 NOTE: The process for generating a second segmentation of the second person can be done using the same steps as the generation of the first segmentation as explained in the rejection of claim 8. and combining the first and second segmentations to generate a combined segmentation “Additionally, the digital image collaging system 102 generates the collage digital images by collaging (e.g., alpha compositing) the background digital images and the foreground digital images together.” – Par 87, Lines 13-16 NOTE: Epstein teaches the obtaining both the foreground and background masks of two images in preparation for alpha compositing. Since these masks naturally segment the images, it would then be obvious to blend the first and second segmentations using the alpha-compositing technique. After the combination, the system for generating a first and second person as taught by Davis can generate their respective first and second segmentations. The alpha compositing technique of Epstein can then be used to create a composite image of the first and second segment. It would have been obvious to one of ordinary skill in the art before the effective filing date of the present invention to modify Davis by incorporating the teachings of Epstein to generate a second segmentation of the second person and combine this segmentation with the first segmentation of the first generated person. One would be motivated to make this combination in order to provide a way to combine multiple objects into one composite image. This would allow a user to take two photos of different people and create a high-quality multi-person image. The segmentations would benefit the blending of these images by preventing any alterations to the pixels that make up the objects . 07-21-aia AIA Claim (s) 7 and 11-12 is/are rejected under 35 U.S.C. 103 as being unpatentable over Davis, Zhang and Epstein and Wang . Regarding claim 7, Davis in view of Zhang teaches the method of claim 1. Davis does not teach obtaining foreground and background masks for each of the first and second artificial personalized images; blending depictions of foregrounds from the first and second artificial personalized images using the foreground masks; and inpainting a background of the blended depictions of the foregrounds from the first and second artificial personalized images using the background masks. However, Epstein teaches obtaining foreground and background masks for each of the first and second artificial personalized images; “To generate the background digital images and the foreground digital images, the digital image collaging system 102 generates background digital image masks and foreground digital image masks, respectively.”- Par 87, Lines 12-15 NOTE: Epstein teaches obtaining two images and generating the foreground and background masks for the images. Alpha compositing (also known as alpha blending) is performed to combine the foreground of one image and the background of another, see Fig 7 blending depictions of foregrounds from the first and second artificial personalized images using the foreground masks; “Additionally, the digital image collaging system 102 generates the collage digital images by collaging (e.g., alpha compositing) the background digital images and the foreground digital images together.” – Par 87, Lines 13-16 NOTE: Blending the generated foregrounds masks together would use the same logic disclosed by Epstein to combine the generated foreground and background masks. After the combination, the diffusion model for combining foreground and background images as taught by Zhang can be modified to use the foreground masks generated by the digital image collaging system taught by Epstein. This modification can then be added to Davis’s system for generating bodies using generative machine learning models. This would allow the foreground and background of the generated bodies from the first and second model to be extracted. Extracting the foreground of each generated person would then be combined into one composite image using the blending technique of Epstein. It would have been obvious to one of ordinary skill in the art before the effective filing date of the present invention to modify Davis by incorporating the teachings of Epstein to obtain foreground and background masks for the first and second personalized images and blend the foregrounds from each first and second personalized images using the foreground masks. One would have been motivated to make this combination to clearly separate the foreground and background area to reduce potential inconsistencies that may occur from the pixels related to the background when blending the foreground masks. Davis in view of Zhang and Epstein still does not teach inpainting a background of the blended depictions of the foregrounds from the first and second artificial personalized images using the background masks. However, Wang teaches inpainting a background of the blended depictions of the foregrounds from the first and second artificial personalized images using the background masks. “We focus on text-guided image inpainting, where a user provides an image, a masked area, and a text prompt and the model fills the masked area, consistent with both the prompt and the image context (Fig. 1)” – Wang Introduction Par 1 and Fig. 1 After the combination, the combined background masks generated from Epstein’s digital image collaging system can be used as input to Wang’s text-guided image inpainting model to inpaint a background based on the text prompt. This modification can be added to Davis’s system for generating a first and second body from a first and second generative model. This would allow for inpainting a background for the composite image of the two generated bodies based on the background mask. It would have been obvious to one of ordinary skill in the art before the effective filing date of the present invention to modify Davis by incorporating the teachings of Epstein to obtain foreground and background masks for each first and second personalized image, blend the foregrounds from the first and second artificial personalized images using the foreground masks, and inpainting a background of the blended foregrounds using the background masks. One would be motivated to make this combination since it is well known in the art that using masks to segment areas in an image gives the generative model context on where objects are in the image. By blending the foreground masks, a composite image can be created of the two generated models created from the first and second models to create a multi-person image. Using the background mask to perform inpainting of a background on this new multi-person image would prevent inpainting to alter the foreground accidentally. Regarding claim 11, Davis in view of Zhang and Epstein teach the method of claim 10. Davis does not teach inpainting the background on the foreground image based on the combined segmentation to generate the new artificial image. However, Wang teaches inpainting the background on the foreground image based on the combined segmentation to generate the new artificial image. “We focus on text-guided image inpainting, where a user provides an image, a masked area, and a text prompt and the model fills the masked area, consistent with both the prompt and the image context (Fig. 1)” – Wang Introduction Par 1 and Fig. 1 NOTE: After the combination, the new foreground image created from combined segmentations generated by Epstein can then be used as input to the inpainting method as disclosed by Wang. This would allow the user to input the image and a background mask that encompasses parts of the image that do not contain the foreground and then generate a background using a text prompt. This modification can then be added to Davis’s system for generating a first and second person with the first and second generative model. The combination would then be able to combine the segments that make up the generated person in each image and inpaint a background to generate a new artificial image. It would have been obvious to one of ordinary skill in the art before effective filing date of the present invention to modify Davis by incorporating the teachings of Wang to perform inpainting of a background on a foreground image based on the combine segmentation to generate a new artificial image. One would be motivated to make this combination because inpainting is a well-known technique in the art for filling in unwanted parts of an image by replacing it with visual attributes a user desires. This helps create visually coherent images and allows modification of unwanted regions without affecting the desired regions. Regarding claim 12, Davis, in view of Zhang, Epstein, and Wang teach the method of claim 11. Davis does not teach wherein the background is generated based on a prompt. However, Wang further teaches wherein the background is generated based on a prompt. “We focus on text-guided image inpainting, where a user provides an image, a masked area, and a text prompt and the model fills the masked area, consistent with both the prompt and the image context (Fig. 1)” – Wang Introduction Par 1 and Fig. 1 PNG media_image3.png 314 1056 media_image3.png Greyscale NOTE: Wang demonstrates an example of providing a text prompt associated to a masked region to modify the region. Fig. 1 shows an example of an image of a dog with a mask that encompasses a region of the background. A text prompt is used to generate a “rocket ship made of cardboard” which the generative AI creates within the mask. One of ordinary skill could also define the mask to encompass everything except the dog to modify the background completely. Wang et al, Fig 1 It would have been obvious to one of ordinary skill in the art before the effective filing date of the present invention to modify Davis by incorporating the teachings of Wang to generate the background based on a prompt. One would be motivated to make this combination to provide an intuitive way for the user to create a background with whatever visual attributes they desire . 07-21-aia AIA Claim (s) 13-14 is/are rejected under 35 U.S.C. 103 as being unpatentable over Davis, Zhang and Jiang et al (Text2Human: Text-Driven Controllable Human Image Generation), hereinafter Jiang . Regarding claim 13, Davis in view of Zhang teaches the method of claim 1. Davis does not teach receiving a first pose image comprising a depiction of an object in an individual pose; receiving a first prompt that defines a first set of visual attributes; and processing the first pose image and the first prompt by the first generative machine learning model to generate the first artificial personalized image. However, Jiang teaches receiving a first pose image comprising a depiction of an object in an individual pose; “To generate the human image, users are required to upload a human pose and texts describing the clothing shapes and textures.” – Pg 6, Fig. 4 Caption PNG media_image4.png 279 490 media_image4.png Greyscale Jiang et al Fig. 4 receiving a first prompt that defines a first set of visual attributes; “We synthesize full-body human images starting from a given human pose with two dedicated steps. 1) With some texts describing the shapes of clothes, the given human pose is first translated to a human parsing map. 2) The final human image is then generated by providing the system with more attributes about the textures of clothes” – Pg 1, Par 1 , Lines 7-12 NOTE: Jiang discloses that the user is able to input a prompt regarding visual attributes and textures of the clothes to be generated on a human. and processing the first pose image and the first prompt by the first generative machine learning model to generate the first artificial personalized image. “We synthesize full-body human images starting from a given human pose with two dedicated steps. 1) With some texts describing the shapes of clothes, the given human pose is first translated to a human parsing map. 2) The final human image is then generated by providing the system with more attributes about the textures of clothes” – Pg 1, Par 1 , Lines 7-12 NOTE: Jiang discloses that a final image is generated using the pose image input and user text description. The final image reflects the pose and clothes desired by the user. After the combination, receiving a pose image of the object and text prompt of visual attributes from the user as taught by Jiang can be added to Davis’s system for generating a first and second person. This modification would allow the generative model of Davis to use a pose image and text input to generate the first and second persons. It would have been obvious to one of ordinary skill in the art before the effective filing date of the present invention to modify Davis by incorporating the teachings of Jiang to receive a pose image depicting the first person, receive a prompt, and process the pose image and prompt to create the artificial image. One would have been motivated to make this combination to provide the user further control and customization for generating the artificial humans. Regarding claim 14, Davis in view of Zhang and Jiang teaches the method of claim 13. Davis does not teach receiving a second pose image comprising a depiction of another object in another pose; receiving a second prompt that defines a second set of visual attributes; and processing the second pose image and the second prompt by the second generative machine learning model to generate the second artificial personalized image. However, Davis in view of Zhang and Jiang teaches receiving a second pose image comprising a depiction of another object in another pose; receiving a second prompt that defines a second set of visual attributes; and processing the second pose image and the second prompt by the second generative machine learning model to generate the second artificial personalized image. There is a prima facie case of obviousness since the limitation is directed to common practices which the court has held normally require only ordinary skill in the art and hence are considered routine expedients are discussed below. See MPEP 2144.04. Duplication of Parts “the court held that mere duplication of parts has no patentable significance unless a new and unexpected result is produced”. Please see rejection of claim 13. It would have been obvious to one of ordinary skill in the art before the effective filing date of the present invention to modify Davis by incorporating the teachings of Jiang to repeat the same steps of claim 13 for a new instance to receive a second pose image depicting an object in a pose, receive a second prompt to define the second visual attributes, and process the second pose and second prompt to create a second artificial image. One would be motivated to make this combination to give the user control over the visual appearance of the second artificial person . 07-21-aia AIA Claim (s) 15 and 18 is/are rejected under 35 U.S.C. 103 as being unpatentable over Davis, Zhang, Ruiz et al (DreamBooth: Fine Tuning Text-to-Image Diffusion Models for Subject-Driven Generation), Li et al (BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation), and Junnan Li et al (US 20240161369 A1), hereinafter Ruiz, Li, and Junnan Li Regarding claim 15, Davis in view of Zhang teaches the method of claim 1. Davis does not teach training the first generative machine learning model by performing a first set of training operations comprising: accessing a first set of training images that each depict the first person with a different background and in different poses; receiving an image comprising a depiction of the first person; receiving a prompt that defines visual attributes of an individual background depicted in an individual training image of the first set of training images and that defines a target pose depicted in the individual training image; processing the image and the prompt by the first generative machine learning model to generate an estimated artificial image that depicts the first person in the target pose and an artificial background having the visual attributes defined by the prompt; computing a deviation between the estimated artificial image and the individual training image; and updating one or more parameters of the first generative machine learning model based on the deviation. However, Ruiz teaches training the first generative machine learning model by performing a first set of training operations comprising: accessing a first set of training images that each depict the first person with a different background and in different poses; receiving an image comprising a depiction of the first person; PNG media_image5.png 593 262 media_image5.png Greyscale PNG media_image6.png 635 175 media_image6.png Greyscale Fig 7 and Fig 8 as shown on Page 8 demonstrates example input images depicting objects in different poses/orientation and backgrounds. Ruiz et al, Fig. 7 (left) and Fig. 8 (right) NOTE: Ruiz discloses using input images of a subject in different poses and backgrounds to generate new artificial images based on a text prompt and the input images. This can be interpreted as accessing a first set of training images depicting a first person with a different background and poses. One of ordinary skill in the art could then simply choose one of the input images which functionally corresponds to receiving an image depicting the first person. After the combination, the collection of input images of the subject in various backgrounds and poses can modify Davis’s system for generating images of artificial bodies to allow the use of multiple source images of the first person in different backgrounds and poses as input to Davis’s synthesized human generation model. It would have been obvious to one of ordinary skill in the art before the effective filing date of the present invention to modify Davis by incorporating the teachings of Ruiz to access a first set of training images that each depict the first person with a different background and in different poses and receive an image comprising a depiction of the first person. One would be motivated to make this combination because it is well known in the art to be a standard procedure for training the generative machine learning model which will improve the model’s ability to create artificial images that closely resemble the subject (first person). Davis in view of Zhang and Ruiz still does not teach receiving a prompt that defines visual attributes of an individual background depicted in an individual training image of the first set of training images and that defines a target pose depicted in the individual training image; processing the image and the prompt by the first generative machine learning model to generate an estimated artificial image that depicts the first person in the target pose and an artificial background having the visual attributes defined by the prompt; computing a deviation between the estimated artificial image and the individual training image; and updating one or more parameters of the first generative machine learning model based on the deviation. However, Li teaches receiving a prompt that defines visual attributes of an individual background depicted in an individual training image of the first set of training images and that defines a target pose depicted in the individual training image; “Specifically, the captioner is an image-grounded text decoder. It is finetuned with the LM objective to decode texts given images. Given the web images Iw, the captioner generates synthetic captions Ts with one caption per image.” – Section 3.3 Pg 4, Par 2 NOTE: Li discloses a vision language pre-training framework for vision-language understanding and generation tasks, see Abstract. The architecture comprises “CapFilt” which comprises a captioner used to generate captions of images and a filter to determine if the captions best fit the image, see section 3.3. Fig 2 demonstrates an example of an input image of a girl holding a cat and the generated caption: “a little girl holding a kitten next to a blue fence”. This functionally corresponds to a prompt that defines visual attributes of an image’s background and pose of the subject in an individual training image. After the combination, the system for providing a prompt based on an individual image can be used to modify Davis’s system for generating images of human bodies based on multiple source images of the first person in different backgrounds and poses. PNG media_image7.png 513 1277 media_image7.png Greyscale Junnan Li et al, Fig 2 It would have been obvious to one of ordinary skill in the art before the effective filing date of the present invention to modify Davis by incorporating the teachings of Li to receive a prompt that defined visual attributes of an individual background and target pose depicted in an individual training image of the first set of training images. One would have been motivated to make this combination to strengthen the learning process of the generative model by supplying a text prompt of what the background and pose should look like. The received image of the first person from the training set will also strengthen the model’s ability to produce a high quality image in the style depicted in the image. Davis in view of Zhang, Ruiz, and Li still does not teach processing the image and the prompt by the first generative machine learning model to generate an estimated artificial image that depicts the first person in the target pose and an artificial background having the visual attributes defined by the prompt; computing a deviation between the estimated artificial image and the individual training image; and updating one or more parameters of the first generative machine learning model based on the deviation. However, Junnan Li teaches processing the image and the prompt by the first generative machine learning model to generate an estimated artificial image that depicts the first person in the target pose and an artificial background having the visual attributes defined by the prompt; “The subject-driven image generation model may be provided an input subject image 102 including a subject, ad subject text 112, and a text prompt 118 to generate an output image 124. ” – Par 34, Lines 1-3 NOTE: Junnan Li teaches processing an input image and text prompt to create an artificial image. After the combination, Davis’s system for generating images of artificial bodies is modified to generate a new artificial person by using the text prompt describing the visual attributes of the background and pose as taught by Li and the image depicting the first person as taught by Ruiz. The text prompt and received image can then be processed as taught by Junnan Li. computing a deviation between the estimated artificial image and the individual training image; and updating one or more parameters of the first generative machine learning model based on the deviation. “The output image may be compared to the ground-truth subject image (e.g., modified image 204) by loss computation 206. The loss computed by loss computation 206 may be used to update parameters of subject-driven image model 130 via backpropagation 208. ” – Par 34, Lines 4-8 NOTE: Junnan Li discloses computing the loss between the generated image and the ground truth. This functionally corresponds to computing the deviation between the estimated artificial image and the training image. The weights of the model are updated accordingly. It would have been obvious to one of ordinary skill before the effective filing date of the present invention to modify Davis by incorporating the teachings of Junnan Li to generate an estimated artificial image of the first person in the target pose and background defined by the prompt, compute the deviation, and update one or more parameters of the generative machine learning model. One would have been motivated to make this combination because it is a standard procedure for generative machine learning models. It is well-known in the art that the training process for generative machine learning models to compare the difference between the training image and the output of the model. This is known as a loss function which aims to optimize the model via supervised learning by updating parameters such as the weights of the neural network. Regarding claim 18, Davis in view of Zhang, Ruiz, Li, and Junnan Li teaches the method of claim 15. Davis does not teach training the second generative machine learning model by performing a second set of training operations comprising: accessing a second set of training images that each depict the second person with a different background and in different poses; receiving a second image comprising a depiction of the second person; receiving a second prompt that defines visual attributes of a second individual background depicted in a second individual training image of the second set of training images and that defines a second target pose depicted in the second individual training image; processing the second image and the second prompt by the second generative machine learning model to generate a second estimated artificial image that depicts the second person in the second target pose and a second artificial background having the visual attributes defined by the second prompt; computing a second deviation between the second estimated artificial image and the second individual training image; and updating one or more parameters of the second generative machine learning model based on the second deviation. However, Davis in view of Zhang, Ruiz, Li, and Junnan Li teaches training the second generative machine learning model by performing a second set of training operations comprising: accessing a second set of training images that each depict the second person with a different background and in different poses; receiving a second image comprising a depiction of the second person; receiving a second prompt that defines visual attributes of a second individual background depicted in a second individual training image of the second set of training images and that defines a second target pose depicted in the second individual training image; processing the second image and the second prompt by the second generative machine learning model to generate a second estimated artificial image that depicts the second person in the second target pose and a second artificial background having the visual attributes defined by the second prompt; computing a second deviation between the second estimated artificial image and the second individual training image; and updating one or more parameters of the second generative machine learning model based on the second deviation There is a prima facie case of obviousness since the limitation is directed to common practices which the court has held normally require only ordinary skill in the art and hence are considered routine expedients are discussed below. See MPEP 2144.04. Duplication of Parts “the court held that mere duplication of parts has no patentable significance unless a new and unexpected result is produced”. Please see rejection of claim 15. It would have been obvious to one of ordinary skill in the art before the effective filing date of the present invention to modify Davis by incorporating the teachings of Zhang, Ruiz, Li, and Junnan Li to perform the steps of claim 15 a second time for the images depicting the second person. One would be motivated to make this combination to strengthen the second generative machine learning model’s ability to artificially create people in the style similar to the training images. This would also prepare both the first and second models to create high quality artificial images to be used as input for a diffusion model configured to combine images of a first and second person to create a high quality composite image . 07-21-aia AIA Claim (s) 16-17 is/are rejected under 35 U.S.C. 103 as being unpatentable over Davis, Zhang, Ruiz, Li, Junnan Li, Epstein, and Wang . Regarding claim 16, Davis in view of Zhang, Ruiz, Li, and Junnan Li teaches the method of claim 15. Davis does not teach capturing a plurality of images of the first person; extracting a plurality of regions of the plurality of images that depicts the first person; obtaining a set of prompts that define different visual attributes of backgrounds; and processing the plurality of regions and the set of prompts by a diffusion model to generate the first set of training images. However, Ruiz teaches capturing a plurality of images of the first person; Fig 7 and Fig 8 as shown on Page 8 demonstrates example input images depicting objects in different poses/orientation and backgrounds NOTE: Multiple images of the subject in different backgrounds and poses are used in Ruiz. This shows that multiple images were captured to depict the first object. Capturing photos of the first person would be obvious to one of ordinary skill in the art to use as training images for generating artificial images of the first person. It would have been obvious to one of ordinary skill in the art before the effective filing date of the present invention to modify Davis in view of Ruiz to teach capturing a plurality of images of a first person. One would be motivated to make this combination because it is a well-known standard step to provide some ground-truth images that can be used to train the generative machine learning models. Davis in view of Zhang and Ruiz still does not teach extracting a plurality of regions of the plurality of images that depicts the first person; obtaining a set of prompts that define different visual attributes of backgrounds; and processing the plurality of regions and the set of prompts by a diffusion model to generate the first set of training images. However, Epstein teaches extracting a plurality of regions of the plurality of images that depicts the first person; “Along these lines, a “foreground digital image” refers to a digital image depicting pixels belonging to a foreground region or a foreground layer of a digital image” – Par 33, Lines 15-17 Epstein NOTE: Epstein teaches a digital image collaging system that finds the foreground and background masks of an image. The foreground can be understood as segment of the image that encapsulates the region of pixels depicting the foreground subject. After the combination, the plurality of images depicting a subject as taught by Ruiz can use the system of Epstein to extract a region depicting the foreground. This modification when combined with Davis’s system for generating artificial images of a first and second person will allow the extraction of the first person from the images captured of a first person in different poses and backgrounds. It would have been obvious to one of ordinary skill in the art before the effective filing date of the present invention to modify Davis by incorporating the teachings of Epstein to extract a region from an image depicting the first person for each captured image. One would be motivated to make this combination to clearly segment the subject of an image. This would allow the training process to determine what part of the image to focus on in order to improve at generating artificial images of the first person. Davis in view of Zhang, Ruiz, and Epstein still does not teach obtaining a set of prompts that define different visual attributes of backgrounds; and processing the plurality of regions and the set of prompts by a diffusion model to generate the first set of training images. However, Wang teaches obtaining a set of prompts that define different visual attributes of backgrounds; and processing the plurality of regions and the set of prompts by a diffusion model to generate the first set of training images. “We focus on text-guided image inpainting, where a user provides an image, a masked area, and a text prompt and the model fills the masked area, consistent with both the prompt and the image context (Fig. 1)” – Introduction Par 1 and Fig. 1 NOTE: Wang teaches a user providing text prompts that modify the image in the masked region that is also provided by the user as input to their diffusion model. One of ordinary skill could provide a mask for the background and modify the image with the text prompt. After the combination, the text prompts used to modify masked areas of the image as taught by Wang can use the image that was segmented into regions using a foreground mask as taught by Epstein as another input to Wang’s diffusion model to then artificially create a new image. This step can be repeated multiple times for each captured image of the first subject as taught by Ruiz. This modification can then be added to Davis to allow for the extraction of the first person in multiple source images which can then be used as input to the diffusion model of Wang to generate a new training images based on the source images and the background text prompt. It would have been obvious to one of ordinary skill before the effective filing date of the present invention to modify Davis by incorporating the teachings of Wang to obtain a set of prompts that define the visual attributes of different backgrounds and process the plurality of regions and set of prompts by a diffusion model to generate the first set of training images. One would be motivated to make this combination to create a varied training data set that encompasses many poses and backgrounds of the first person to improve the model’s performance and accurately generate images in the same likeness of the training images. Regarding claim 17, Davis in view of Zhang, Ruiz, Li, Junnan Li, Epstein, and Wang teaches the method of claim 16. Davis does not teach zooming in on the first person in the plurality of images to extract the plurality of regions. However, Epstein teaches zooming in on the first person in the plurality of images to extract the plurality of regions. “Along these lines, a “foreground digital image” refers to a digital image depicting pixels belonging to a foreground region or a foreground layer of a digital image” – Par 33, Lines 15-17 NOTE: Wang teaches extracting the foreground region by using a foreground mask that encompasses pixels that make up the subject of the foreground. Masking/segmentation functionally corresponds to zooming in on the first person since masking/segmenting clearly indicates the focus by isolating the region of interest. After the combination, the plurality of captured images of the first person as taught by Ruiz can use the methods of Epstein to extract the region of the first person in each image. This modification can then be added to Davis system for generating artificial images of a first and second person. The combination would allow for each region extracted by masking to functionally work as zooming in on the first person. It would have been obvious to one of ordinary skill before the effective filing date of the present invention to modify Davis by incorporating the teachings of Epstein to extract the plurality of regions by zooming in on the first person. One would be motivated to make this combination to isolate and focus on the subject (the first person) and reduce inconsistencies from background pixels prior to generating the new artificial images. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to DAVID V. NGUYEN whose telephone number is (571)272-6111. The examiner can normally be reached M-F 7:30-5:00. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, King Y Poon can be reached at 571-270-0728. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /DAVID VAN NGUYEN/Examiner, Art Unit 2617 /KING Y POON/ Supervisory Patent Examiner, Art Unit 2617 Application/Control Number: 18/947,989 Page 2 Art Unit: 2617 Application/Control Number: 18/947,989 Page 3 Art Unit: 2617 Application/Control Number: 18/947,989 Page 4 Art Unit: 2617 Application/Control Number: 18/947,989 Page 5 Art Unit: 2617 Application/Control Number: 18/947,989 Page 6 Art Unit: 2617 Application/Control Number: 18/947,989 Page 7 Art Unit: 2617 Application/Control Number: 18/947,989 Page 8 Art Unit: 2617 Application/Control Number: 18/947,989 Page 9 Art Unit: 2617 Application/Control Number: 18/947,989 Page 10 Art Unit: 2617 Application/Control Number: 18/947,989 Page 12 Art Unit: 2617 Application/Control Number: 18/947,989 Page 13 Art Unit: 2617 Application/Control Number: 18/947,989 Page 14 Art Unit: 2617 Application/Control Number: 18/947,989 Page 15 Art Unit: 2617 Application/Control Number: 18/947,989 Page 16 Art Unit: 2617 Application/Control Number: 18/947,989 Page 17 Art Unit: 2617 Application/Control Number: 18/947,989 Page 18 Art Unit: 2617 Application/Control Number: 18/947,989 Page 19 Art Unit: 2617 Application/Control Number: 18/947,989 Page 20 Art Unit: 2617 Application/Control Number: 18/947,989 Page 21 Art Unit: 2617 Application/Control Number: 18/947,989 Page 22 Art Unit: 2617 Application/Control Number: 18/947,989 Page 23 Art Unit: 2617 Application/Control Number: 18/947,989 Page 24 Art Unit: 2617 Application/Control Number: 18/947,989 Page 25 Art Unit: 2617 Application/Control Number: 18/947,989 Page 26 Art Unit: 2617 Application/Control Number: 18/947,989 Page 27 Art Unit: 2617 Application/Control Number: 18/947,989 Page 28 Art Unit: 2617 Application/Control Number: 18/947,989 Page 29 Art Unit: 2617 Application/Control Number: 18/947,989 Page 30 Art Unit: 2617 Application/Control Number: 18/947,989 Page 31 Art Unit: 2617 Application/Control Number: 18/947,989 Page 32 Art Unit: 2617
Read full office action

Prosecution Timeline

Nov 14, 2024
Application Filed
May 26, 2026
Non-Final Rejection mailed — §103
Aug 26, 2026
Response Filed
Sep 28, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12711586
STATIC DISTORTION CORRECTION FOR HEAD MOUNTED DISPLAY
2y 9m to grant Granted Aug 18, 2026
Patent 12705684
TIME SLICING
2y 0m to grant Granted Aug 11, 2026
Patent 12694574
ATTRIBUTE CODING AND UPSCALING FOR POINT CLOUD COMPRESSION
2y 3m to grant Granted Jul 28, 2026
Patent 12573160
INTIMACY-BASED MASKING OF THREE DIMENSIONAL (3D) FACE LANDMARKS
3y 3m to grant Granted Mar 10, 2026
Study what changed to get past this examiner. Based on 4 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
88%
Grant Probability
99%
With Interview (+14.6%)
2y 5m (~6m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 8 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month