Prosecution Insights
Last updated: October 02, 2026
Application No. 18/584,022

MASK-FREE COMPOSITE IMAGE GENERATION

Non-Final OA §102§103
Filed
Feb 22, 2024
Examiner
CONNER, SEAN M
Art Unit
2663
Tech Center
2600 — Communications
Assignee
Adobe Inc.
OA Round
1 (Non-Final)
79%
Grant Probability
Favorable
1-2
OA Rounds
0m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 79% — above average
79%
Career Allowance Rate
368 granted / 468 resolved
+16.6% vs TC avg
Strong +27% interview lift
Without
With
+26.7%
Interview Lift
resolved cases with interview
Typical timeline
2y 8m
Avg Prosecution
8 currently pending
Career history
483
Total Applications
across all art units

Statute-Specific Performance

§101
9.5%
-30.5% vs TC avg
§103
54.7%
+14.7% vs TC avg
§102
11.0%
-29.0% vs TC avg
§112
18.7%
-21.3% vs TC avg
Black line = Tech Center average estimate • Based on career data from 468 resolved cases

Office Action

§102 §103
DETAILED ACTION The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Applicant’s election without traverse of Invention I in the reply filed on 5/20/26 is acknowledged. The Amendment filed 5/20/26 has been entered and considered. Claims 10-16 have been canceled, and claims 21-27 have been added. Claims 1-9 and 17-27, all the claims pending in the application, are rejected. Claim Objections Claim 18 is objected to because of the following informalities: the claim recites “trained to identifying” which should be amended to recite “trained to identify”. Appropriate correction is required. Claim Rejections - 35 USC § 102 The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. Claims 1-3, 5-7, 9, 17, 19, 21-23, and 25-26 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by “ObjectStitch: Generative Object Compositing” by Song et al. (cited in the IDS filed 2/22/24; hereinafter “Song”). As to independent claim 1, Song discloses a method (Abstract discloses that Song is directed to “object compositing” using “diffusion models”) comprising: obtaining a first image depicting a background scene (Fig. 2: “input background”) and a second image depicting a foreground element (Fig. 2: “input subject”); generating, using an adapter network, a guidance embedding based on the second image (Section 3 and Fig. 2, reproduced below, disclose that the input subject image is encoded by a vision transformer ViT and adapted by “content adaptor” to produce a “descriptive embedding”); and generating, using an image generation model, a synthetic image depicting the foreground element and the background scene based on the first image and the guidance embedding (Section 3 and Fig. 2 disclose a “pretrained txt2img diffusion model” which inputs the input background image and the descriptive embedding and outputs a blended image depicting the input subject in the background scene), wherein the image generation model determines a location of the foreground element within the synthetic image in light of the background scene (Section 3 and Fig. 2 disclose a masked area M which provides a “soft constraint” of the location of the input subject in the background scene, wherein the mask “not only fully covers the object, but also extends to its neighboring area (providing room for shadow generation)” and “is flexible enough for the model to apply spatial transformations, synthesize novel views and generate shadows and reflection”; emphasis added; the mask provides a general location for the foreground object, but the model itself determines the specific location within the background scene; see Fig. 6 and 10-11). PNG media_image1.png 264 1008 media_image1.png Greyscale As to claim 2, Song further discloses that generating the guidance embedding comprises: encoding, using an image encoder, the second image to obtain an image embedding, wherein the guidance embedding is generated based on the image embedding (Section 3 and Figs. 2-3 disclose vision transformer ViT which is a “CLIP image encoder” that outputs an “image embedding” to the content adaptor that generates the adaptive embedding). As to claim 3, Song further discloses that generating the synthetic image comprises: combining the first image with a noise map to obtain a noisy image, wherein the image generation model takes the noisy image as input (Section 3 and Fig. 2 discloses that the input to the U-Net Generator of the txt2img diffusion model is adjusted to “contain the original background image outside the hole and noise inside the hole”). As to claim 5, Song further discloses that generating the synthetic image comprises: performing a reverse diffusion process (Sections 2-3 and Fig. 2 discloses that the blended image is generated by a “diffusion model” which include a “forward process” and a “reverse process”). As to claim 6, Song further discloses obtaining location information for the foreground element of the second image, wherein a location of the foreground element in the synthetic image is determined based on the location information (Section 3 and Fig. 2 disclose a masked area M which provides a “soft constraint” of the location of the input subject in the background scene for the final blended image). As to claim 7, Song further discloses that a size of the foreground element is determined by the location information (Section 3 and Fig. 2 disclose a masked area M which provides a “soft constraint of the location and scale of the composited object”; emphasis added). As to claim 9, Song further discloses that the image generation model is trained to combine multiple images using training data that includes a foreground image, a background image, and a ground truth image that combines the foreground image and the background image (Section 3 discloses a “training data generation” process in which the training images are formed by extracting the object (foreground) from training images using a segmentation mask to create the background image and input image for training, and the original image (including both the background and the foreground) is used as the corresponding “ground truth”). As to independent claim 17, Song discloses an apparatus (Abstract discloses that Song is directed to “object compositing” using “diffusion models” which are necessarily implemented by an apparatus such as a computer) comprising: at least one processor; at least one memory storing instructions executable by the at least one processor (Section 4 discloses that the algorithm is implemented on “GPUs” necessarily running computer code stored in memory); an image encoder comprising parameters stored in the at least one memory and trained to encode a foreground image to obtain an image embedding (Section 3 and Figs. 2-3 disclose a foreground “input subject” image which is input to a vision transformer ViT “CLIP image encoder” which generates an “image embedding”, wherein the vision transformer includes a neural network necessarily having parameters stored in memory); an adapter network comprising parameters stored in the at least one memory and trained to generate a guidance embedding based on the image embedding (Section 3 and Figs. 2-3 disclose that the “image embedding” is input to a “Content adaptor” which outputs a “descriptive” or “adaptive embedding”, wherein the content adaptor is shown to include an “MLP” which is “train[ed]” and necessarily has parameters stored in memory; see Section 3.3.2); and an image generation model comprising parameters stored in the at least one memory and trained to generate a synthetic image based on a background image and the guidance embedding (Section 3 and Fig. 2 disclose a “pretrained txt2img diffusion model” which inputs the input background image and the descriptive embedding and outputs a blended image depicting the input subject in the background scene, wherein the pretrained model includes trained parameters necessarily stored in memory), wherein the synthetic image depicts a foreground element and a first portion of the background scene (Fig. 2 shows that the final “blended image” includes the foreground object from the “input subject” image within the background scene depicted in the “input background”), and wherein the foreground element is located at a position of the synthetic image that is determined by the image generation model (Section 3 and Fig. 2 disclose a masked area M which provides a “soft constraint” of the location of the input subject in the background scene, wherein the mask “not only fully covers the object, but also extends to its neighboring area (providing room for shadow generation)” and “is flexible enough for the model to apply spatial transformations, synthesize novel views and generate shadows and reflection”; emphasis added; the mask provides a general location for the foreground object, but the model itself determines the specific location within the background scene; see Fig. 6 and 10-11). As to claim 19, Song further discloses a segmentation component comprising parameters stored in the at least one memory and trained to perform object segmentation on an image to generate a segmentation mask (Section 3 discloses that synthetic training data uses an “object instance segmentation model” CenterMask which performs “panoptic segmentation” and generates a “segmentation mask”). Independent claim 21 recites a non-transitory computer readable medium storing code for image processing, the code comprising instructions that, when executed by at least one processor, cause the at least one processor to perform operations (Section 4 discloses that the algorithm is implemented on “GPUs” necessarily running computer code stored in memory) comprising the steps recited in the method of independent claim 1. Accordingly, claim 21 is rejected for reasons analogous to those discussed above in conjunction with claim 1. Claims 22-23 and 25-26 recite features nearly identical to those recited in claims 2-3 and 6-7, respectively. Accordingly, claims 22-23 and 25-26 are rejected for reasons analogous to those discussed above in conjunction with claims 2-3 and 6-7, respectively. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention. Claims 4 and 24 are rejected under 35 U.S.C. 103 as being unpatentable over Song in view of “ControlCom: Controllable Image Composition using Diffusion Model” by Zhang et al. (hereinafter “Zhang”). As to claim 4, Song does not expressly disclose that the noise map having a same resolution as the first image. Zhang, like Song, is directed to “image composition” by “synthesizing a realistic composite image from a pair of foreground and background images” using a “diffusion model” (Abstract and Fig. 1). Zhang discloses that the diffusion model inputs 11 channels including the background image and noisy latent code zt, the inputs having the same resolution of 64 x 64 (Section 4 and Fig. 2). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Song’s diffusion model to input a combination of the background image and a noise map having the same resolution, as taught by Zhang, to arrive at the claimed invention discussed above. Such a modification is the result of combining prior art elements according to known methods to yield predictable results. It is predictable that the proposed modification would have facilitated the iterative denoising process of the diffusion model from zt to z0, the finally obtained composite image (Section 4, paragraph 1 of Zhang). Claim 24 recites features nearly identical to those recited in claim 4. Accordingly, claim 24 is rejected for reasons analogous to those discussed above in conjunction with claim 4. Claims 8 and 27 are rejected under 35 U.S.C. 103 as being unpatentable over Song in view of U.S. Patent Application Publication No. 2023/0325996 to Zhang et al. (cited in IDS filed 8/21/25; hereinafter “Zhang2”). As to claim 8, Song does not expressly disclose obtaining an additional image depicting an additional foreground element, wherein the synthetic image is generated to depict the additional foreground element. Zhang2, like Song, is directed to a trained compositing model for combining a foreground image and background image (Abstract and [0058]). Zhang2 discloses the capability of generating a first composite image 2002 from a background image 2004 and a first foreground object image 2006, then generating a second composite image 2008 using the first composite image 2002 and a second foreground object image 2010, then generating a third composite image 2014 using the second composite image 2008 and a third foreground object image 2016 ([0241 and Figs. 20A-C, reproduced below). PNG media_image2.png 404 652 media_image2.png Greyscale PNG media_image3.png 406 660 media_image3.png Greyscale PNG media_image4.png 406 662 media_image4.png Greyscale It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Song to be implemented in a graphic user interface that allows for additional foreground elements to be added to the original blended image, as taught by Zhang2, to arrive at the claimed invention discussed above. Such a modification is the result of combining prior art elements according to known methods to yield predictable results. It is predictable that the proposed modification would have improved user experience by virtue of providing a more flexible tool. Claim 27 recites features nearly identical to those recited in claim 8. Accordingly, claim 27 is rejected for reasons analogous to those discussed above in conjunction with claim 8. Claim 18 is rejected under 35 U.S.C. 103 as being unpatentable over Song in view of “Beyond Generation: Harnessing Text to Image Models for Object Detection and Segmentation” by Ge et al. (hereinafter “Ge”). As to claim 18, Song does not expressly disclose an object detection component comprising parameters stored in the at least one memory and trained to identifying the foreground element from an image. Ge, like Song, is directed to image synthesis using diffusion models (Abstract). In particular, Ge is directed to “foreground-background segmentation” to generate the foreground object masks used in such models (Abstract). Ge discloses that this segmentation is performed using “detectors” which are “trained” and thus include parameters necessarily stored in memory (Abstract and Section 4.1). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Song to perform the object segmentation for generating training images using a trained object detection network, as taught by Ge, to arrive at the claimed invention discussed above. Such a modification is the result of combining prior art elements according to known methods to yield predictable results. It is predictable that the proposed modification would have improved the generation of the training data. Claim 20 is rejected under 35 U.S.C. 103 as being unpatentable over Song in view of “Making Images Real Again: A Comprehensive Survey on Deep Image Composition” by Niu et al. (hereinafter “Niu”). As to claim 20, Song does not expressly disclose an inpainting component comprising parameters stored in the at least one memory and trained to generate the background image based on the segmentation mask. Niu, like Song, is directed to “image composition” networks which “combine the foreground from one image and another background image” (Abstract). Niu discloses that training images for an image composition network such as Song’s can involve “removing the foreground objects” based on “segmentation masks”, then restoring the background images to “complete background images by using image inpainting techniques” so as to generate “triplets of foregrounds, backgrounds, and ground-truth composite images” for training, similar to the triplets generated by Song for training (Section II(C) of Niu and Section 3.3.1 of Song). Notably, the references cited by Niu for the inpainting techniques include “Image Inpainting with Deep Generative Models” (reference [170] to Yeh et al.), wherein deep models are necessarily trained to learn parameters stored in memory. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Song to use the segmentation masks to extract the foreground object an inpainting deep model to restore the background image so that the two images can be used in training the deep foreground-background image composition network, as taught by Niu, to arrive at the claimed invention discussed above. Such a modification is the result of combining prior art elements according to known methods to yield predictable results. It is predictable that the proposed modification would have improved the quality of the training data, thus resulting an improved model. Pertinent Art The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. “Semantic Image Inpainting with Deep Generative Models” by Yeh et al. discloses deep inpainting models, as referenced by Niu above. U.S. Patent Application Publication No. 20250014148 to Cho et al. discloses a deep learning framework for background-foreground composition in which a first neural network is trained to estimate the position and size of an object and a second neural network places the foreground object into the background scene based on the estimated position and size. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to SEAN M CONNER whose telephone number is (571)272-1486. The examiner can normally be reached 10 AM - 6 PM Monday through Friday, and some Saturday afternoons. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Greg Morse can be reached at (571) 272-3838. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /SEAN M CONNER/Primary Examiner, Art Unit 2663
Read full office action

Prosecution Timeline

Feb 22, 2024
Application Filed
Sep 23, 2026
Non-Final Rejection mailed — §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12725406
TASK-ORIENTED CLUSTERING USING PROMPT LEARNING
3y 2m to grant Granted Sep 01, 2026
Patent 12718532
IMAGE CLASSIFICATION METHOD AND APPARATUS, AND METHOD AND APPARATUS FOR IMPROVING TRAINING OF AN IMAGE CLASSIFIER
4y 4m to grant Granted Aug 25, 2026
Patent 12705759
SELECTIVELY IDENTIFYING DATA BASED ON MOTION DATA FROM A DIGITAL VIDEO TO PROVIDE AS INPUT TO AN IMAGE PROCESSING MODEL
3y 8m to grant Granted Aug 11, 2026
Patent 12701275
Methods and Systems for Extracting Sport-Related Information from Digital Video Frames
2y 10m to grant Granted Aug 04, 2026
Patent 12688646
MEASUREMENT METHOD, MEASUREMENT DEVICE, AND RECORDING MEDIUM
4y 9m to grant Granted Jul 21, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
79%
Grant Probability
99%
With Interview (+26.7%)
2y 8m (~0m remaining)
Median Time to Grant
Low
PTA Risk
Based on 468 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month