Prosecution Insights
Last updated: October 01, 2026
Application No. 18/957,772

SUBJECT DRIVEN IMAGE EDITING

Non-Final OA §103
Filed
Nov 24, 2024
Examiner
GE, JIN
Art Unit
2619
Tech Center
2600 — Communications
Assignee
Adobe Inc.
OA Round
1 (Non-Final)
80%
Grant Probability
Favorable
1-2
OA Rounds
8m
Est. Remaining
98%
With Interview

Examiner Intelligence

Grants 80% — above average
80%
Career Allowance Rate
440 granted / 552 resolved
+17.7% vs TC avg
Strong +19% interview lift
Without
With
+18.8%
Interview Lift
resolved cases with interview
Typical timeline
2y 6m
Avg Prosecution
26 currently pending
Career history
572
Total Applications
across all art units

Statute-Specific Performance

§101
10.6%
-29.4% vs TC avg
§103
62.0%
+22.0% vs TC avg
§102
11.0%
-29.0% vs TC avg
§112
9.5%
-30.5% vs TC avg
Black line = Tech Center average estimate • Based on career data from 552 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Information Disclosure Statement The information disclosure statements (IDS) submitted on 01/20/2026 and 11/24/2024 are in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner. Election/Restrictions Applicant's election with traverse of claims 1-9 and 16-20 in the reply filed on 06/29/2026 is acknowledged. This is not found persuasive because the identified group, as described in the previously mailed Election/Restrictions Requirement, belong to distinctly different inventions in the art ((claim different subject matter) even though they may be have some common points. Particularly, Group I: Claims 1-9 and 16-20 are directed towards generating, using an image generation model, a synthetic image based on the concept feature through performing a style transform using a concept input, a source image, and an input mask. Group II: Claims 9-15 are directed towards generating, using an image generation model, a synthetic image based on the concept feature and a background features using a concept input and a source image, wherein 1. The concept features are generated from different source (a: concept features are generated from input mask, b:concept features are generated from source image, 2. Even consider claim 3, the background features are generated from different source (a: background features are generated from preliminary background features and mask, b: background features are generated from the source image). The species are independent or distinct because claims to the different species recite mutually exclusive characteristics because recites devices for accomplishing two different tasks (directed towards to generating, using an image generation model, a synthetic image based on the concept feature through performing a style transform using a concept input, a source image, and an input mask vs generating, using an image generation model, a synthetic image based on the concept feature and a background features using a concept input and a source image). In addition, these species are not obvious variants of each other based on the current record. There is a search and/or examination burden for the patentably distinct species as set forth above because at least the following reason(s) apply: the species or groupings of patentably indistinct species require a different field of search (e.g., searching different classes/subclasses or electronic resources, or employing different search strategies or search queries). So the restriction is proper. Claim Interpretation The following is a quotation of 35 U.S.C. 112(f): (f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph: An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked. As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph: (A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function; (B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and (C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function. Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function. Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function. Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitation(s) is/are: “the image generation model comprises a location adaptation module, a style adaptation module including an instance normalization component, a scale adaptation module, and a content adaptation module” in claim 18. Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof. If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1-5, 16, and 19-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over U.S. PGPubs 2025/0349054 to Akerlund et al. in view of U.S. PGPubs 2021/0110588 to Adamson, III. Regarding claim 1, Akerlund et al. teach a method comprising (par 0012): obtaining a concept input (Fig 2A, par 0091-0093, “the user can submit, simultaneously, the image 210A and the user input 210B requesting one or more image edits to the image 210A, to the client device 20”), a source image (Fig 2A, par 0090-0092, “ The user can provide a user input 210B requesting one or more image edits to the image 210A”), and an input mask (Fig 2C, par 0094-0096, “referring to FIG. 2C, the automatically generated image mask (e.g., 230) can be rendered to the user via the user interface 200 of the chat application to receive user confirmation. For instance, a prompt 210C (e.g., “Confirm image mask outlined by the dashed line?”) can be rendered to seek user input to confirm (or modify) the position of the automatically generated image mask”), wherein the concept input represents a concept (Fig 2A, par 0091-0093, “the user can submit, simultaneously, the image 210A and the user input 210B requesting one or more image edits to the image 210A, to the client device 20 ….. the user input 210B can be: “change dog to cat”.”), the source image depicts a scene (Fig 2A, par 0090-0092, “ The user can provide a user input 210B requesting one or more image edits to the image 210A”), and the input mask indicates a location for the concept in the scene (Fig 2C, par 0094-0096, “an image mask can be generated based on the location description that indicates the location for the source object (e.g., dog) present in the image 210A. The image mask can, for instance, mask the source object that needs to be edited. Alternatively, the image mask can, for instance, mask all areas of the image 210A except for the source object, where content of the masked area(s) of the image 210A can be preserved (e.g., not modified) during processing of the image 210A to generate an edited image (e.g., 220B in FIG. 2B) that has a background visually the same as, or similar to, the image 210A (e.g., also referred to as “source image”). ….referring to FIG. 2C, the automatically generated image mask (e.g., 230) can be rendered to the user via the user interface 200 of the chat application to receive user confirmation. For instance, a prompt 210C (e.g., “Confirm image mask outlined by the dashed line?”) can be rendered to seek user input to confirm (or modify) the position of the automatically generated image mask”); generating concept features from the source image to the concept input based on the input mask (par 0014-0015, “The one or more image editing instructions can, additionally or alternatively, include or indicate an edit to the source image. The edit to the source image can be derived from the user request to edit the image. For example, the user request to edit the image can be a request to replace a source object in the image to be edited with a target object (e.g., “replace the dog with a white cat”). In this example, the edit to the source image (in the one or more image editing instructions) can be, for instance, “generate a white cat at a position of the region that is masked, and don't change other image content from the original image that is outside of the region that is masked”. As another example, the user request to edit the image can be a request to modify a characteristic (e.g., color, size, location, etc.) of a source object present in the source image, such as, “replace the color of the dog from black to white”. In this example, the edit to the source image can be, for instance, “change the color of the dog within the region that is masked from black to white”)”, par 0050-0057, “the image understanding engine 142 can process a source image, and/or a user query (or instead of the user query, a first text prompt derived from the user query), using the visual language model and/or the object detection & classification model, to generate a text representation of the source image. For instance, the source image can be a painting (or a photo uploaded by a user from an electronic album) showing a white butterfly sitting on top of a native pink milkweed. The user query can include or indicate one or more edits to the source image. In some implementations, the user query can, but does not necessarily need to, identify one or more source objects in the source image to be edited or modified. In some implementations, additionally or alternatively, the user query can include or indicate a target object to be generated in the edited image based on modifying or replacing one of the one or more source objects (or other image content) in the source image. In some implementations, additionally or alternatively, the user query can include or indicate a modification or edit to a property (e.g., color, size, shape, etc.) of a source object/image content in the source image. Descriptions of the user query, however, are not limited herein. In some implementations, the aforementioned first text prompt (to be processed, along with the source image, by the image understanding engine 142) can include, for instance, a first instruction to identify a location of source object(s) in the source image based on the user query ….the image understanding engine 142 can process the source image and/or the user query, to generate a text representation of the source image. The text representation of the source image can, for instance, include a description of image content of the source image (e.g., a white butterfly sitting on top of a native pink milkweed, or a more detailed description). The text representation of the source image can further indicate, for instance, a position (e.g., positions of pixels) of a source object (which may be identified based on the user query) to be edited (e.g., modified to have a different property such as color, or replaced with a target object) ….In response to determining that a particular object (e.g., white butterfly) is determined to belong to the same category as the target object (e.g., monarch butterfly), the image understanding engine 142 can determine the particular object as the source object to be edited in the source image and/or determine locations or pixels corresponding to the particular object. ….. the image mask can alternatively mask image content to be preserved or retained. The one or more image editing instructions can, additionally or alternatively, indicate the target object (e.g., monarch butterfly) to replace the source object (e.g., white butterfly), or a property (e.g., color, shape, etc.) of the source object in the source image to be edited or modified”); and generating, using an image generation model (par 0004-0005, “an image generation model”), a synthetic image based on the concept features, wherein the synthetic image depicts the concept from the concept input within the scene from the source image at the location indicated by the input mask (par 0006, “some of the various implementations do not require a user to specify a region of the source image that is to be edited, for target content (e.g., target object) to replace original content (e.g., original object) within the specified region”, par 0057, “the image mask can alternatively mask image content to be preserved or retained. The one or more image editing instructions can, additionally or alternatively, indicate the target object (e.g., monarch butterfly) to replace the source object (e.g., white butterfly), or a property (e.g., color, shape, etc.) of the source object in the source image to be edited or modified.”, par 0092, par 0099-0100, “The second prompt can be processed, for instance, a large language model (“LLM”) configured for text generation, to generate the one or more image editing instructions for image 210A. Continuing with the non-limiting example above, the second prompt can be, for instance, “generate image editing instruction(s) for the image 210A, considering the user input 210B”, or “generate image editing instruction(s) for the image 210A, considering the user input 210B. also generate a text reply to the user input 210B”. In this example, the one or more image editing instructions can be, for instance, “given the image 210A and using the image mask masking all regions of the image 210A except for the source object, replace the source object in the image 210A with the target object”. The text reply (e.g., 220A in FIG. 2B) can be, for instance, “check out the edited image below”, “Here you are, a cat riding a bike”, etc.” In some implementations, the one or more image editing instructions and the image 210A can be processed as input, using a second machine learning (ML) model, to generate an edited image (e.g., 220B in FIG. 2B). The second ML model can be, for instance, an image-generation model trained to generate image(s) based on text descriptions. Continuing with the non-limiting example above, the edited image 220B can have same background as the source image 210A, except for the source object of “dog” being replaced with the target object of “cat”. In other words, the edited image 220B can show a cat riding a bike in a country road, with a background of mountains and/or clouds, where object(s) (e.g., clouds, mountains, bike, country road) in the edited image 220B that are not the target object (e.g., cat) are the same as those in the source image 210A.). But Akerlund et al. keeps silent for teaching generating concept features by performing a style transfer from the source image to the concept input based on the input mask. In related endeavor, Adamson, III teaches generating concept features by performing a style transfer from the source image to the concept input based on the input mask (par 0036, “The way we do this is we apply image segmentation to the image to mark the presence of a known object type in the image (e.g., via a pixel-wise mask generated for each image and that can be combined with the color pixel image and depth pixel image via a bit-wise operation and the combination depicted) and to label the image segment with an indication of a known class type.”, par 0043, “Image segmentation 120 is configured, generally, to identify and segment an object present in a source image 130 from a background of the source image 130 to obtain image segment 132 of such object and background. In some instances, an image segment 132 of an object may include a source image 130 and an image segment mask (e.g., a pixel-wise mask, without limitation) that marks an object (e.g., marks a location of an outline or area of an object in source image 130) and attaches a label that categorically classifies the object”, par 0045-0046, “Style effects 124 may be configured, generally, to apply style transfer effects to a segmented object. Style effects 124 may include one or more machine learning models for style transfer or style transformation. In the case of style transfer, a machine learning model has been trained to discern and apply the style of a style image to the content of a content image—stated another way, change expressive elements of a segmented content image to more closely match expressive elements of a segmented style image. In the case of style transformation, a machine learning model has been trained to change expressive elements of a segmented content image, typically of a specific recognized categorical class (e.g., human, building, animal, automobile, plant, without limitation). In the case of style transformation, the machine learning model is trained using e.g., a photo and a painting or graphical representation of the photo having a desired style”, par 0066-0067, “Object segmenters 208 may be configured to generate object segment mask 218 of image segments 220 and object isolators 222 may be configured to generate isolated objects masks 224 from segmented objects, e.g., generated by applying object segment mask 218 to source image 202 …. to segment the object, those instances may be isolated by Deep Learning models. So here we would be applying both together. In a contemplated operation, isolate just the face of the segmented portrait of a person and apply style transfer to just the isolated face”, par 0101, “Style transfers may be used to transfer a style from an original image of a scene to the doctored image in an attempt to obfuscate indicators (e.g., irregularities due to differences in the types of image capture or generation devices, lighting, or something else, without limitation) that might visibly tip off a viewer that an image was edited”). It would have been obvious to a person of ordinary skill in the art at the time before the effective filing data of the claimed invention to modified Akerlund et al. to include generating concept features by performing a style transfer from the source image to the concept input based on the input mask as taught by Adamson, III to deploying machine learning models for object classification and style transfer modules of a mobile application that were trained using image segments to perform image synthesis. Regarding claim 2, Akerlund et al. as modified by Adamson, III teach all the limitation of claim 1, and Akerlund et al. further teach wherein obtaining the input mask comprises: obtaining a text prompt describing an element of the source image (par 0003, “changing a text prompt from “photo of yellow dog riding on a bicycle” to “photo of white dog riding on a bicycle” can result in a completely different generated image, such as one that changes the dog's shape, which can be undesired to a user”, par 0014, “ the user request to edit the image can be a request to replace a source object in the image to be edited with a target object (e.g., “replace the dog with a white cat”). In this example, the edit to the source image (in the one or more image editing instructions) can be, for instance, “generate a white cat at a position of the region that is masked, and don't change other image content from the original image that is outside of the region that is masked””); and generating the input mask based on the source image and the text prompt, wherein the input mask is based on a region of the element described by the text prompt (par 0012, “The one or more image editing instructions can be, but does not necessarily need to be, in the form of a text prompt processable using an image generation model, where the text prompt can be, for instance, “using the source image to generate an edited image by changing image content within the bounding box that is generated for the source image and that has location information of [ . . . ] with a white cat, preserve image content outside the bounding box” “, par 0072, “the first text prompt can be, for instance, “generate a description of the provided image” or “generate a description of the provided image and identify positions of all objects recognized in the provided image”. In some implementations, the first text prompt can be derived based on the user query 151A. For example, given the user query 151A being “change cat to dog”, the first text prompt can be, for instance, “generate a description of the image, add a bounding box for cat in the image if there is any”.”, Fig 2C, par 0094, “Techniques described in the present disclosure allows automatic generation of such image mask, without a user manually defining a mask area of the image mask, which can be time-consuming and requires accurate definition of the mask area from a human user. Optionally, referring to FIG. 2C, the automatically generated image mask (e.g., 230) can be rendered to the user via the user interface 200 of the chat application to receive user confirmation. For instance, a prompt 210C (e.g., “Confirm image mask outlined by the dashed line?”) can be rendered to seek user input to confirm (or modify) the position of the automatically generated image mask. Generation of the edited image utilizing the automatically generated image mask can be performed in response to receiving user confirmation of the automatically generated image mask 230”). Regarding claim 3, Akerlund et al. as modified by Adamson, III teach all the limitation of claim 1, and Akerlund et al. further teach further comprising: generating preliminary background features based on the source image; and generating background features based on the preliminary background features and the input mask, wherein the synthetic image is generated based on the background features (par 0013, “the user request to edit the image can be: “change the background of the image from beach to grass”. In the example, the region of the image to be masked can be a region to be preserved and can be indicated, for instance, using a bounding box surrounding a target object (e.g., a tourist), where image content (e.g., the tourist) within the bounding box is to be preserved and image content (e.g., the background, e.g., beach) outside the bounding box is to be edited. As another example, the user request to edit the image can be: “add a rabbit in the grass”. In this example, the region of the image to be masked can be a region to be edited and can be indicated, for instance, using a bounding box surrounding a portion of the grass that is to be placed with a rabbit, where image content outside of the bounding box is preserved for inclusion in the edited image“, par 0052, “the user query can also be, for instance, “change to a monarch butterfly” which does not explicitly identify the source object (e.g., “white butterfly”) to be edited in the source image, but identifies a target object to be introduced into the source image, to generate an edited image that shares certain image content (e.g., an image background such as the pink milkweed plant) with the source image and that includes additional image content generated based on the user query“, par 0094, “the image mask can, for instance, mask all areas of the image 210A except for the source object, where content of the masked area(s) of the image 210A can be preserved (e.g., not modified) during processing of the image 210A to generate an edited image (e.g., 220B in FIG. 2B) that has a background visually the same as, or similar to, the image 210A (e.g., also referred to as “source image”, par 0098, “he text representation for the image 210A can be a description of all objects present in the image 210A, such as “The image depicts a dog riding a bike in a country road, with clouds and mountains in the background”. The text representation for the image 210A can also include location information of the objects (all or only the source object) present in the image 210A, such as “The image depicts a dog riding a bike in a country road, with clouds and mountains in the back, a location of the dog is indicated by . . . ”. “, par 0100, “The second ML model can be, for instance, an image-generation model trained to generate image(s) based on text descriptions. Continuing with the non-limiting example above, the edited image 220B can have same background as the source image 210A, except for the source object of “dog” being replaced with the target object of “cat”. In other words, the edited image 220B can show a cat riding a bike in a country road, with a background of mountains and/or clouds, where object(s) (e.g., clouds, mountains, bike, country road) in the edited image 220B that are not the target object (e.g., cat) are the same as those in the source image 210A”, Adamson, III: par 0043, “Image segmentation 120 is configured, generally, to identify and segment an object present in a source image 130 from a background of the source image 130 to obtain image segment 132 of such object and background. In some instances, an image segment 132 of an object may include a source image 130 and an image segment mask (e.g., a pixel-wise mask, without limitation) that marks an object (e.g., marks a location of an outline or area of an object in source image 130) and attaches a label that categorically classifies the object. Such a segment mask may be applied to the source image 130 change values of pixels in the source image 130 that are not associated with the object thereby defining the object in the resultant image (e.g., an edited source image or a copy). Similarly, image segment 132 of a background may include an image segment mask that marks a background (e.g., marks a location of an outline of a background in source image 130)”, par 0047, “Image synthetization 126 may be configured, generally, to combine one or more of a styled image segment 134, a background image, an overlay image, and display information (e.g., text, a stamp, a logo, or an icon, without limitation) into a synthesized image 136. In a contemplated use, a background present in a source image 130 may be “replaced” with an image of a new background (e.g., using convolutional techniques for combining images, without limitation). Overlays and display information may be combined after or before style transfer effects are applied”, par 0096, “FIG. 9 is a diagram depicting a process flow of image data obtained by performing image segmentation, classification, style transfer and background modification in accordance with one or more embodiments. A portrait object 902 and background 904 are present in the image segmented. A classification label is applied to the recognized image. In the third frame, style effects have been applied to the portrait object 902 to generate styled portrait 908, based on the classification label 906. In the fourth frame, background 904 has been replaced with new background 910. In the fifth frame, alpha has been adjusted on the styled portrait 908 so adjusted portrait 912 appears somewhat translucent in the synthesized image”). Regarding claim 4, Akerlund et al. as modified by Adamson, III teach all the limitation of claim 3, and Akerlund et al. further teach further comprising: combining the concept features and the background features to obtain target features, wherein the synthetic image is generated based on the target features (par 0094, “the image mask can, for instance, mask all areas of the image 210A except for the source object, where content of the masked area(s) of the image 210A can be preserved (e.g., not modified) during processing of the image 210A to generate an edited image (e.g., 220B in FIG. 2B) that has a background visually the same as, or similar to, the image 210A (e.g., also referred to as “source image”, par 0100, “The second ML model can be, for instance, an image-generation model trained to generate image(s) based on text descriptions. Continuing with the non-limiting example above, the edited image 220B can have same background as the source image 210A, except for the source object of “dog” being replaced with the target object of “cat”. In other words, the edited image 220B can show a cat riding a bike in a country road, with a background of mountains and/or clouds, where object(s) (e.g., clouds, mountains, bike, country road) in the edited image 220B that are not the target object (e.g., cat) are the same as those in the source image 210A”). Regarding claim 5, Akerlund et al. as modified by Adamson, III teach all the limitation of claim 1, and teach wherein generating the concept features comprises: generating preliminary concept features based on the concept input; generating preliminary background features based on the source image; performing the style transfer based on the preliminary concept features, the preliminary background features, and the input mask to obtain refined preliminary concept features; and generating the concept features based on the refined preliminary concept features and the input mask (Akerlund et al.: par 0016, “The particular image generation model can be specified in the user request to edit the image (i.e., “the source image”), or can be determined based on the user request to edit the image. For instance, based on the user request to edit the image being a request to modify a style of the source image, the one or more image editing instructions can include a model selection instruction that specifies an image generation model trained or fine-tuned to perform image style transfer and/or an address of such image generation model for image style transfer. As another example, based on the user request to edit the image being a request to remove a source object from the source image, the one or more image editing instructions can include a model selection instruction that specifies an image generation model trained or fine-tuned to replace the source image with image content consistent with a background of the source image”, par 0052, “the user query can also be, for instance, “change to a monarch butterfly” which does not explicitly identify the source object (e.g., “white butterfly”) to be edited in the source image, but identifies a target object to be introduced into the source image, to generate an edited image that shares certain image content (e.g., an image background such as the pink milkweed plant) with the source image and that includes additional image content generated based on the user query. In some implementations, it is noted that, depending on one or more factors (correlation between the source image and the user query, a degree of image edit, etc.), a new image can be generated without using the source image, instead of the edited image which is generated utilizing the source image”, par 0071, “The first text prompt can additionally, or alternatively, include an instruction to identify a source object (or source content, or source area) of the image 151B to be edited (e.g., deleted, modified, added) and/or a position of the source object (or source content/area), in the text description. In some implementations, optionally, the source object (or source content) can be identified based on processing the user query 151A and/or the image 151B. In some implementations, the image understanding model 191 can be fine-tuned using one or more training instances 180A, so that the first text prompt no longer needs to be provided to the image understanding model 191 “, par 0092-0094, “the user input 210B requesting one or more image edits to the image 210A can identify the target object (e.g., cat), without identifying the source object (e.g., dog) in the image 210A that is to be replaced with the target object. For instance, as shown in FIG. 2A, the user input 210B can be: “change to cat”. In this case, the image-editing system can determine the source object (e.g., dog) in the image 210A to be replaced with the target object (e.g., cat). For instance, the image-editing system can determine the source object (e.g., dog) in the image 210A by: classifying object(s) present in the image 210A (e.g., using the visual language model 190 191 or an object recognition and classification model); and determining a particular object as the source object based on the particular object and the target object belonging to the same or similar classification/category (e.g., animal, or living life, etc.) …. the user input 210B can include a source object (e.g., dog in FIG. 2A) in the image 210A to be edited. In this case, the location description can indicate a location of the source object in the image 210A (without identifying locations for all objects present in the image 210A). ….the image mask can, for instance, mask all areas of the image 210A except for the source object, where content of the masked area(s) of the image 210A can be preserved (e.g., not modified) during processing of the image 210A to generate an edited image (e.g., 220B in FIG. 2B) that has a background visually the same as, or similar to, the image 210A (e.g., also referred to as “source image”). Techniques described in the present disclosure allows automatic generation of such image mask, without a user manually defining a mask area of the image mask, which can be time-consuming and requires accurate definition of the mask area from a human user “, Adamson, III: par 0043, “Image segmentation 120 is configured, generally, to identify and segment an object present in a source image 130 from a background of the source image 130 to obtain image segment 132 of such object and background. In some instances, an image segment 132 of an object may include a source image 130 and an image segment mask (e.g., a pixel-wise mask, without limitation) that marks an object (e.g., marks a location of an outline or area of an object in source image 130) and attaches a label that categorically classifies the object. Such a segment mask may be applied to the source image 130 change values of pixels in the source image 130 that are not associated with the object thereby defining the object in the resultant image (e.g., an edited source image or a copy). Similarly, image segment 132 of a background may include an image segment mask that marks a background (e.g., marks a location of an outline of a background in source image 130 “, par 0047, “ Image synthetization 126 may be configured, generally, to combine one or more of a styled image segment 134, a background image, an overlay image, and display information (e.g., text, a stamp, a logo, or an icon, without limitation) into a synthesized image 136. In a contemplated use, a background present in a source image 130 may be “replaced” with an image of a new background (e.g., using convolutional techniques for combining images, without limitation). Overlays and display information may be combined after or before style transfer effects are applied “, par 0058,” image or video capture used to generate a source image may not generate a depth pixel image, depth information may not be included with a source image 202, or included but just not useable or available. In some embodiments, an object present in an image may be segmented by a deep learning model trained to generate substitute depth pixel image, or by applying a segmentation model trained to segment objects and/or background without depth information. In some embodiments, object segmenters 208 may include trained image segmentation models configured to generate segment mask 216 in response to color pixel images, and more specifically, configured to segment objects present in source image 202 in response to color pixel image 20 “, par 0096, “FIG. 9 is a diagram depicting a process flow of image data obtained by performing image segmentation, classification, style transfer and background modification in accordance with one or more embodiments. A portrait object 902 and background 904 are present in the image segmented. A classification label is applied to the recognized image. In the third frame, style effects have been applied to the portrait object 902 to generate styled portrait 908, based on the classification label 906. In the fourth frame, background 904 has been replaced with new background 910. In the fifth frame, alpha has been adjusted on the styled portrait 908 so adjusted portrait 912 appears somewhat translucent in the synthesized image”). Regarding claim 16, Akerlund et al. teach a system comprising: a memory component; and a processing device coupled to the memory component, the processing device configured to perform operations (par 0026). The remaining limitations of the claim are similar in scope to claim 1 and rejected under the same rationale. Regarding claim 19, Akerlund et al. as modified by Adamson, III teach all the limitation of claim 16, and teach wherein the processing device is further configured to perform operations comprising: generating, using a segmentation model, the input mask based on the source image and a text prompt, wherein the input mask is based on a region of an element described by the text prompt (Akerlund et al.: par 0003, “Image editing is one of the most fundamental tasks in computer graphics, encompassing the process of modifying an input image through the use of an auxiliary input, such as a label, scribble, mask, or reference image. As described above, the current LLMs do not provide simple editing means for a given image, and generally lack control over specific semantic regions of the given image (e.g., using text guidance only). For example, even the slightest change in the textual prompt may lead to a completely different image being generated “, par 0013, “the user request to edit the image can be: “change the background of the image from beach to grass”. In the example, the region of the image to be masked can be a region to be preserved and can be indicated, for instance, using a bounding box surrounding a target object (e.g., a tourist), where image content (e.g., the tourist) within the bounding box is to be preserved and image content (e.g., the background, e.g., beach) outside the bounding box is to be edited. As another example, the user request to edit the image can be: “add a rabbit in the grass”. In this example, the region of the image to be masked can be a region to be edited and can be indicated, for instance, using a bounding box surrounding a portion of the grass that is to be placed with a rabbit, where image content outside of the bounding box is preserved for inclusion in the edited image “, par 0073, “the text representation 153 of the image 151B can further indicate location information for one or more objects (e.g., location of the boundary of the lawn, location of boundary of the dogwood tree, etc.) in the image 151B. In some implementations, the location information for the one or more objects can be applied to generate one or more image masks for subsequent use in generating an edited image or a new image, so that pixel values of certain pixels of the image 151B that correspond to image content of the image to be covered by the one or more image masks can be preserved (i.e., not being modified) and be included in the edited image (or the new image) “, par 0094-0095, “an image mask can be generated based on the location description that indicates the location for the source object (e.g., dog) present in the image 210A. The image mask can, for instance, mask the source object that needs to be edited. Alternatively, the image mask can, for instance, mask all areas of the image 210A except for the source object, where content of the masked area(s) of the image 210A can be preserved (e.g., not modified) during processing of the image 210A to generate an edited image (e.g., 220B in FIG. 2B) that has a background visually the same as, or similar to, the image 210A (e.g., also referred to as “source image”) “, Adamson, III: “Trained segmentation, classification, and style transformation models are used to apply style effects to image segments corresponding to objects (e.g., a portrait of a person) present in an image. When depth information is included with a source image it may be used to segment, classify, and/or apply style transformation to an image or image segments (as the case may be)”, par 0036-0037, “The way we do this is we apply image segmentation to the image to mark the presence of a known object type in the image (e.g., via a pixel-wise mask generated for each image and that can be combined with the color pixel image and depth pixel image via a bit-wise operation and the combination depicted) and to label the image segment with an indication of a known class type “, par 0043, “Image segmentation 120 is configured, generally, to identify and segment an object present in a source image 130 from a background of the source image 130 to obtain image segment 132 of such object and background. In some instances, an image segment 132 of an object may include a source image 130 and an image segment mask (e.g., a pixel-wise mask, without limitation) that marks an object (e.g., marks a location of an outline or area of an object in source image 130) and attaches a label that categorically classifies the object”). Regarding claim 20, Akerlund et al. as modified by Adamson, III teach all the limitation of claim 16, and further teach wherein the processing device is further configured to perform operations comprising: generating, using an inversion component, preliminary background features based on the source image (Akerlund et al.: par 0013, “the user request to edit the image can be: “change the background of the image from beach to grass”. In the example, the region of the image to be masked can be a region to be preserved and can be indicated, for instance, using a bounding box surrounding a target object (e.g., a tourist), where image content (e.g., the tourist) within the bounding box is to be preserved and image content (e.g., the background, e.g., beach) outside the bounding box is to be edited. As another example, the user request to edit the image can be: “add a rabbit in the grass”. In this example, the region of the image to be masked can be a region to be edited and can be indicated, for instance, using a bounding box surrounding a portion of the grass that is to be placed with a rabbit, where image content outside of the bounding box is preserved for inclusion in the edited image“, par 0052, “the user query can also be, for instance, “change to a monarch butterfly” which does not explicitly identify the source object (e.g., “white butterfly”) to be edited in the source image, but identifies a target object to be introduced into the source image, to generate an edited image that shares certain image content (e.g., an image background such as the pink milkweed plant) with the source image and that includes additional image content generated based on the user query“, par 0094, “the image mask can, for instance, mask all areas of the image 210A except for the source object, where content of the masked area(s) of the image 210A can be preserved (e.g., not modified) during processing of the image 210A to generate an edited image (e.g., 220B in FIG. 2B) that has a background visually the same as, or similar to, the image 210A (e.g., also referred to as “source image”, par 0098, “he text representation for the image 210A can be a description of all objects present in the image 210A, such as “The image depicts a dog riding a bike in a country road, with clouds and mountains in the background”. The text representation for the image 210A can also include location information of the objects (all or only the source object) present in the image 210A, such as “The image depicts a dog riding a bike in a country road, with clouds and mountains in the back, a location of the dog is indicated by . . . ”. “, par 0100, “The second ML model can be, for instance, an image-generation model trained to generate image(s) based on text descriptions. Continuing with the non-limiting example above, the edited image 220B can have same background as the source image 210A, except for the source object of “dog” being replaced with the target object of “cat”. In other words, the edited image 220B can show a cat riding a bike in a country road, with a background of mountains and/or clouds, where object(s) (e.g., clouds, mountains, bike, country road) in the edited image 220B that are not the target object (e.g., cat) are the same as those in the source image 210A”, Adamson, III: par 0043, “Image segmentation 120 is configured, generally, to identify and segment an object present in a source image 130 from a background of the source image 130 to obtain image segment 132 of such object and background. In some instances, an image segment 132 of an object may include a source image 130 and an image segment mask (e.g., a pixel-wise mask, without limitation) that marks an object (e.g., marks a location of an outline or area of an object in source image 130) and attaches a label that categorically classifies the object. Such a segment mask may be applied to the source image 130 change values of pixels in the source image 130 that are not associated with the object thereby defining the object in the resultant image (e.g., an edited source image or a copy). Similarly, image segment 132 of a background may include an image segment mask that marks a background (e.g., marks a location of an outline of a background in source image 130)”, par 0119, “a system and method synthesizes an output image from an input image or input video frames by performing image segmentation to identify a foreground object and a background in the input image. In some embodiments, the image segmentation is performed by using a depth pixel matte of the input image. The system and method applies a style transfer effect to the foreground object only using the depth map”). Claim(s) 6 is/are rejected under 35 U.S.C. 103 as being unpatentable over U.S. PGPubs 2025/0349054 to Akerlund et al. in view of U.S. PGPubs 2021/0110588 to Adamson, III, further in view of Park et al. (Park, T., Liu, M.Y., Wang, T.C., Zhu, J.Y.: Semantic image synthesis with spatially-adaptive normalization. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 2337–2346 (2019)). Regarding claim 6, Akerlund et al. as modified by Adamson, III teach all the limitation of claim 5, but keep silent for teaching wherein: the style transfer comprises a masked adaptive instance normalization. In related endeavor, Huang et al. teach wherein: the style transfer comprises a masked adaptive instance normalization (abstract, “We propose spatially-adaptive normalization, a simple but effective layer for synthesizing photorealistic images given an input semantic layout. Previous methods directly feed the semantic layout as input to the deep network, which is then processed through stacks of convolution, normalization, and nonlinearity layers. We show that this is suboptimal as the normalization layers tend to “wash away” se mantic information. To address the issue, we propose using the input layout for modulating the activations in normalization layers through a spatially-adaptive, learned trans formation”, Fig 1, section 1-3, “We assume the training dataset contains registered segmentation masks and images. With the proposed spatially-adaptive normalization, our compact network achieves better results compared to leading methods”). It would have been obvious to a person of ordinary skill in the art at the time before the effective filing data of the claimed invention to modified Akerlund et al. as modified by Adamson, III to include wherein: the style transfer comprises a masked adaptive instance normalization as taught by Huang et al. to propose spatially-adaptive normalization, a simple but effective layer for synthesizing photorealistic images given an input semantic layout to allows user control over both semantic and style to demonstrate the advantage of the proposed method over existing approaches, regarding both visual fidelity and alignment with input layouts. Claim(s) 8 is/are rejected under 35 U.S.C. 103 as being unpatentable over U.S. PGPubs 2025/0349054 to Akerlund et al. in view of U.S. PGPubs 2021/0110588 to Adamson, III, further in view of U.S. PGPubs 2023/0267663 to Chopra et al. Regarding claim 8, Akerlund et al. as modified by Adamson, III teach all the limitation of claim 1, but keep silent for teaching further comprising: performing boundary smoothing on the input mask to obtain a modified mask, wherein the concept features are generated based on the modified mask. In related endeavor, Chopra et al. teach further comprising: performing boundary smoothing on the input mask to obtain a modified mask, wherein the concept features are generated based on the modified mask (par 0050-0051, “As illustrated in the representation 400, the combination module 204 uses the aggregate per-pixel displacement map to generate the warped garment image I.sub.wrp. To do so in one example, the combination module 204 uses the aggregate per-pixel displacement map to warp the garment image I.sub.p and a mask M.sub.p to generate the warped garment image I.sub.wrp and a warped binary garment mask M.sub.wrp, respectively. Additionally, intermediate flow maps f.sub.l for l∈{0, . . . , K} are used to produce intermediate warped images I.sub.wrp.sup.l and intermediate warped masks M.sub.wrp.sup.l. Each of the warped images (final and intermediate) are subject to an L1 loss and a perceptual similarity loss with respect to garment regions of the first digital image 302. Each predicted warped mask is subject to a reconstruction loss with respect to a ground truth mask. The predicted flow maps are subjected to a total variation loss to ensure spatial smoothness of flow predictions. As shown, pixels depicting the garment in the second digital image 304 are displaced to align with the pose of the person to generate the warped garment image I.sub.wrp. For example, the combination module 204 generates the warped garment data 216 as describing the warped garment image I.sub.wrp.”, par 0056, “the output module 208 includes a machine learning model such as a convolutional network 608 and the output module 208 implements the convolutional network 608 to generate a digital image 610 that depicts the person in the pose wearing the garment I.sub.tryon. In this example, the convolutional network 608 includes six encoder and decoder layers and the convolutional network 608 processes the warped garment data 216, the segment mask data 218, and the additional prior data 212 to generate the digital image 610 that depicts the person in the pose wearing the garment I.sub.tryon. For example, this is representable as: I.sub.tryon=M.sub.out*I.sub.wrp+(1−M.sub.out)*I.sub.rp where: M.sub.out is generated by the convolutional network 608 and is a composite mask for garment pixels in the try-on output; and I.sub.rp is generated by the convolutional network 608 and is a rendered person including all pixels depicting the person except the garment in the try-on output.“). It would have been obvious to a person of ordinary skill in the art at the time before the effective filing data of the claimed invention to modified Akerlund et al. as modified by Adamson, III to include further comprising: performing boundary smoothing on the input mask to obtain a modified mask, wherein the concept features are generated based on the modified mask as taught by Chopra et al. to generate a warped garment image by combining the candidate appearance flow maps as an aggregate per-pixel displacement map using a convolutional gated recurrent network to ensure spatial smoothness of flow predictions to accurately predict three-dimensional geometries (e.g., body-part ordering) based on two-dimensional digital images. Claim(s) 9 and 17 is/are rejected under 35 U.S.C. 103 as being unpatentable over U.S. PGPubs 2025/0349054 to Akerlund et al. in view of U.S. PGPubs 2021/0110588 to Adamson, III, further in view of U.S. Patent 11995803 to Karpman et al.. Regarding claim 9, Akerlund et al. as modified by Adamson, III teach all the limitation of claim 1, but keep silent for teaching wherein generating the synthetic image comprises: obtaining a noise map; and denoising the noise map based on the concept features. In related endeavor, Karpman et al. teach wherein generating the synthetic image comprises: obtaining a noise map; and denoising the noise map based on the concept features (col 3:15-36, “Text-to-image diffusion model 112 can execute the base image diffusion model 120 (and the high-resolution diffusion models 116) on the assembled training set (e.g., text-image pairs) to infer and/or encode custom parameters for iteratively transforming randomly sampled visual noise into a visually appealing synthetic image that aligns with visual concepts described by a text promp “, col 5:15-36, “the base image diffusion model 120 defines a deep learning network (e.g., a convolutional neural network, a residual neural network, etc.) configured (e.g., through the training described) to generate images from random (e.g., Gaussian) noise based on text prompts and/or descriptions. The base image diffusion model 120 can include a U-net architecture (e.g., Efficient U-Net) defined from residual and multi-head attention blocks that enable the base image diffusion model 120 to progressively denoise (e.g., infill, generate, augment) image data according to cross-attention inputs based on the text prompt. The base image diffusion model 120 can therefore: receive one or more text embeddings from the set of pre-trained text encoders 118; receive and/or initialize a (randomly sampled) noise distribution at a preset resolution (e.g., 64 pixels by 64 pixels); and transform the noise distribution into a base image at the preset resolution based on the one or more text embeddings and parameters, weights, and/or paths corresponding to an iterative denoising process learned by the base image diffusion model 120 during training”, col 14:32-38, “ the base image diffusion model 120 can automatically infer, derive, and/or produce a set of initial weights, parameters, and/or paths within the model architecture that therefore enable the base image diffusion model 120 to transform randomly sampled noise into a base image that is semantically aligned with an input text prompt”). It would have been obvious to a person of ordinary skill in the art at the time before the effective filing data of the claimed invention to modified Akerlund et al. as modified by Adamson, III to include wherein generating the synthetic image comprises: obtaining a noise map; and denoising the noise map based on the concept features as taught by Karpman et al. to advance in language modeling and diffusion models have significantly increased the capabilities of deep learning networks in generating customizable photorealistic and/or particularly stylized images from natural language prompts, leading to significant public and industry attention on their capability to rapidly generate high quality digital artwork. Regarding claim 17, Akerlund et al. as modified by Adamson, III teach all the limitation of claim 16, but keep silent for teaching wherein: the image generation model comprises a diffusion U-Net. In related endeavor, Karpman et al. teach wherein: the image generation model comprises a diffusion U-Net (col 5:15-62, “the base image diffusion model 120 defines a deep learning network (e.g., a convolutional neural network, a residual neural network, etc.) configured (e.g., through the training described) to generate images from random (e.g., Gaussian) noise based on text prompts and/or descriptions. The base image diffusion model 120 can include a U-net architecture (e.g., Efficient U-Net) defined from residual and multi-head attention blocks that enable the base image diffusion model 120 to progressively denoise (e.g., infill, generate, augment) image data according to cross-attention inputs based on the text prompt”). It would have been obvious to a person of ordinary skill in the art at the time before the effective filing data of the claimed invention to modified Akerlund et al. as modified by Adamson, III to include wherein: the image generation model comprises a diffusion U-Net as taught by Karpman et al. to advance in language modeling and diffusion models have significantly increased the capabilities of deep learning networks in generating customizable photorealistic and/or particularly stylized images from natural language prompts, leading to significant public and industry attention on their capability to rapidly generate high quality digital artwork. Allowable Subject Matter Claims 7 and 18 are objected to as being dependent upon a rejected base, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. The following is a statement of reasons for the indication of allowable subject matter: The cited prior art fails to teach the combination of elements recited in claim 7, including " further comprising: identifying a shape of the concept using a cross-attention layer of the image generation model; and computing shape guidance based on the shape and the input mask, wherein the synthetic image is generated based on the shape guidance". The following is a statement of reasons for the indication of allowable subject matter: The cited prior art fails to teach the combination of elements recited in claim 18, including " wherein: the image generation model comprises a location adaptation module, a style adaptation module including an instance normalization component, a scale adaptation module, and a content adaptation module". Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to Jin Ge whose telephone number is (571)272-5556. The examiner can normally be reached 8:00 to 5:00. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Jason Chan can be reached at (571)272-3022. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. JIN . GE Examiner Art Unit 2619 /JIN GE/ Primary Examiner, Art Unit 2619
Read full office action

Prosecution Timeline

Nov 24, 2024
Application Filed
Jul 30, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12749262
THREE-DIMENSIONAL MODEL GENERATION METHOD, THREE-DIMENSIONAL MODEL GENERATION DEVICE, AND NON-TRANSITORY COMPUTER READABLE MEDIUM
3y 4m to grant Granted Sep 29, 2026
Patent 12743815
ON COMPRESSION OF A MESH WITH MULTIPLE TEXTURE MAPS
2y 1m to grant Granted Sep 22, 2026
Patent 12731176
Method for creating digital art from photos and videos of coins and various methods of presenting the art to be viewed.
3y 8m to grant Granted Sep 08, 2026
Patent 12731352
Video System with Scene-Based Object Insertion Feature
3y 0m to grant Granted Sep 08, 2026
Patent 12718500
DELIVERING VIRTUALIZED CONTENT
2y 4m to grant Granted Aug 25, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
80%
Grant Probability
98%
With Interview (+18.8%)
2y 6m (~8m remaining)
Median Time to Grant
Low
PTA Risk
Based on 552 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month