Prosecution Insights
Last updated: October 02, 2026
Application No. 18/939,499

SYSTEMS AND METHODS OF IMAGE EDITING BASED ON MULTIMODAL LARGE LANGUAGE MODELS

Non-Final OA §103
Filed
Nov 06, 2024
Priority
Aug 05, 2024 — provisional 63/679,597
Examiner
SONNERS, SCOTT E
Art Unit
2613
Tech Center
2600 — Communications
Assignee
Samsung Electronics Co., Ltd.
OA Round
1 (Non-Final)
69%
Grant Probability
Favorable
1-2
OA Rounds
1y 4m
Est. Remaining
81%
With Interview

Examiner Intelligence

Grants 69% — above average
69%
Career Allowance Rate
271 granted / 392 resolved
+7.1% vs TC avg
Moderate +12% lift
Without
With
+12.2%
Interview Lift
resolved cases with interview
Typical timeline
3y 3m
Avg Prosecution
15 currently pending
Career history
407
Total Applications
across all art units

Statute-Specific Performance

§101
9.3%
-30.7% vs TC avg
§103
38.2%
-1.8% vs TC avg
§102
27.1%
-12.9% vs TC avg
§112
16.2%
-23.8% vs TC avg
Black line = Tech Center average estimate • Based on career data from 392 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claim(s) 1-3, 9-14 and 18-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Wang et al1 (“Wang”) in view of Zou et al2 (“SEEM”). Regarding claim 1, Wang teaches a method of image editing (see below steps which comprise steps for a method of image editing) comprising: generating image tokens from an input image and word tokens from an editing prompt (note that a “token” as broadly recited is interpreted to be any discrete representation of a type of data and thus any representation of an image or word which is discrete and able to be processed by some other component for some purpose may be considered a token; see Wang, section 3.2 teaching “Given a user instruction I as input” which is an editing prompt such as “change the cat to a fox”, there is use of “a large language model…to extract segmentation prompt q for the segmenter, and an input caption…and an edited caption…for the image editor” such that this creates word tokens from an editing prompt (note that additionally word tokens would be inherently generated when utilizing ChatGPT as in order to process the text input such text necessarily must be tokenized and further processed into embeddings relating to such tokens, though simply the outputs produced by the LLM may be considered tokens as well) and image tokens are generated in numerous ways where as can be seen in figure 2 and as explained in section 2.2 and 3.2, the input image is used to generated tokens related to the image such as “Description d of the image” making these image tokens functionally as they are tokens describing an image, and furthermore additionally and alternatively, the input image is turned into tokens when consumed by the “segmenter” in “Grounding DINO” as Grounding DINO is understood by a PHOSITA to consume an input image such that at the very least the pixels of the image being processed would constitute image tokens and given that it is disclosed as using “an open-set object detection which combines the Transformer-based detector DINO…with grounded pre-training to detect arbitrary objects with human inputs such as category names or referential expressions” this use of a transformer to process the image inherently requires that the input to the transformer has been tokenized in some manner; furthermore note such inherency is supported by the teachings of Grounding DINO provided in the pertinent art section below evidencing that use of Grounding DINO as in Wang inherently requires to generate tokens of some form to be fed to the model, else one would not be using the Grounding DINO technique ); generating a mask token based on an artificial intelligence model processing the image tokens and the word tokens (note that a token is interpreted broadly as above and a mask token is any token that can be associated with a mask or is used in generating a mask; thus see Wang, sections 3.2-3.3 teaching that “Grounding DINO is first applied to get a bounding box for a given segmentation prompt q” by inputting the input image and segmentation prompt such that this bounding box may be considered a token as it is a discrete form of data fed to another processing module and the artificial intelligence model comprises the “large language model ChatGPT” as well as the “Segmenter” which as explained above uses “an open-set object detection which combines the Transformer-based detector DINO…with grounded pre-training to detect arbitrary objects with human inputs such as category names or referential expressions” such that such models working together may be considered an artificial intelligence model processing the image tokens and the word token through the “Segmenter” portion of the AI model and Grounding DINO take in the input image “xo” and the token “q” and generates the mask token in the form of the bounding box which provides a coarse localization that is used to generate the ”per-pixel binary mask” by another processing stage; here then the image tokens can be considered the input image “Img” passed as “xo” to Grounding DINO which are used along with the word token “q” to generate the mask token, or the image tokens can be considered to be the “description d” from the image which is processed by the LLM of the AI model to guide the segmentation prompt q as that segmentation prompt is the result of processing the description d by the LLM as well as the input editing); generating an editing mask based on a mask decoder processing the mask token, word embeddings of the editing prompt, and visual embeddings of the input image (note that an “embedding” is considered to be see Wang, section 2.2 and 3.3 where the mask token bounding box is input to SAM which functions as a mask decoder which takes in mask token “b” and input image “xo” to generate a “per-pixel binary mask M” which as in section 3.4 it uses for “the mask-guided image editing” such that this “bounding box” is “refined” to the mask to identify the masked area for editing such that the ”region within the mask will have the changes guided by the edited caption, while the region outside the mask will be mapped back to the original pixels” making this the mask actually used for editing the image; note that Wang teaches use of SAM as in sections 2.2 and 3.3, and SAM (as provided in the Pertinent Art section below) inherently requires encoding of the input image into visual embeddings along with encoding the guiding prompt into word embeddings if it is words or box embeddings if the guiding prompt is a box such as would be passed from the Grounding DINO portion of the model, however note that such inherency need not necessarily be relied upon as the SEEM reference combined below squarely teaches such use of both types of embeddings and a box mask token to provide a per-pixel mask that can be used to guide image-editing); generating a correlation map that correlates the editing mask to a set of one or more words of the editing prompt (note that the generating of the correlation map is not recited as performed by any particular model or component named nor is the manner of generation limited nor form of the correlation map limited such that a correlation map is any data structure establishing a correspondence between at least two things, in this case the editing mask and one or more words of the editing prompt; see Wang, section 3.4 teaching “the mask-guided image editing” which establishes a relationship between which areas of the image defined by the editing mask are to be changed in relation to one or more words of the editing prompt ‘’c” such that “region within the mask will have the changes guided by the edited caption, while the region outside the mask will be mapped back to the original pixels” such that the mask-guided denoising step applies that caption-conditioned prediction only within the mask according to equation 4 and where as in section 3.1 it is explained “c is the condition of the diffusion model” and “we only consider c to be a text prompt of a text-guided diffusion model” such that the diffusion model is able to condition generation on c as guided by the editing mask); and generating an output image based on the correlation map, the output image comprising an edited version of the input image according to the editing prompt (see Wang, section 3.4, teaching “the mask-guided image editing” to “get the edited image Img(e)” after “iteratively applying Eq. 4” which generates this image based on the correlation map formed from utilizing equation 4 to correlate the edit caption words to the editing mask ensuring the output image is an edited version of the input image according to the editing prompt such as seen in figure 2 for example where the edited image is a version of the input image with a cat edited to be a fox). Wang teaches all of the above, but fails to teach specifically that the mask decoder takes in all of the recited inputs of the mask token, word embeddings of the editing prompt, and visual embeddings of the image. Rather as explained above, the SAM model utilized in Wang does teach a mask decoder which can take in a box as well visual embeddings of an input image to provide a per pixel editing mask, but it is not disclosed as utilizing all of a mask token, word embeddings, and visual embeddings of the image as rather it takes in the image and bounding box from Grounding DINO without any text embedding. Thus Wang stands as a base system upon which the claimed system can be considered an improvement through a mask decoder that can utilize a mask token, word embeddings, and visual embeddings of an input image to generate a mask for editing such that use of such word embeddings could be considered to better guide the mask refinement to generate a per-pixel mask. In the same field of endeavor relating to processing an input image to obtain prompt-guided masking and segmentation into per-pixel masks (see SEEM, section 1 and figure 2 teaching “a new prompting scheme that can encode various user intents into prompts in a joint visual-semantic space, enabling strong flexibility for various segmentation tasks and generalization capability to unseen prompts or their combinations”), SEEM teaches a mask decoder which generates a refined mask from a mask token such as a bounding box as well as word embeddings of a prompt and visual embeddings of an input image (see SEEM, figure 2, “SEEM encodes image, text, and human inputs into joint visual-semantic space as queries, features, and prompts, and then decodes queries to class and mask embeddings” and as can be seen in figure 2 the text prompt input is input to a “text encoder” and the image is input to the image encoder and the visual sampler takes in the mask token such a box in order to present them as encodings or embeddings to the “Joint Image-Text Representation Space” and as in section 3.1, “Given an input image I… an image encoder is first used to extract image features Z. Then, SEEM-Decoder predicts the masks M and semantic concepts C based on the query outputs Om h (mask embeddings) and Oc h (class embeddings), which interact with text, visual, and memory prompts” as in equation 1, and “we use visual prompts to unify all non-textual prompts and align them with textual prompts” and “our model becomes familiar with all prompt types and supports a variety of compositions, such as no prompts, one prompt type, or both visual and textual prompts using the same model and weights. In particular, the visual and textual prompts can be simply concatenated and fed to SEEM-Decoder” and “our visual prompt features are aligned with textual features in a joint visual-semantic space”). Thus SEEM teaches a known technique applicable to the base system of Wang. Therefore it would have been obvious for one of ordinary skill in the art before the effective filing date of the invention to modify Wang to utilize the SEEM model in connection with Grounding DINO model instead of the SAM model as doing so would be no more than application of a known teaching to a base system ready for improvement where such modification would yield predictable results and result in an improved system. The predictable result of the combination would be that instead of using the SAM model only utilizing the visual and box mask token prompt to generate the per-pixel editing mask, the SEEM model would instead be used and take in the bounding box mask token from Grounding DINO already being supplied as well as the input image already being supplied, and would also take in the guiding prompt as word embeddings to supply a refined mask to base the editing on. The predictable result would be that the editing mask is generated by a mask decoder that takes in a mask token, word embeddings of an editing prompt and visual embeddings of the input image to generate a refined mask, which then would be passed to the Image Editor in Wang just as the mask is already taught as being passed from SAM. This would result in an improved system and one having ordinary skill in the art would have been motivated to apply the techniques, as the generated masks could be tied even more closely to the correct words of the edit prompt and SEEM teaches that it is an improvement over SAM as it allows to provide segmentations with semantic meaning from the text prompts and other inputs to guide the segmentation (see SEEM, section 2, “Though SAM demonstrates strong zero-shot performance, it produces segmentations without semantic meaning” and as in section 3.1 “our visual prompt features are aligned with textual features in a joint visual-semantic space”). Regarding claim 2, Wang as modified teaches all that is required as applied to claim 1 above and further teaches wherein the mask token is generated based on the artificial intelligence model determining the set of one or more words of the editing prompt are applicable to the input image based on at least one of the image tokens correlating to at least one of the word tokens (see Wang as modified where Wang as in figure 2, section 2.2 and 3.2, where as explained above, the “Description d” provided by BLIP2 may be considered an image token as it’s a token representing the image, and this is fed along with the word tokens to the ChatGPT processing stage which allows it to provide the set of one or more words applicable to the editing prompt such as “cat” where “vision-language model BLIP2 can process the image and is able to answer questions about its content. BLIP2 can provide a short description of the original image, which can assist ChatGPT to provide prompts for the segmenter and captions for the image editor” and “query BLIP2 to obtain a description d of the image. Given an input image, we first ask BLIP2 “Is this a photo, a painting or another kind of art?”. We denote the answer as ρ and reuse it in another query to BLIP2 composed as “ρ of” to obtain a completed sentence describing the image as input prompt to ChatGPT. ChatGPT in turn can refine the prompt by identifying which object to edit and provide more details even when the user instruction does not specify the content to be edited or the description is incomplete to unambiguously refer to the intended objects in the image” such that here BLIP2 and ChatGPT determine the applicable words of the editing prompt to the input image based on image tokens which such applicable word is then used by the Segmenter to segment applicable words). Regarding claim 3, Wang as modified teaches all that is required as applied to claim 1 above and further teaches wherein generating the editing mask is based on feeding the word embeddings to a first transformer decoder layer of the mask decoder and feeding the visual embeddings to a second transformer decoder layer of the mask decoder, the mask decoder being trained to generate the editing mask based on the mask decoder processing the mask token, word embeddings of the editing prompt, and visual embeddings of the input image (see Wang as modified by SEEM where SEEM is combined to already teach generating the editing mask and as combined it teaches such decoder layers being fed to respective inputs as in SEEM at section 3.1 and as in figure 2 and Algorithm 1, the image features or word embeddings go to a cross-attention step layer whereas the text prompt features go to the self-step such that the visual embeddings enter the cross-attention sub-layer, the word embedding enters a second prompt-attention sub-layer and the mask token is carried in the same prompt set as the word embedding such that the mask decoder is trained to generate the mask based on the mask decoder processing the mask token, word embeddings and visual embeddings of the input image). Regarding claim 4, Wang as modified teaches all that is required as applied to claim 1 The method of claim 1, wherein generating the correlation map is based on matrix multiplication between the mask token and the word embeddings (note that as the correlation map is “based on” such matrix multiplication, but does not specifically identify where, when or what performs such matrix multiplication upon which the correlation generated is based, this means that if matrix multiplication between the mask token and the word embeddings takes place at any time and the results of such multiplication are used in generating the correlation map then the limitation is met; see Wang as modified by SEEM, where SEEM teaches a mask token and word embeddings involved in a matrix multiplication as in section ). Regarding claim 9, Wang as modified teaches all that is required as applied to claim 1 above and further teaches wherein a word embedder generates the word embeddings from the editing prompt and a visual encoder generates the visual embeddings from the input image (see Wang as modified by SEEM where SEEM provides such functioning already in combination as in section 1 teaching “we propose to encode points, masks, text, boxes, and even a referred region from another image into prompts in the same joint visual-semantic space” where such encodings correspond to embeddings of each type of input and the model is comprised of “a simple Transformer encoder-decoder architecture [31, 6] with an extra text encoder” and “image encoder and text encoder are used as the prompt encoder to encode all types of queries, which are fed into the decoder” and “encode all spatial queries, namely, points, boxes, scribbles and masks into visual prompts by pooling their corresponding visual features from the image encoder, and use the text encoder to convert text queries into text prompts” such that here a word embedder such as the text encoder generates the usable word embeddings to use along with the visual embeddings from the image encoder). Regarding claim 10, Wang as modified teaches all that is required as applied to claim 1 above and further teaches wherein a diffusion model generates the output image based on the diffusion model processing the correlation map, the input image, and the editing prompt (see Wang, note title includes “Diffusion-based Image editing” and as in the abstract “adopt Stable Diffusion and the mask-guided generation from DiffEdit” and as in section 3.1 and 3.4 teaching as explained above the correlation map corresponding to the application of the editing mask as applied in the mask-guided denoising and where the diffusion model generates the output image as in connection with equations 1-4 where in equation 4 the diffusion model processes the correlation map showing where to edit the image as well as the input image and editing prompt provided from the language processor to condition the “text-guided diffusion model” such that the output image is generated where the “region within the mask will have the changes guided by the edited caption, while the region outside the mask will be mapped back to the original pixels” such changes being provided by the diffusion model; note section 4.1 discloses to “use Stable Diffusion 1.5 as the backbone of the image editor” such that again it is clear a diffusion model generates the outputs based on processing the above inputs as explained above). Regarding claim 11, Wang as modified teaches all that is required as applied to claim 1 above and further teaches wherein the artificial intelligence model comprises a multimodal large language model (as explained above the AI model may be considered a combination of AI models working together, and additionally a multimodal large language model any LLM that can process and reason across multiple types of data or modalities and note that the claim does not specify exactly how the MLLM is utilized; see Wang as modified where as in Wang section 2.2-3.4 and as in figure 2 it can be seen that the AI model would comprise the “BLIP2” model as well as the ChatGPT, Grounding DINO and SEEM models all working in concert where as in section 3.2, the “Description d” provided by BLIP2 may be considered an image token as it’s a token representing the image, and this is fed along with the word tokens to the ChatGPT processing stage which allows it to provide the set of one or more words applicable to the editing prompt such as “cat” where “vision-language model BLIP2 can process the image and is able to answer questions about its content. BLIP2 can provide a short description of the original image, which can assist ChatGPT to provide prompts for the segmenter and captions for the image editor” and “query BLIP2 to obtain a description d of the image. Given an input image, we first ask BLIP2 “Is this a photo, a painting or another kind of art?”. We denote the answer as ρ and reuse it in another query to BLIP2 composed as “ρ of” to obtain a completed sentence describing the image as input prompt to ChatGPT. ChatGPT in turn can refine the prompt by identifying which object to edit and provide more details even when the user instruction does not specify the content to be edited or the description is incomplete to unambiguously refer to the intended objects in the image” such that as BLIP2 is a “vision language model” it is a multimodal LLM as it can reason across multiple types of data instead of simply text). Regarding claims 12-14, the instant claims recite a “device comprising: one or more processors; and memory storing instructions that, when executed by the one or more processors, cause the device to…” perform the same steps as in the method claims 1-3. Wang as modified teaches the steps of the method and further teaches that a device is caused to perform the method where the device comprises one or more processors and a memory storing instructions (see Wang as modified where Wang’s technique and instructions causing a device to execute the technique is taught as “code at https://github.com/QianWangX/InstructEdit” where such instructions are taught as performed by a device comprising a processor programmed to execute instructions given the field of endeavor and as explained in section 4.1 teaching a processor such as “a single NVIDIA A100” by which the techniques “are performed”). In light of this, the limitations of claims 12-14 correspond to the limitations of claims 1-3, respectively; thus they are rejected on the same grounds as claims 1-3, respectively. Regarding claims 18-20, the instant claims are directed toward a device in the form of a “non-transitory computer-readable medium storing code that comprises instructions executable by a processor to” perform the method as in claims 1-3, respectively. Similar to the above analysis, Wans as modified teaches the steps of the method and further teaches a device such as a non-transitory computer-readable medium storing code that comprises instructions executable by a processor to perform the method (see Wang as modified above teaching the method steps and Wang teaching such instructions in the form of “code at https://github.com/QianWangX/InstructEdit” and as in section 4.1 “the experiments are performed on a single NVIDIA A100. We use Stable Diffusion v1.5 as the backbone of the image editor. We use the model’s weights and implementation of Grounded Segment Anything from https://github.com/IDEA-Research/Grounded-Segment-Anything” such that here a GPU such as a NVIDIA A100 is a processor which runs instructions such as stored in the non-transitory computer-readable mediums of the github repository at the links such that in order to run the technique on an A100 a person having ordinary skill in the art recognizes this requires such loading of such instructions to run the experiments as disclosed). Allowable Subject Matter Claims 4-8 and 15-17 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. The following is a statement of reasons for the indication of allowable subject matter: the prior art of record fails to teach or suggest the respective claim limitations when considered as a whole. Regarding claim 4, claim 4 requires, inter alia, “wherein generating the correlation map is based on matrix multiplication between the mask token and the word embeddings.” Here such narrowing of how the correlation map is generated not only further specifies how the correlation map must be generated but also requires certain aspects of independent claim 1 to be read more narrowly as well. This is because given that there must be a matrix multiplication between whatever is determined to be a mask token and a word embedding, this requires that the generated mask token as required to be generated must also be output as a mask token in a format in which such a matrix multiplication (would in order to be operatively defined, have a definition), would require the word embeddings to be in the form of an embedding that can be matrix multiplied. Thus the components upstream of the correlation map generation, such as the artificial intelligence model that generates the mask token, produce and pass the mask token in some embedded format that can be multiplied against the previously generated word embeddings. The prior art fails to teach or suggest such limitations when considered as a whole. Wang produces its mask through a segmenter that detects a bounding box that can be considered a mask token and combined with SEEM, such mask token is consumed along with word embeddings in order to output a mask. However, neither Wang nor SEEM teach that the mask token is matrix multiplied against word embeddings as Wang’s grounding of the mask in the prompt is achieved already through Grounding DINO and in SEEM the mask embedded as a virtual prompt is never matrix multiplied against the word embeddings of a prompt, rather both are concatenated and fed to the decoder. Note that Wang as modified by SEEM provides the closest prior art in the Examiner’s search which provides a similar pipeline wherein a mask token is produced and then consumed by a downstream component to further refine the generated mask. The Examiner is unable to find any teaching or suggestion in the prior art of such a mask token generated as required and utilized as required that is matrix multiplied by word embeddings of an editing prompt to generate a correlation map that correlates an editing mask to one or more words of the editing prompt and is then used in generating the output image. Thus the claims contain allowable subject matter. Similar reasoning applies to claim 15 which contains similar allowable subject matter. Regarding claim 5, the instant claim requires, inter alia, “generating a negative token based on the artificial intelligence model processing the image tokens and the word tokens.” Wang as modified fails to teach or suggest any negative token also being generated by the AI model processing the image tokens and word tokens. While negative tokens are known in the art such as the ability to feed conditions to a diffusion model such as Stable Diffusion to guide it away from certain latent spaces, such that tokens that encode these conditions and are fed to a diffusion model could be considered negative tokens. These are not negative tokens that are generated based on processing the image tokens and word tokens, but rather serve a different purpose to serve as negative tokens, but not to be generated from the image tokens and word tokens. Note that SEEM actually does teach a “Neg_mask” as in Algorithm 1 on page 5 and “’Negative’ means adding negative tokens during interactive segmentation” on page 7 explaining various ways the model was experimented with. However, these negative tokens are supplied by a user and correspond to inputs as in Algorithm 1 and are not generated by an AI model based on image and word tokens. Furthermore, as noted above Wang, as modified by SEEM, is the closest compatible prior art which generates a mask token and then generates an editing mask from the mask token using a mask decoder as recited, and does not teach generation of a negative token by the same AI model. Other relevant prior art such as for example as described in Nguyen et al3 (Nguyen), also fails to teach or suggest such generation of negative tokens and also does not teach a generation of a mask token which is defined and functions as in the claimed invention (see Nguyen, entire paper, note especially sections 3.1-.3.4 teaching relevant techniques utilizing LLMs and MLLMs and 4.2 teaching instruction guided diffusion models). Xu et al (US PGPUB No. 20220067992) provides another teaching in the same field of endeavor which also seeks to utilize natural language prompts to edit an image and attempts to ground the process by associating a mask with different editing operations (see Xu, paragraphs 0035-0061). From an editing prompt the prompt is encoded into tokens or embeddings which allow the system to understand the type of editing operation to be performed and an area may be masked to apply a recognized editing operation such as an “inpaint” operation. However, like the prior art above, there is no teaching or suggestion of the same process or generating a mask token and then generating an editing mask as recited from the mask token using a mask decoder. Furthermore, Xu does not teach generating a correlation map in the same manner as rather once a masked area is determined, it will be filled according to the operation and will not specifically further attend to words of the prompt or the visual embeddings as the mask has already been created. Thus the claims contains allowable subject matter. Note that claim 16 contains the same allowable subject matter identified with respect to claim 5 and is allowable at least for the same reasons. Note that dependent claims 6-8 and 17 are considered to contain allowable subject matter at least because they contain the allowable subject matter of their respective parent claims. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. See Liu et al4 (“Liu”) – teaching the Grounding DINO model utilized by Wang, note that Liu confirms the Examiner’s assertion that the inputs to the Grounding DINO model are word and image embeddings that are used to output a mask token such as a bounding box as in section 3 and figure 3 where “input text” is passed to a “text backbone” for “text feature extraction” which produces word embeddings and “image backbone” provides image features or embeddings which are then used to have the text guide the output and scoring of appropriate bounding boxes which are generated with the output box “with the largest scores as the output”. However Grounding DINO does not teach or suggest any of the above noted allowable subject matter. See Kirillov et al5 (“SAM”) – teaching the SAM model utilized by Wang, but swapped out for SEEM in the combination above. Of note is that, as taught in section 3, while SAM can be provided with a box or mask as or text as a prompt, along with the input image, to achieve the pixel level segmentation editing mask, there is no actual teaching in SAM of utilizing both types of input or using one to guide another and thus SAM cannot be relied upon to teach a mask decoder that utilizes all of a text embedding, mask token and visual embedding. However section 3 teaches that use of SAM would involve embeddings of the image as well as embeddings of the text, which further shows that SEEM is compatible with Wang in combination as both SEEM and SAM operate similarly, with SEEM explicitly providing additional functionality to accept multiple types of inputs simultaneously. However Kirillov does not teach or suggest any of the of the above noted allowable subject matter. See Guo et al6 (“Guo”) - teaching Focus On Your Instruction (FOI), which also seeks to guide editing of an image with a text prompt and which involves masking and in some way correlating a text prompt with masks to guide the image editing and generation process and is meant for “precisely extracting regions of interest for each instruction” and “guiding the denoising process to concentrate within these regions of interest” (see Abstract). Guo teaches as in section 4 and figure 4 that a mask is generated from cross attention layers in early stages of the denoising process being guided by an editing prompt and masks are extracted from these cross attention layers that correspond to different editing keywords. These masks are then “broadcast” across its corresponding subinstruction from the text and these masks are “concatenated for all sub-instructions” and this “mask is then adaptively interpolated across each cross-attention layer”. Thus instead of a mask decoder functioning as claimed, the mask is generated from cross-attention layer as is known in the art and while an editing mask is generated, it is not generated from a mask token taking in a mask token and word and visual embeddings as instead positions for the masks are used to guide the attention modulation with respect to a certain keyword. Thus there is not similar teaching of the mask token being used by a mask decoder along with the appropriate inputs required and furthermore, the architecture of Guo differs from that used in Wang and that claimed as the mask process is not upstream of any generation of an editing mask as the initial masks generated and extracted are used in their same format and need not be transformed to any editing mask format. Any inquiry concerning this communication or earlier communications from the examiner should be directed to SCOTT E SONNERS whose telephone number is (571)270-7504. The examiner can normally be reached Mon-Friday 9-5. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Xiao Wu can be reached at (571) 272-7761. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /SCOTT E SONNERS/Examiner, Art Unit 2613 /XIAO M WU/Supervisory Patent Examiner, Art Unit 2613 1 Wang Q, Zhang B, Birsak M, Wonka P. Instructedit: Improving automatic masks for diffusion-based image editing with user instructions. arXiv preprint arXiv:2305.18047. 2023 May 29. 2 Zou X, Yang J, Zhang H, Li F, Li L, Wang J, Wang L, Gao J, Lee YJ. Segment everything everywhere all at once. Advances in neural information processing systems. 2023 Dec 15;36:19769-82. 3 Nguyen TT, Ren Z, Pham T, Huynh TT, Nguyen PL, Yin H, Nguyen QV. Instruction-guided editing controls for images and multimedia: A survey in llm era. arXiv preprint arXiv:2411.09955. 2024 Nov 15. 4 Liu S, Zeng Z, Ren T, Li F, Zhang H, Yang J, Jiang Q, Li C, Yang J, Su H, Zhu J. Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection. https://arxiv.org/abs/2303.05499v5. 2024 July 19. 5 Kirillov A, Mintun E, Ravi N, Mao H, Rolland C, Gustafson L, Xiao T, Whitehead S, Berg AC, Lo WY, Dollár P. Segment Anything. arXiv preprint arXiv:2304.02643. 2023 Apr 5. 6 Guo Q, Lin T. Focus on Your Instruction: Fine-grained and Multi-instruction Image Editing by Attention Modulation. arXiv preprint arXiv:2312.10113. 2023 Dec 15.
Read full office action

Prosecution Timeline

Nov 06, 2024
Application Filed
Jul 28, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12743819
TRAINING-FREE CONSISTENT TEXT-TO-IMAGE GENERATION
2y 4m to grant Granted Sep 22, 2026
Patent 12743793
CAMERA TRACKING VIA DYNAMIC PERSPECTIVES
2y 0m to grant Granted Sep 22, 2026
Patent 12737967
METHOD AND APPARATUS FOR REPRESENTING DYNAMIC NEURAL RADIANCE FIELDS FROM UNSYNCHRONIZED VIDEOS
2y 1m to grant Granted Sep 15, 2026
Patent 12718418
METHOD FOR DECODING 3D CONTENT, ENCODER, AND DECODER
2y 5m to grant Granted Aug 25, 2026
Patent 12711590
Automated Clash Detection Using Two-Dimensional Drawings
2y 6m to grant Granted Aug 18, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
69%
Grant Probability
81%
With Interview (+12.2%)
3y 3m (~1y 4m remaining)
Median Time to Grant
Low
PTA Risk
Based on 392 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month