Prosecution Insights
Last updated: October 02, 2026
Application No. 18/527,016

IMPLEMENTING DIALOG-BASED IMAGE EDITING

Non-Final OA §103
Filed
Dec 01, 2023
Examiner
TSWEI, YU-JANG
Art Unit
2614
Tech Center
2600 — Communications
Assignee
Lemon Inc.
OA Round
3 (Non-Final)
84%
Grant Probability
Favorable
3-4
OA Rounds
0m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 84% — above average
84%
Career Allowance Rate
388 granted / 464 resolved
+21.6% vs TC avg
Strong +16% interview lift
Without
With
+16.0%
Interview Lift
resolved cases with interview
Typical timeline
2y 3m
Avg Prosecution
44 currently pending
Career history
507
Total Applications
across all art units

Statute-Specific Performance

§101
5.9%
-34.1% vs TC avg
§103
72.8%
+32.8% vs TC avg
§102
6.0%
-34.0% vs TC avg
§112
7.4%
-32.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 464 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . This action is in response to the Amendment filed on 4/21/2026. Claims 1-20 are pending. Claims 1, 12, 17 have been amended. Continued Examination Under 37 CFR 1.114 A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 5/21/2026 has been entered. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1, 4, 12, 14, 17, 19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Wei et al. (“DialogPaint A Dialog-based Image Editing Model”, 20231018, arXiv, hereinafter Wei) in view of Song et al. (“LLM-Planner Few-Shot Grounded Planning for Embodied Agents with Large Language”, 2023, ICCV, hereinafter Song) and further in view of Yamaura et al. (US 6,590,584 B1, hereinafter Yamaura). Regarding Claim 12, Wei teaches a system, comprising: at least one processor; and at least one memory comprising computer-readable instructions that upon execution by the at least one processor cause the system to perform operations comprising (Wei, Page 6, Section 4.1 Experimental Setup, " we initialized our model using the weights from Instruct-Pix2Pix…The model was adapted to our image editing dataset using 8 Nvidia Tesla A100 40G GPUs" indicating execution by processors; “we initialized our model using the weights from Instruct-Pix2Pix” implying stored model weights/instructions in memory that are loaded and executed; Page 1, Abstract, "We introduce … a novel framework that bridges conversational interactions with image editing, enabling users to modify images through natural dialogue): receiving text indicating a task of editing an image (Wei, Page 1, Abstract, "enabling users to modify images through natural dialogue"; Page 5, Section 3.2.1, “Given a natural language prompt describing an image and an editing task”); generating a list of one or more objects and one or more attributes associated with each of the one or more objects based on the text and the image, wherein the one or more objects are comprised in the image (Wei, Page 6, Section 3.2.2, "The textual instruction T is encoded through CLIP to obtain a latent vector representation CT. Concurrently, the input image I is processed through an encoder to derive its latent representation CI"; Page 7, Section 4.2, "the underlying architecture and training strategy of DialogPaint, which emphasizes context preservation and object-specific edits”), determining operations to be performed on each of the one or more objects, (Wei, Abstract, "handling tasks such as object replacement, style transfer, and color modification"; Page 6, Section 3.2.1, "The model's capability to generate such precise instructions is crucial, as it bridges the gap between user intent and the subsequent image editing operations performed by the Image Editing Model"),), [[ determining an order of performing the operations on an object-by-object basis ]]; generating a plan of implementing the task based on the text and the order of performing the operations, wherein the plan comprises information indicating a set of algorithm tools selected for the task (Wei, Page 4, Section 3.2, "Subsequently, the Image Editing Model is invoked to perform image editing based on the explicit textual instructions derived from the dialogue interactions." Page 6, Section 4.1 Experimental Setup, "employs the Stable Diffusion architecture, leveraging the foundational principles from lnstructPix2Pix, to execute image editing based on the explicit textual instructions <read on algorithm tools>"); generating an edited image based at least in part on the plan (Wei, Page 1, Abstract, "employs these instructions, along with the input image, to produce the desired output <read on edited image>" Page 6, Section 4.2 Results, "the model edits images according to dialog instructions."); [[ and storing the plan, wherein the plan is configured to be accessed to perform another task of editing a different image]]. But Wei does not explicitly disclose determining an order of performing the operations on an object-by-object basis. However, Song teaches generating a list of objects (Song, Page 2099, “we add the list of objects perceived in the environment so far into the prompt”; Page 3002, “The prompt begins with an intuitive explanation of the task and the list of allowable high-level actions” “further constrain the output space of the LLM to the allowed set of actions and objects.”); determining operations to be performed on each of the objects (Song, Page 2999, Section 1, “We use LLMs to generate high-level plans (HLPs), i.e., a sequence of subgoals (e.g., [Navigation potato, Pickup potato…”]); determining an order of performing the operations on an object-by-object basis (Song, Page 2099, Section 1, “We use LLMs to generate high-level plans (HLPs), i.e., a sequence <read on order> of subgoals … that the agent needs to achieve, in the specified order”); generating a plan of implementing the task based on the text and the order of performing the operations, wherein the plan comprises information indicating a set of algorithm tools selected for the task (Song, Page 3002, Section 4.2 Prompt Design, "The prompt begins with an intuitive explanation of the task and the list of allowable high-level actions <read on algorithm tools> "; Page 3002, Section 4.1, "We also use logit biases to further constrain the output space of the LLM to the allowed set of actions and objects"; Page 3002, Section 4.1, "generate high-level plans (HLPs) ... as an ordered list of actions/subgoals"); determining an order of performing the operations on an object-by-object basis (Song, Page 3002, Section 4.1, "a sequence <read on order> of subgoals ... in the specified order"; Page 2999, Section 1, "we add the list of objects perceived in the environment so far into the prompt... A new continuation ... will be generated ... based on the observed objects"). Song and Wei are analogous since both are dealing with natural language driven task execution grounded in perceptual entities leading to ordered actions that operate on visual inputs. Wei provided a way of dialog-based image editing that maps natural language to editing actions and produces edited images. Song provided a way of generating ordered, grounded plans that select actions from a set of tools/skills and execute them by using explicit object lists and ordered objected-based plans for selectable actions. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to incorporate ordered plan generation and explicit tool-set selection taught by Song into modified invention of Wei such that [a system with at least one processor and memory forms an explicit plan that orders per-object operations and identifies which algorithm tools to use, then executes to generate the edited image to improve multi-step task reliability and to provide explicit ordered execution of object edits which will enhance interpretability of editing operations. The motivation is to ensure reliable, multi-step task execution grounded to detected objects and available tools, improving interpretability and performance which is discussed by Song in Section 1 and Section 3. The combination of Wei, Song and Yamaura does not explicitly disclose storing the plan, wherein the plan is configured to be accessed to perform another task of editing a different image. However, Yamaura teaches storing the plan (Yamaura, Column 7, Line 12-17, "the command string registerer 14 sequentially stores a series of editing commands designated from the user during a period of editing one image, and manages them as macros in a format in correspondence with associated information (to be described later) related to the edited images"; Column 10, Line 34-37, "When the command string is registered, the command string registerer 14 stores the series of editing commands used for the input image at step S14, as a macro “command string”, with the scan parameters set at step S11"), wherein the plan is configured to be accessed (Yamaura, Column 11, Line 24-29, "In the parameter designation area of the copy start image upon selection of the copy based on editing history, the associated information of registered command strings, are displayed, as buttons, in list form as shown in FIG. 8 (step S21)"; Column 11, Line 33-35, "The user can designate a desired command string by the click or touch of a button with a serial number on the parameter designation area (step S22)"), to perform another task of editing a different image (Yamaura, Column 12, Line 11-14, "The input images of the respective original pages are sequentially transferred, and the controller 11 sequentially applies command strings designated at step S24 to these images (step S28)"; Column 12, Line 22-26 "According to the present invention, plural editing commands are handled as a macro, and the editing operation can be repeatedly applied to a subsequent edited image simply by calling the macro"; Column 13, Line 48-52, "In some document file generation, a series of editing operations performed in the past is to be applied to plural image information sequentially input via e.g. a scanner, with almost no change (or only by changing values)"). Yamaura and Wei are analogous since both concern image-editing workflows in which a sequence of editing operations derived from user input is applied to an image. Wei provides dialog-driven generation and execution of an image-editing sequence on a single image, and Song adds ordered object-grounded planning of the operations. Yamaura provides an explicit mechanism for persisting the resulting sequence of editing operations as a stored command string (macro) associated with the image it was derived from, and for later accessing that stored sequence via a selectable list so that it can be applied to a subsequent, different image. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Yamaura's command-string storage and later-access mechanism into the modified Wei + Song system such that the plan produced by Wei + Song for editing one image is stored and configured to be accessed to perform another task of editing a different image. The motivation is expressly disclosed by Yamaura, which states "the editing operation can be repeatedly applied to a subsequent edited image simply by calling the macro" and "the usability can be greatly improved," providing predictable improvements in efficiency, consistency of edits across images, and reduction of repetitive user input. Regarding Claim 1, it recites limitations similar in scope to the limitations of Claim 12 but as a method and the combination of Wei, Song and Yamaura teaches all the limitations as of Claim 12. Therefore, it is rejected under the same rationale. Regarding Claim 4, the combination of Wei, Song and Yamaura teaches the invention in Claim 1. The combination further teaches further comprising: determining a plurality of visual algorithms based on the image and the task; and generating the list of objects (Wei, Page 6, Section 3.2.2, " Our approach employs the Stable Diffusion architecture, leveraging the foundational principles from InstructPix2Pix, to execute image editing based on the explicit textual instructions Page 4, Section 3.2.2, " The textual instruction T is encoded through CLIP to obtain a latent vector representation CT . Concurrently, the input image I is processed through an encoder to derive its latent representation CI") Wei does not explicitly disclose but Song teaches the attributes associated with each of the objects using the plurality of visual algorithms (Song, Page 3002, Section 4.2, "The prompt begins with an intuitive explanation of the task and the list of allowable high-level actions <read on plurality of ... tools/algorithms>" Page 2999, Section 1, "We add the list of objects perceived in the environment so far into the prompt”). Song and Wei are analogous since both are dealing with natural language driven task execution grounded in perceptual entities leading to ordered actions that operate on visual inputs. Wei provided a way of dialog-based image editing that maps natural language to editing actions and produces edited images. Song provided a way of using different actions and tasks to the prompt during the process. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to incorporate action and tasks taught by Song into modified invention of Wei such that when selecting among multiple algorithms (Stable Diffusion, lnstructPix2Pix) can be based on the image and task and organizes operation per object, a predictable integration within system architecture to improve selection. Regarding Claim 14, it recites limitations similar in scope to the limitations of Claim 4 and therefore is rejected under the same rationale. Regarding Claim 17, it recites limitations similar in scope to the limitations of claim 12 and the combination of Wei, Song and Yamaura teaches all the limitations as of Claim 12. And Wei discloses these features can be implemented on a computer-readable storage medium (Wei, Page 7, “we initialized our model using the weights from Instruct- Pix2Pix”; Page 6, “the model was trained on 8 Nvidia Tesla A100 40G GPUs; Page 9, “We constructed a unique dataset containing both dialogue and image editing samples, which played a pivotal role in training our model to understand and execute user instructions effectively”; it is noted the 8 Nvidia Tesla A100 40G GPUs have massive combined memory which can store the instructions for execution). Regarding Claim 19, it recites limitations similar in scope to the limitations of Claim 4 and therefore is rejected under the same rationale. Claim(s) 2 is/are rejected under 35 U.S.C. 103 as being unpatentable over Wei et al. (“DialogPaint A Dialog-based Image Editing Model”, 20231018, arXiv, hereinafter Wei) in view of Song et al. (“LLM-Planner Few-Shot Grounded Planning for Embodied Agents with Large Language”, 2023, ICCV, hereinafter Song) and further in view of Yamaura et al. (US 6,590,584 B1, hereinafter Yamaura) as applied to Claim 1 above and further in view of Aziz et al. (US 20240045780 A1, hereinafter Aziz). Regarding Claim 2, the combination of Wei, Song and Yamaura teaches the invention in Claim 1. The combination further teaches comprising: causing to display at least one sentence in natural language in response to receiving the text (Wei, Page 6, Section 3.2.1, “Given a natural language prompt describing an image and an editing task, the model is trained to generate a series of dialogue responses” Page 1, Abstract , "modify images through natural dialogue < read on sentence in natural language>", Page 3, Section 3.1 Framework Overview, references to "dialogue interactions" where the system engages users via text messages during the editing workflow < read on causing to display at least one sentence in natural language in response to receiving the text>”). The combination does not explicitly disclose but Aziz teaches the at least one sentence configured to guide a user to upload the image (Aziz, Paragraph [0109], “Popup window 1000 may include additional features such as a button to silence notifications from the robot, a button for the user to select only those notifications that the user deems important, or an upload button that allows the user to upload a media file instructing” Aziz and Wei are analogous since both are dealing with the Large Language Modeling process. Wei provided a way of dialog-based image editing that maps natural language to editing actions and produces edited images. Aziz provided a way of allowing user to upload the media during the LLM process. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to incorporate user guided media upload taught by Aziz into modified invention of Wei such that during the image editing, system will be able to guide the user to upload the media in order for further process which will provide more intuitive and user friendly editing environment. Claim(s) 3, 13, 18 is/are rejected under 35 U.S.C. 103 as being unpatentable over Wei et al. (“DialogPaint A Dialog-based Image Editing Model”, 20231018, arXiv, hereinafter Wei) in view of Song et al. (“LLM-Planner Few-Shot Grounded Planning for Embodied Agents with Large Language”, 2023, ICCV, hereinafter Song) and further in view of Yamaura et al. (US 6,590,584 B1, hereinafter Yamaura) as applied to Claim 1, 12 above and further in view of Hsu et al. (US 7177798 B2, hereinafter Hsu). Regarding Claim 3, the combination of Wei, Song and Yamaura teaches the invention in Claim 1. The combination further teaches further comprising: causing to display at least one sentence in natural language based on determining additional information is needed to complete the task (Wei, Page 3, Section 3.1 Framework Overview, references to "dialogue interactions" where the system issues messages during the interaction when more detail is needed to proceed <read on determining additional information is needed> and those messages are in text form <read on causing to display at least one sentence in natural language> that function to "request" specifics < read on request a user to input the additional information>) Page 1, Abstract, "dialog <read on sentence in natural language>")., The combination does not explicitly disclose but Hsu teaches the at least one sentence configured to request a user to input the additional information (Hsu, Column 10, Line 10-11, “system 101 may prompt the user to "Please enter a search (natural language or keyword)”). Hsu and Wei are analogous since both are dealing with natural language driven task execution grounded in perceptual entities leading to ordered actions that operate on visual inputs. Wei provided a way of dialog-based image editing that maps natural language to editing actions and produces edited images. Hsu provided a way of generating next action based on user input natural language of additional information Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to incorporate ordered user input of additional information taught by Hsu into modified invention of Wei such that when dealing with natural language driven task, system will be able to provide user with chance to input their request through input prompt which will add additional function of interactive process and increase the efficiency to generate the edited image. Regarding Claim 13, it recites limitations similar in scope to the limitations of Claim 3 and therefore is rejected under the same rationale. Regarding Claim 18, it recites limitations similar in scope to the limitations of Claim 3 and therefore is rejected under the same rationale. Claim(s) 5, 6, 9, 10, 15, 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Wei et al. (“DialogPaint A Dialog-based Image Editing Model”, 20231018, arXiv, hereinafter Wei) in view of Song et al. (“LLM-Planner Few-Shot Grounded Planning for Embodied Agents with Large Language”, 2023, ICCV, hereinafter Song) and further in view of Yamaura et al. (US 6,590,584 B1, hereinafter Yamaura) as applied to Claim 1 and further in view of Shen et al. (“HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face”, 2023, hereinafter Shen). Regarding Claim 5, the combination of Wei, Song and Yamaura teaches the invention in Claim 1. The combination does not explicitly disclose but Shen teaches generating textual descriptions corresponding to algorithm tools based on specifications of the algorithm tools (Shen, §3 .2 Model Selection, "we use model descriptions as the language interface lo connect each model ... we first gather the descriptions of expert models from the ML community (e. g. , Hugging Face)"), and wherein a description corresponding to each algorithm tool comprises partial information about each algorithm tool (Shen, §5 Limitations, "maximum token length is always limited ... how to briefly and effectively summarize model descriptions is also worthy of exploration"), and a specification comprises complete information about each algorithm tool (Shen, Appx. A.1.2 Model Descriptions, " These descriptions encompass various aspects of the model, such as its function, architecture, supported languages and domains, licensing, and other relevant details. These comprehensive model descriptions play a crucial role "). Shen and Wei are analogous since both deal with LLM-mediated orchestration of external models via a language interface. Wei provided dialog-based editing and execution of image models; Shen provided description-from-specification and summarization for token-limited LLM inputs. Therefore, it would have been obvious to one of ordinary skill to incorporate Shen's description-generation scheme into Wei so that each tool has a complete specification and a short text description. The motivation is to enable an LLM controller to reason over tool capabilities while respecting context limits (HuggingGPT §3 .2 and §5). Regarding Claim 6, the combination of Wei, Song and Yamaura and Shen teaches the invention in Claim 5. The combination further teaches inputting the descriptions into a large language model (Shen, §3.2 Model Selection,"available models are presented as options within a given context ... HuggingGPT is able to select the most appropriate model"), in response to determining that a total size of the descriptions is less than or equal to an input limit of the large language model (Shen, Page 6, “Due to the limits of maximum context length, it is not feasible to encompass the information of all relevant models within one prompt”); and selecting one or more algorithm tools related to the task based on the descriptions (Shen, Introduction, "the LLM acts as a controller to manage Al models ... select models according to their function descriptions and execute each subtask'') . Shen and Wei are analogous since both involve LLM-driven model control for image editing. Wei provided dialogue input and model execution; Shen provided feeding model descriptions into an LLM and selection from them. Therefore, it would have been obvious to incorporate Shen's description-based tool selection into Wei so that the LLM receives tool descriptions and selects the appropriate editing algorithm(s). The motivation is to improve automatic matching between user instruction and tool choice (HuggingGPT §3.2, Intro). Regarding Claim 9, the combination of Wei, Song and Yamaura teaches the invention in Claim 1. The combination does not explicitly disclose but Shen teaches generating information indicating a particular object in the list to which each of the set of algorithm tools is applied (Shen 's task-parsing/specification and resource-dependency records constitute the claimed " information": "a standardized template for tasks ... ' task', ' id', 'dep', and ' args"' (template; <read on information>) and "resource dependency ... sets this symbol (i.e., <resource>-task_id) to the corresponding resource subfield in the arguments ... dynamically replaces this symbol with the resource generated by the prerequisite task" ( <read on indicating a particular object in the list ... to which each ... tool is applied>); the worked example shows object detection with "bounding box" and a chain of selected models (<read on set of algorithm tools applied to objects)), Shen and Wei are analogous- both orchestrate LLM-driven pipelines that select and apply tools/models to visual tasks. Wei provides dialogue-driven image editing with object-specific edits; Shen provides an explicit task/argument/resource record that captures which tool acts on which detected object (e.g., bounding boxes) in a plan. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to incorporate Shen's task-specification and resource-dependency recording into Wei so that DialogPaint generates information indicating a particular object for which each tool is applied, improving traceability and determinism of multi-tool edits. Regarding Claim 10, the combination of Wei, Song and Yamaura teaches the invention in Claim 1. The combination further teaches further comprising: generating executable code based at least in part on the plan, wherein the generating executable code based at least in part on the plan (Wei, Page 1, Abstract, "employs these instructions, along with the input image, to produce the desired output <read on edited image>" Page 6, Section 4.2 Results, "the model edits images according to dialog instructions.") The combination does not explicitly disclose but Shen teaches generating the executable code based on a complete specification corresponding to each of the set of algorithm tools (Shen, §5 Limitations, "maximum token length is always limited ... how to briefly and effectively summarize model descriptions is also worthy of exploration" Appx. A.1.2 Model Descriptions, " These descriptions encompass various aspects of the model, such as its function, architecture, supported languages and domains, licensing, and other relevant details. These comprehensive model descriptions play a crucial role"; Page 3000, “we show that LLMPlanner can generate complete and high-quality high-level plans that are grounded in the current environment with a fraction of labeled data.”). Shen and Wei are analogous since both deal with LLM-mediated orchestration of external models via a language interface. Wei provided dialog-based editing and execution of image models; Shen provided a LLM-Planner with complete code generation for set of tools. Therefore, it would have been obvious to one of ordinary skill to incorporate Shen's description-generation scheme into Wei so that each tool has a complete specification and a short text description which will enable an LLM controller to reason over tool capabilities. Regarding Claim 15, the combination of Wei, Song and Yamaura and Shen teaches the invention in Claim 12. The combination further teaches inputting the descriptions into a large language model (Shen, §3.2 Model Selection,"available models are presented as options within a given context ... HuggingGPT is able to select the most appropriate model"), and selecting one or more algorithm tools related to the task based on the descriptions (Shen, Introduction, "the LLM acts as a controller to manage Al models ... select models according to their function descriptions and execute each subtask''). generating descriptions corresponding to algorithm tools based on specifications of the algorithm tools, wherein a description corresponding to each algorithm tool (Shen, §3 .2 Model Selection, "we use model descriptions as the language interface lo connect each model ... we first gather the descriptions of expert models from the ML community (e. g. , Hugging Face)"), comprises partial information about each algorithm tool (Shen, §5 Limitations, "maximum token length is always limited ... how to briefly and effectively summarize model descriptions is also worthy of exploration"), and a specification comprises complete information about each algorithm tool (Shen, Appx. A.1.2 Model Descriptions, " These descriptions encompass various aspects of the model, such as its function, architecture, supported languages and domains, licensing, and other relevant details. These comprehensive model descriptions play a crucial role "); inputting the descriptions into a large language model (Shen, §3.2 Model Selection,"available models are presented as options within a given context ... HuggingGPT is able to select the most appropriate model") in response to determining that a total size of the descriptions is less than or equal to an input limit of the large language model (Shen, Page 6, “Due to the limits of maximum context length, it is not feasible to encompass the information of all relevant models within one prompt”); and selecting one or more algorithm tools related to the task of editing the image based on the descriptions (Shen, Introduction, "the LLM acts as a controller to manage Al models ... select models according to their function descriptions and execute each subtask''). Shen and Wei are analogous since both involve LLM-driven model control for image editing. Wei provided dialogue input and model execution; Shen provided feeding model descriptions into an LLM and selection from them. Therefore, it would have been obvious to incorporate Shen's description-based tool selection into Wei so that the LLM receives tool descriptions and selects the appropriate editing algorithm(s). The motivation is to improve automatic matching between user instruction and tool choice (HuggingGPT §3.2, Intro). Regarding Claim 20, it recites limitations similar in scope to the limitations of Claim 15 and therefore is rejected under the same rationale. Claim(s) 7, 16 is/are rejected under 35 U.S.C. 103 as being unpatentable over Wei et al. (“DialogPaint A Dialog-based Image Editing Model”, 20231018, arXiv, hereinafter Wei) in view of Song et al. (“LLM-Planner Few-Shot Grounded Planning for Embodied Agents with Large Language”, 2023, ICCV, hereinafter Song), and further in view of Yamaura et al. (US 6,590,584 B1, hereinafter Yamaura) and Shen et al. (“HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face”, 2023, hereinafter Shen) as applied to Claim 5 above and further in view of Wu et al. (“Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models”, hereinafter Wu). Regarding Claim 7, the combination of Wei, Song and Yamaura and Shen teaches the invention in Claim 5. The combination further teaches dividing the algorithm tools into a plurality of batches . .. input limit ... (Shen, §3.2 In-context Task-model Assignment, "due to the limits of maximum context length ... not feasible to encompass ... within one prompt. To mitigate this issue, we first filter ... then select the top-K models as the candidates ... substantially reduce the token usage in the prompt"); and selecting one or more algorithm tools in each ... based on the descriptions (Shen, §3.2, "select the most appropriate model for each task ... use model descriptions as the language interface") Shen and Wei are analogous since both concern LLM-driven model orchestration under context length limits. Wei provided dialog-editing without input-limit handling; Shen provided subset (top-K) prompt filtering for token limits; Therefore, it would have been obvious to one of ordinary skill to incorporate Shen's techniques into Wei so that the system divides tool descriptions into manageable batches, inputs them sequentially, and selects tools per batch. The motivation is to ensure LLM processing within token limits while preserving accurate tool selection. The combination does not explicitly disclose but Wu teaches sequentially inputting descriptions ... input limit of a large language model (Wu, §3, "we truncate the dialogue history with a maximum length threshold to meet the input length of ChatGPT model" ), and describes a Prompt Manager that "records the interaction history and controls the inputs/outputs of various visual foundation models" (§3 Prompt Manager ). Wu and Wei are analogous since both deal with dialogue-based multi-model image editing and explicit input-length management. Wei provided dialog-editing without input-limit handling; Wu provided history truncation and prompt manager sequencing. Therefore, it would have been obvious to one of ordinary skill to incorporate Wu's techniques into Wei so that the system can adjust the data based on the limit for inputs when selects tools per batch which maintain the suitability of the system. Regarding Claim 16, the combination of Wei, Song and Yamaura and Shen teaches the invention in Claim 12. The combination does not explicitly disclose but Shen teaches generating descriptions corresponding to algorithm tools based on specifications of the algorithm tools (Shen, §3 .2 Model Selection, "we use model descriptions as the language interface lo connect each model ... we first gather the descriptions of expert models from the ML community (e. g. , Hugging Face)"),, wherein a description corresponding to each algorithm tool comprises partial information about each algorithm tool, and a specification comprises complete information about each algorithm tool (Shen, §5 Limitations, "maximum token length is always limited ... how to briefly and effectively summarize model descriptions is also worthy of exploration"), and a specification comprises complete information about each algorithm tool (Shen, Appx. A.1.2 Model Descriptions, " These descriptions encompass various aspects of the model, such as its function, architecture, supported languages and domains, licensing, and other relevant details. These comprehensive model descriptions play a crucial role "); dividing the algorithm tools into a plurality of batches . .. input limit ... (Shen, §3.2 In-context Task-model Assignment, "due to the limits of maximum context length ... not feasible to encompass ... within one prompt. To mitigate this issue, we first filter ... then select the top-K models as the candidates ... substantially reduce the token usage in the prompt")' and selecting one or more algorithm tools in each ... based on the descriptions (Shen, §3.2, "select the most appropriate model for each task ... use model descriptions as the language interface"). Shen and Wei are analogous since both concern LLM-driven model orchestration under context length limits. Wei provided dialog-editing without input-limit handling; Shen provided subset (top-K) prompt filtering for token limits; Therefore, it would have been obvious to one of ordinary skill to incorporate Shen's techniques into Wei so that the system divides tool descriptions into manageable batches, inputs them sequentially, and selects tools per batch. The motivation is to ensure LLM processing within token limits while preserving accurate tool selection. The combination does not explicitly disclose but Wu teaches the combination does not explicitly disclose but Wu teaches sequentially inputting descriptions ... input limit of a large language model (Wu, §3, "we truncate the dialogue history with a maximum length threshold to meet the input length of ChatGPT model" ), and describes a Prompt Manager that "records the interaction history and controls the inputs/outputs of various visual foundation models" (§3 Prompt Manager ). Wu and Wei are analogous since both deal with dialogue-based multi-model image editing and explicit input-length management. Wei provided dialog-editing without input-limit handling; Wu provided history truncation and prompt manager sequencing. Therefore, it would have been obvious to one of ordinary skill to incorporate Wu's techniques into Wei so that the system can adjust the data based on the limit for inputs when selects tools per batch which maintain the suitability of the system. Claim(s) 8 is/are rejected under 35 U.S.C. 103 as being unpatentable over Wei et al. (“DialogPaint A Dialog-based Image Editing Model”, 20231018, arXiv, hereinafter Wei) in view of Song et al. (“LLM-Planner Few-Shot Grounded Planning for Embodied Agents with Large Language”, 2023, ICCV, hereinafter Song) and further in view of Yamaura et al. (US 6,590,584 B1, hereinafter Yamaura) as applied to Claim 1 respectively and further in view of Masuoka et al. ( US 20040230636 A1, hereinafter Masuoka). Regarding Claim 8, the combination of Wei, Song and Yamaura teaches the invention in Claim 1. The combination does not explicitly disclose but Shen teaches establishing a mapping relationship between a plurality of editing tasks and a plurality of sets of algorithm tools selected for the plurality of editing tasks; and determining, based on the mapping relationship, a set of selected algorithm tools related to one of the plurality of editing tasks [[in response to receiving an editing task that is the same or similar to the one of the plurality of editing tasks]]" (Shen, Page 4, "workflow includes four stages: task planning, model selection, task execution, and response generation" <read on mapping/selection>; Page 14, "To format the parsed task, we define the template [{"task": task, "id", task_id, "dep": dependency task_ids, "args":" <read on mapping relationship>; Page 6, "selecting the most appropriate model for each task in the parsed task list" <read on determining ... Jet of selected algorithm tools>”). Shen and Wei are analogous since both deal with LLM-mediated orchestration of external models via a language interface. Wei provided dialog-based editing and execution of image models; Shen provided a way of adapt mapping relationship to the task during the LLM data processing. Therefore, it would have been obvious to one of ordinary skill to incorporate Shen's mapping relationship into Wei so that system will be able to create a concrete ask-tool assignment. The combination does not explicitly disclose but Masuoka teaches in response to receiving an editing task that is the same or similar to the one of the plurality of editing tasks (Masuoka, Paragraph [0053], "if the user needs to perform the same or similar task in the future, she will have to do all of that again"); Masuoka and Wei are analogous since both are dealing with user-driven, multi-step task execution pipelines that interpret user intent and orchestrate a sequence of operations/tools to accomplish the task. Wei provided a way of performing dialogue-based image editing by interpreting user instructions and invoking selected editing operations (algorithm tools) within a plan. Masuoka provided a way of recognizing and handling recurring tasks by expressly addressing when a later task is the "same or similar" to a previous task and motivating reuse of prior task configurations/steps to avoid re-specification (e.g., " if the user needs to perform the same or similar task in the future, she will have to do all of that again. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to incorporate the reuse trigger taught by Masuoka into the modified invention of Wei such that in response to receiving an editing task that is the same or similar to a previously handled editing task, the system consults the stored mapping relationship and determines the set of selected algorithm tools for that task. Claim(s) 11 is/are rejected under 35 U.S.C. 103 as being unpatentable over Wei et al. (“DialogPaint A Dialog-based Image Editing Model”, 20231018, arXiv, hereinafter Wei) in view of Song et al. (“LLM-Planner Few-Shot Grounded Planning for Embodied Agents with Large Language”, 2023, ICCV, hereinafter Song) and further in view of Yamaura et al. (US 6,590,584 B1, hereinafter Yamaura) as applied to Claim 1 and further in view Chan et al. (US 20180081417 A1, hereinafter Chan). Regarding Claim 11, the combination of Wei, Song and Yamaura teaches the invention in Claim 1. The combination does not explicitly disclose but Chan teaches sharing or storing the plan; uploading the plan to a server computing system; or exporting the plan to another platform for creating an effect in the another platform (Chan, Paragraph [0065], “,planning program 200 is offered to a user as-a-service that includes a sharing of plans of activities” [0030], “UI 122 receives input in response to a user of device 120 utilizing natural language, such as written words or spoken words, that device 120 identifies as information and/or commands” [0028], “device 120 may upload one or more plans of activities from user plans 126 to system 102 on a periodic basis and/or as dictated by a user of device 120”). Chan and Wei are analogous since both are dealing with natural language modeling process. Wei provided a way of dialog-based image editing that maps natural language to editing actions and produces edited images. Chan provided a way of allow user to sharing and transferring data in between different system with natural language data. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to incorporate data sharing between systems taught by Chan into modified invention of Wei such that during the image editing, system will be able to allow user to input the natural language input and sharing in between different system and/or platform which increase the flexibility of the system. Response to Arguments Applicant’s arguments with respect to claim 1, 12, 17, filed on 4/21/2026, with respect to rejection under 35 USC § 103 regarding prior art combination does not teaches the limitation “"storing the plan, wherein the plan is configured to be accessed to perform another task of editing a different image" have been considered but is in moot. They are now taught by the combination of Wei, Song and Yamaura. Applicant’s arguments with respect to claim 1, 12, 17, filed on 4/21/2026, with respect to rejection under 35 USC § 103 regarding prior art combination does not teaches the limitation "generating a list of one or more objects and one or more attributes associated with each of the one or more objects based on the text and the image, wherein the one or more objects are comprised in the image." Have been considered but is not persuasive. In response to the argument, as described in the rejection of Claim 1 above, this limitation is taught by the combination of Wei and Song. In particular, Wei in Page 7, Section 4.2 teaches "the underlying architecture and training strategy of DialogPaint, which emphasizes context preservation and object-specific edits," which shows Wei's system identifies specific objects in the image and their editable characteristics (i.e., attributes such as color, style, expression) in order to perform "object replacement, style transfer, and color modification" (Wei, Page 1, Abstract). Wei further teaches in Page 6, Section 3.2.2 that "The textual instruction T is encoded through CLIP to obtain a latent vector representation CT. Concurrently, the input image I is processed through an encoder to derive its latent representation CI," demonstrating that both the text and the image are jointly analyzed to identify what objects to edit and what attributes of those objects to modify. Additionally, Song expressly teaches generating a list of objects perceived from the visual input for use in planning, stating in Page 2999, Section 1: "For grounding, we add the list of objects perceived in the environment so far into the prompt as a simple but effective description of the current environment." Song further teaches at Page 3002, Section 4.2 that "The prompt begins with an intuitive explanation of the task and the list of allowable high-level actions," and at Page 3002, Section 4.1 that "We also use logit biases to further constrain the output space of the LLM to the allowed set of actions and objects." Song therefore explicitly generates and uses a list of objects together with characterizing information (allowable actions applicable to each object, object status/state in the environment) that constitutes the "attributes associated with each of the objects" required by the claim. The claim recites "one or more objects and one or more attributes associated with each of the one or more objects," which is satisfied when the combined system identifies at least one object and at least one attribute associated therewith. Wei's object-specific editing (which necessarily identifies object-level characteristics such as color, expression, and style in order to modify them) combined with Song's explicit list of perceived objects annotated with allowable actions/states, together generate the claimed list of objects and their attributes. The amendment inserting "one or more" before "objects" and "attributes" does not narrow the claim in a way that avoids the combination, since a list of objects and attributes generated by the modified Wei + Song system inherently reads on a list of "one or more" objects and "one or more" attributes. In regard to Claims 2-11, 13-16, 18-20, they directly/indirectly depend on independent Claim 1, 12, 17 respectively. Applicant does not argue anything other than the independent claim 1, 12, 17. The limitations in those claims in conjunction with combination previously established as explained. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. US 10282055 B2 Ordered processing of edits for a media editing application US 20050216841 A1 System and method for editing digitally represented still images Any inquiry concerning this communication or earlier communications from the examiner should be directed to YUJANG TSWEI whose telephone number is (571)272-6669. The examiner can normally be reached 8:30am-5:30pm EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kent Chang can be reached at (571)272-7667. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /YuJang Tswei/Primary Examiner, Art Unit 2614
Read full office action

Prosecution Timeline

Show 1 earlier event
Nov 05, 2025
Non-Final Rejection mailed — §103
Feb 05, 2026
Response Filed
Feb 23, 2026
Final Rejection mailed — §103
Apr 21, 2026
Response after Non-Final Action
Apr 22, 2026
Response after Non-Final Action
May 21, 2026
Request for Continued Examination
May 26, 2026
Response after Non-Final Action
Sep 10, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12749275
DIRECT MANIPULATION OF IMPLICITLY DEFINED DIGITAL 3D SHAPES
2y 4m to grant Granted Sep 29, 2026
Patent 12743679
SYSTEMS AND METHODS FOR TEMPLATE IMAGE EDITS
2y 5m to grant Granted Sep 22, 2026
Patent 12743795
Determining Object Structure Using Camera Devices With Views Of Moving Objects
2y 4m to grant Granted Sep 22, 2026
Patent 12718420
INFORMATION PROCESSING DEVICE AND METHOD
2y 2m to grant Granted Aug 25, 2026
Patent 12675993
AUGMENTED, VIRTUAL AND MIXED-REALITY CONTENT SELECTION & DISPLAY FOR BANK NOTE
4y 4m to grant Granted Jul 07, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
84%
Grant Probability
99%
With Interview (+16.0%)
2y 3m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 464 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month