DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Amendment / Arguments
Applicant amended the independent claims to add “automatically generating” at least one prompt (previously, it was “obtaining” at least one prompt). The applied reference, Wu, teaches this feature as claimed. See para. 179, for example, which teaches that based on a text input and input image, the system/methods of Wu can generate a “rewritten prompt”, which teaches “automatically generating at least one prompt”, as recited in the independent claims. The rewritten prompt is/was automatically generated.
Wu is also not limited to what types of image editing can be done via its systems and methods. For example, a lot of Wu’s figures show the following “Describe your edit” entry area, which allows users creative freedom to describe what edit they want (i.e. replace a background, for example, or other types of edits).
PNG
media_image1.png
180
216
media_image1.png
Greyscale
The remaining amendments to the independent claims are more stylistic or redundant language added to the claims (i.e. generate a replacement background image that excludes the foreground object.). As mapped in the independent claims, Wu teaches segmentation of foreground and background (i.e. para. 115), and an user image editing request to replace the background of an image, will cause the system/method of Wu to generate a new background (excluding foreground, because this object isn’t being replaced or edited). The similar addition to the “receiving” step of claim 1 is also redundant, basically re-iterating that the “replacement background image” is separate from the segmented image.
To Applicant’s arguments, Applicant argues the following (Remarks, page 9):
PNG
media_image2.png
202
684
media_image2.png
Greyscale
The examiner disagrees with Applicant’s arguments because Applicant’s claim interpretation of Applicant’s own claims is respectfully incorrect:
Applicant argues that “the claimed…model does not generate the final composite image. Instead [the model of Applicant’s independent claims’ generates an intermediate background-only image” (see above).
This is not true. Applicant’s claimed model does generate a final composite image. See Applicant’s claim 1, “generating a new composite image” step.
This is also not true for another reason: there is no step in claim 1 that generates an intermediate image. No “intermediate background image” is claimed. Instead, there is a step of receiving, from the model, the replacement background image (not displaying, not an intermediate image), and the only image that is generated for the user is the new composite image.
The steps of providing and receiving in claim 1 are model-driven steps.
Also, Applicant’s arguments in the above-referenced paragraph is respectfully somewhat contradictory. There is a sentence that begins with saying that Applicant’s claims do not generate a final composite image, and then the next sentence seems to contradict that argument and admit that a final composite image is actually generated (because, respectfully, it is, per the language claim 1).
If Applicant wants claim 1 to require an intermediate background image for display (which seems to be what Applicant’s arguments head toward), then the independent claims need to be amended to have that “intermediate” image, generated and displayed. Otherwise, Wu teaches receiving a replacement background image, from the machine learning model, for generating a new composite image.
Stated differently, there is no generating of a separate intermediate background image in Applicant’s independent claims. The only “generating” is of the new composite image. Applicant should add a step such as “generating an intermediate background image” to be consistent with Applicant’s arguments, because right now that isn’t claimed.
Finally, other of Applicant’s dependent claims do positively recite generating multiple background images (see e.g. claims 10-12, mapped in the prior office action and herein).
Accordingly, the claims stand rejected under 103.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1, 9, 10, 14, 18 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Wu (U.S. Patent App. Pub. No. 2026/0045012 A1).
Regarding claim 1:
Wu teaches: a computer-implemented method (claim 1, a computer-implemented method), comprising:
obtaining a composite image (claim 6, “receiving an initial image”) depicting an identifiable foreground object (para. 199, the initial image (input data) can have a foregoing and a background);
performing image segmentation to isolate a foreground object from the composite image, the image segmentation yielding a segmented composite image (para. 199, “segmentation of the initial image into a foreground and a background”);
determining object attributes of the foreground object based on image analysis of the segmented composite image (para. 189, determine types of objects, the “type” corresponding to object attributes based on image analysis of object recognition)
(another example teaching: para. 180, segmentation of foreground object dog, the object attributes are pixels and boundaries associated with the dog so user can edit, such as replace or move the dog object. See para. 181);
automatically generating at least one prompt for a text-to-image model using the determined object attributes of the foreground object, the at least one prompt defining an intended replacement background for the composite image (Fig. 19D: text reads “Replace the background with an autumn forest”. The image has a dog in the foreground. Regarding “automatically generating” the at least one prompt, see e.g. paras. 179, 183, the method/systems of Wu can automatically generate a “rewritten prompt” that is provided as input to a machine learning model. The “rewritten prompt” of Wu teaches Applicant’s claimed “automatically generating at least one prompt” as claimed);
providing, to the text-to-image model, instructions to generate a replacement background image that excludes the foreground object based on the at least one prompt (e.g. claim 1: “providing the request and the prompt as input to the selected machine-learning model”. A request similar to the one mapped in the “automatically generating” step, to replace the background with an autumn forest, would include instructions to generate a background image (that isn’t the autumn forest)). ** Also, Wu is not limited to what type of “image editing” can be user driven via text, or images. See Figs. 20B-20D, for example, which all show “Describe your edit”. A user providing a text description of, generate a new background, or something similar, would also satisfy this);
receiving, from the text-to-image model, the replacement background image as an image separate from the segmented composite image (claim 1: generating, by the selected machine-learning model, the output image that satisfies the request and the prompt.” This is done as part of this step; segmentation was mapped above. See also para. 115, 199); and
generating a new composite image, the generating including compositing portions of the composite image corresponding to the foreground object with the replacement background image (claim 1: “generating, by the selected machine-learning model, the output image that satisfies the request and the prompt.” See also Fig. 19F).
Accordingly, it would have been obvious for one of ordinary skill in the art to have further modified the applied reference(-s), in view of same, to have obtained the above, and the results of the modification would have been obvious and predictable to one of ordinary skill in the art as of the effective filing date of the claimed invention. See MPEP §2143(A).
The prior art included each element recited in claim 1, although not necessarily in a single embodiment, with the only difference being between the claimed element and the prior art being the lack of actual combination of certain elements in a single prior art embodiment, as described above.
One of ordinary skill in the art could have combined the elements as claimed by known methods, and in that combination, each element merely performs the same function as it does separately. One of ordinary skill in the art would have also recognized that the results of the combination were predictable as of the effective filing date of the claimed invention.
Regarding claim 9:
Wu teaches: the method of claim 1, further comprising extracting the foreground object from the segmented composite image,
wherein the new composite image is generated based on combining the extracted foreground object with the replacement background image (Figs. 19D-F, the foreground dog is extracted and a new composite image is generated with the dog and replacement background (autumn forest)).
It would have been obvious for one of ordinary skill in the art, as of the effective filing date of Applicant’s claims, to have further modified the applied reference(-s) in view of same to have obtained the above, motivated to assist with interactive editing of images.
Regarding claim 10:
Wu teaches: the method of claim 1, wherein automatically generating the at least one prompt for the text-to-image model comprises providing, to an LLM, instructions to generate one or more candidate prompts based on the object attributes of the foreground object (see paras. 206-08, a prompt engine (generated with a LLM that is part of the prompt engine), generates a rewritten prompt (“one or more candidate prompts). Wu is not particularly limited to what the initial text prompt can be; prompts and therefore rewritten or candidate prompts based on foreground (as part of the initial image) is an obvious and taught embodiment of Wu), and
wherein providing instructions to the text-to-image model to generate the replacement background image comprises: receiving, from the LLM, the generated candidate prompts (Id.); and
for each generated candidate prompt, providing, to the text-to-image model, instructions to generate a candidate background image corresponding to the candidate prompt (paras 206-19, the rewritten prompt (i.e. generated candidate prompt) is provided to generate a background image corresponding to the candidate prompt. An embodiment whereby the prompt includes background replacement, the replaced background being a “candidate background image” corresponding to the “candidate prompt”, as also mapped in claim 1 regarding background replacements, is an obvious embodiment over Wu). See also Wu, para. 11 and claim 2.
It would have been obvious for one of ordinary skill in the art, as of the effective filing date of Applicant’s claims, to have further modified the applied reference(-s) in view of same to have obtained the above, motivated to assist with interactive editing of images.
Regarding claim 14: see also claim 1.
Wu teaches: a computing system, comprising: a processor; and a memory coupled to the processor, the memory storing computer-executable instructions that, when executed by the processor, configure the processor (claim 15, system having processors, memory with instructions, to cause the processor to perform operations) to:
The instructions correspond to the method of claim 1; the same rationale for rejection applies.
Regarding claim 18: see claim 10.
These claims are similar; the same rationale for rejection applies.
Regarding claim 20: see also claim 1.
Wu teaches: a non-transitory, computer-readable medium storing instructions that, when executed by a processor, configure the processor (claim 18, non-transitory CRM with processor executable instructions) to:
The instructions correspond to the method of claim 1; the same rationale for rejection applies.
Claim(s) 2-5, 11, 15 and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Wu in view of Benedetto (U.S. Patent App. Pub. No. 2024/0193351 A1).
Regarding claim 2:
Wu teaches: the method of claim 1, wherein automatically generating the at least one prompt for the text-to-image model comprises: determining current text input in an input field of a graphical user interface (see paras. 205-06. User enters a text prompt in the user interface module 202, such as the text prompt/text input in Fig. 19D, as a non-limiting example);
providing, to a large language model (LLM), instructions to generate predictive text data based on the object attributes of the foreground object and the current text input in the input field (see paras. 206, a prompt engine (generated with a LLM that is part of the prompt engine), generates a rewritten prompt (i.e. “predictive text data” based on the prompt, or “current text input” as claimed. Wu is not particularly limited to what the initial text prompt can be; predictive text based on foreground (as part of the initial image) is an obvious and taught embodiment of Wu).
Re: presenting the predictive text data via the graphical user interface, Benedetto teaches that it is known, in the context of generating predictive text (what Benedetto calls “keyword variations”), to present them to a user interface (see claims 7, 9 and Fig. 4A-1: 106, “cloudy, rainy and night” are presented a predictive text data for the input “dark”.
Modifying Wu, in view of Benedetto such to have presented the user with predictive text (i.e. as variations or adjusted prompts), is all of taught/suggested by the prior art, and the results of the modification would have been obvious and predictable to one of ordinary skill in the art as of the effective filing date of the claimed invention. See MPEP §2143(A).
One of ordinary skill in the art could have combined the elements as claimed by known methods, and in that combination, each element merely performs the same function as it does separately. One of ordinary skill in the art would have also recognized that the results of the combination were predictable as of the effective filing date of the claimed invention.
Regarding claim 3:
Wu and/or Benedetto teach: the method of claim 2, wherein the predictive text data comprises a next word suggestion (Wu, paras. 205-06, next word as to prompt) (Benedetto, next work as to keyword variation. See Figs. 4A-1: 106, or Fig. 4A-2: 107).
It would have been obvious for one of ordinary skill in the art, as of the effective filing date of Applicant’s claims, to have further modified the applied reference(-s) in view of same to have obtained the above, motivated to assist with interactive optimization of user inputs.
Regarding claim 4:
Benedetto teaches :the method of claim 2, further comprising: receiving, via the graphical user interface, user selection of first text from the predictive text data (Fig. 4A-1: 106, user can select a predictive text); and
combining the current text input in the input field with the selected first text to obtain the at least one prompts Fig. 4A:2: user selected “night”, this selected text was combined with the current text input in the “Search Field”) See also Fig. 5A .
It would have been obvious for one of ordinary skill in the art, as of the effective filing date of Applicant’s claims, to have further modified the applied reference(-s) in view of same to have obtained the above, motivated to assist with interactive optimization of user inputs.
Regarding claim 5:
Benedetto teaches :the method of claim 2, wherein the predictive text data comprises a plurality of word suggestions and wherein each of the plurality of word suggestions is presented as a selectable option via the graphical user interface (e.g. Figs. 4A-1 and 4A-2).
It would have been obvious for one of ordinary skill in the art, as of the effective filing date of Applicant’s claims, to have further modified the applied reference(-s) in view of same to have obtained the above, motivated to assist with interactive optimization of user inputs.
Regarding claim 11:
It would have been obvious for one of ordinary skill in the art to have further modified the applied reference(-s), in view of same, to have obtained: the method of claim 10, further comprising presenting, via a graphical user interface, the candidate background images corresponding to the one or more candidate prompts as selectable options,
wherein the replacement background image comprises a selection of one of the candidate background images, and the results of the modification would have been obvious and predictable to one of ordinary skill in the art as of the effective filing date of the claimed invention. See MPEP §2143(A).
Benedetto teaches that it is known to provide users with selectable options via GUI (see Figs. 4A1-4A2, here the selectable options are related to keyword variations). Wu teaches that it is known for users to desire image modifications that include background replacement (see mapping to claim 1). Modifying the applied references, in view of same, such to include selectable options for background replacement (i.e. per Wu, selectable options of background autumn forests, as per Wu, Fig. 19D), whereby Wu also teaches GUI with selectable options, see Figs. 13A, 13B, 13C, and/or 16A, is all of taught/suggested by the prior art, and the results of the modification would have been obvious and predictable to one of ordinary skill in the art.
One of ordinary skill in the art could have combined the elements as claimed by known methods, and in that combination, each element merely performs the same function as it does separately. One of ordinary skill in the art would have also recognized that the results of the combination were predictable as of the effective filing date of the claimed invention.
Regarding claim 12:
It would have been obvious for one of ordinary skill in the art to have further modified the applied reference(-s), in view of same, to have obtained: the method of claim 10, further comprising: receiving, via the graphical user interface, a selection of one of the candidate background images and user input of modifications to the candidate prompt associated with the selected candidate background image;
providing, to the text-to-image model, instructions to generate a modified candidate background image based on the modified candidate prompt, and the results of the modification would have been obvious and predictable to one of ordinary skill in the art as of the effective filing date of the claimed invention. See MPEP §2143(A).
See mapping to claim 11 which teaches/claims/maps selectable background images. RE: user input of modifications to the candidate prompt, see Benedetto, Figs. 4A-1-4A2, selectable modifications to prompt teaches this feature. Modifying the applied references, such to include the selectable background image and selected modifications to candidate prompt, and provide them to the model of Wu, as mapped in claim 1, to generate a modified image, is all of taught/suggested by the prior art, and the results of the modification would have been obvious and predictable to one of ordinary skill in the art.
One of ordinary skill in the art could have combined the elements as claimed by known methods, and in that combination, each element merely performs the same function as it does separately. One of ordinary skill in the art would have also recognized that the results of the combination were predictable as of the effective filing date of the claimed invention.
Regarding claim 15: see claim 2.
These claims are similar; the same rationale for rejection applies.
Regarding claim 16: see claim 4.
These claims are similar; the same rationale for rejection applies.
Regarding claim 19: see claim 12.
These claims are similar; the same rationale for rejection applies.
Claim(s) 6, 7 and 13 are rejected under 35 U.S.C. 103 as being unpatentable over Wu in view of Green (U.S. Patent App. Pub. No. 2024/0282130 A1).
Regarding claim 6:
It would have been obvious for one of ordinary skill in the art to have combined and modified the applied reference(-s), in view of same, to have obtained: the method of claim 1, wherein determining the object attributes of the foreground object comprises providing at least a portion of the segmented composite image to a language model that is trained to output attribute labels of input images, and the results of the modification would have been obvious and predictable to one of ordinary skill in the art as of the effective filing date of the claimed invention. See MPEP §2143(A).
Wu teaches segmentation of images, as mapped in claim 1. Green teaches a language model (e.g. para. 44, OpenAI) trained to output attribute labels of input images (e.g. paras. 22, 28-29 and Fig. 3). Modifying the applied references, such to input the segmented image of Wu to the classifier language model of Green, is all of taught and suggested by the prior art, and would have been obvious and predictable to one of ordinary skill in the art.
One of ordinary skill in the art could have combined the elements as claimed by known methods, and in that combination, each element merely performs the same function as it does separately. One of ordinary skill in the art would have also recognized that the results of the combination were predictable as of the effective filing date of the claimed invention.
Regarding claim 7:
Green teaches: the method of claim 6, wherein the language model is fine-tuned using a dataset of images depicting first objects and attribute labels associated with the first objects (para. 29, “Training and/or learning may be supervised using known and true outputs (e.g., known labels of features of objects) associated with the training data (objects found in images).”).
It would have been obvious for one of ordinary skill in the art, as of the effective filing date of Applicant’s claims, to have further modified the applied reference(-s) in view of same to have obtained the above, motivated to facilitate image editing and analysis, via inclusion of labels.
Regarding claim 13:
It would have been obvious for one of ordinary skill in the art to have further modified the applied reference(-s), in view of same, to have obtained: the method of claim 1, wherein performing the image segmentation comprises identifying a class of the foreground object using a classifier, and the results of the modification would have been obvious and predictable to one of ordinary skill in the art as of the effective filing date of the claimed invention. See MPEP §2143(A).
Green teaches a classifier (para. 28) that can identify “attributes” which can be “related to what is present in the image, such as the coloring, lighting, people, buildings, scenery, objects, themes, etc.” (para. 22). Labels that are for “people”, “buildings” , “scenery” all teach class, and applying this to foreground objects, per both references (see mapping to claim 1 for Wu), is all of aught/suggested by the prior art, and the results of the modification would have been obvious and predictable to one of ordinary skill in the art.
One of ordinary skill in the art could have combined the elements as claimed by known methods, and in that combination, each element merely performs the same function as it does separately. One of ordinary skill in the art would have also recognized that the results of the combination were predictable as of the effective filing date of the claimed invention.
Claim(s) 8 and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Wu in view of Benedetto and further in view of Aggarwal (U.S. Patent App. Pub. No. 2025/0278816 A1) and Zadeh (U.S. Patent App. Pub. No. 2025/0315691 A1).
Regarding claim 8:
It would have been obvious for one of ordinary skill in the art to have combined and modified the applied reference(-s), in view of same, to have obtained: the method of claim 2, further comprising: computing input embeddings using the object attributes of the foreground object and the current text input; and
determining a ranking of outputs of the LLM based on a similarity measure between embeddings of predictive text candidates and the input embeddings, and the results of the modification would have been obvious and predictable to one of ordinary skill in the art as of the effective filing date of the claimed invention. See MPEP §2143(A).
Aggarwal teaches that it is known to generate input embeddings for input images and text (e.g. claim 1, generating image embeddings for the multiple images and a text embedding for the text input). This teaches the above computing step. Aggarwal also teaches that the image embedding represents the “content and visual features” of the input image, i.e. attributes of foreground objects. See e.g. paras. 45, 53. Moreover, an input image as per Aggarwal is an embedding computed “using the object attributes of the foreground object”, because this is part of the input image.
Re: the determining step, Zadeh teaches that it is known to compute similarity scores or measures between embeddings, such similarity scores between embeddings of a user prompt and stored prompts, the user prompt embedding corresponding to Applicant’s claimed “input embeddings”, and the stored prompts corresponding to “embeddings of predictive text candidates”. See Zadeh, Fig. 2 and paras. 47-48. The stored prompts are ranked based on similarity scores. Zadeh, Fig. 2. This ranking of Zadeh teaches the claimed “ranking of outputs of the LLM”, said outputs corresponding to predictive text/stored prompts. Modifying the applied references, in view of same, such to have included embedding, per Aggarwal, and similarity comparisons/ranking, per Zadeh, and to use both embeddings per Agarwal, as relevant to the predictive prompt (which is relevant to both prompt and input image for modification, per Wu), is all of taught, suggested and motivated by the prior art. Motivation would be to better elicit the desired outputs and formulate better prompts (Zadeh, para. 11).
The prior art included each element recited in claim 8, although not necessarily in a single embodiment, with the only difference being between the claimed element and the prior art being the lack of actual combination of certain elements in a single prior art embodiment, as described above.
One of ordinary skill in the art could have combined the elements as claimed by known methods, and in that combination, each element merely performs the same function as it does separately. One of ordinary skill in the art would have also recognized that the results of the combination were predictable as of the effective filing date of the claimed invention.
Regarding claim 17: see claim 8.
These claims are similar; the same rationale for rejection applies.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
U.S. Patent App. Pub. No. 20250078453: Systems and methods for training a multi-modal machine learning architecture for content generation.
* * * * *
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
* * * * *
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Sarah Lhymn whose telephone number is (571)270-0632. The examiner can normally be reached M-F, 9:00 AM to 6:00 PM EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Xiao Wu can be reached at 571-272-7761. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
Sarah Lhymn
Primary Examiner
Art Unit 2613
/Sarah Lhymn/Primary Examiner, Art Unit 2613