DETAILED ACTION
This action is in response to the application filed 12/3/2024.
Claims 1-10 have been submitted for examination.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1, 2, and 5- 10 is/are rejected under 35 U.S.C. 103 as being unpatentable over Zhang et al. (US 2023/0230198), hereinafter Zhang, in view of Boyd (US 2024/0378251).
As per claim 1, Zhang teaches the following:
a method that is executed by an information processing apparatus, (see Fig. 1), the method comprising:
storing, in a memory, (see Fig. 10), a first prompt including an instruction for processing of a first image. As Zhang teaches in paragraph [0047], and corresponding Fig. 2, at step 202 an image generation system 106 receives a natural language command indicating a targeted image element,
the first prompt having been received as an input that is provided to a first trained model which performs processing on and outputs an image. As Zhang teaches in paragraph [0049], and corresponding Fig. 2, at step 204 the image generation system uses the textual features from the natural language command to condition a generative neural network to generate an image. Zhang teaches in paragraph [0021] that the neural network is trained and in the abstract that a train model is utilized; and
generating a second prompt including at least an instruction for processing of a second image which has been specified, . As Zhang teaches in paragraphs [0050] and [0051], and corresponding Fig. 2, at steps 206 and 208, additional natural language commands may be received to modify the generated image, thus forming a “second image”.
However, Zhang does not explicitly teach of reusing at least part of the first prompt. In a similar field of endeavor, Boyd teaches of a method of automatically generating media assets using a machine-learned generator (see abstract). Boyd further teaches in paragraph [0081], that the user may enter a command to generate additional assets similar to a prior result, which causes the system to re-use the prompt used to generate the previously-generated asset.
It would have been obvious to one of ordinary skill in the art before the effective filing date of applicant’s claimed invention to have modified the input commands of Zhang with the reuse of Boyd. One of ordinary skill would have been motivated to have made such modification because as Boyd teaches in paragraph [0082], such iterative generation benefits users in creating more desirable assets based on user feedback.
Regarding claim 2, modified Zhang teaches the method of claim 1 as described above. However, as described above, Zhang does not explicitly teach of reusing at least part of the first prompt. Boyd further teaches the following:
receiving an instruction that reuses the first prompt, wherein, when the instruction for reusing the first prompt is received, the second prompt is generated by reusing at least a part of the first prompt. As Boyd teaches in paragraph [0081], “the prompt used to generate the previously-generated asset can be re-used”.
It would have been obvious to one of ordinary skill in the art before the effective filing date of applicant’s claimed invention to have modified the input commands of Zhang with the reuse of Boyd. One of ordinary skill would have been motivated to have made such modification because as Boyd teaches in paragraph [0082], such iterative generation benefits users in creating more desirable assets based on user feedback.
Regarding claim 5, modified Zhang teaches the method of claim 1 as described above. However, as described above, Zhang does not explicitly teach of reusing at least part of the first prompt. Boyd further teaches the following:
the second prompt is generated by reusing at least one of metadata of the first image and metadata of the second image, and the first prompt. As Boyd teaches in paragraph [0073], generated assets can be associated with metadata. Boyd further teaches in paragraph [0081] that to generate more assets like a presented asset, the model can input the existing asset as part of the prompt. Therefore, the metadata of the existing asset is utilized in the second prompt.
It would have been obvious to one of ordinary skill in the art before the effective filing date of applicant’s claimed invention to have modified the input commands of Zhang with the reuse of Boyd. One of ordinary skill would have been motivated to have made such modification because as Boyd teaches in paragraph [0082], such iterative generation benefits users in creating more desirable assets based on user feedback.
Regarding claim 6, modified Zhang teaches the method of claim 1 as described above. However, as described above, Zhang does not explicitly teach of reusing at least part of the first prompt. Boyd further teaches the following:
the second prompt is generated by inputting the first prompt and an instruction for processing of the first prompt to a second trained model that generates a prompt based on an input instruction. As Boyd teaches in paragraph[0081], one revision option includes inputting, to the model, the existing asset plus a prompt. Therefore, this is interpreted as the second prompt is generated via inputting the first prompt plus a selected asset to form a new prompt, wherein the second prompt is generated via the model.
It would have been obvious to one of ordinary skill in the art before the effective filing date of applicant’s claimed invention to have modified the input commands of Zhang with the reuse of Boyd. One of ordinary skill would have been motivated to have made such modification because as Boyd teaches in paragraph [0082], such iterative generation benefits users in creating more desirable assets based on user feedback.
Regarding claim 7, modified Zhang teaches the method of claim 1 as described above. However, as described above, Zhang does not explicitly teach of presenting the first prompt in association with the first image when the specification of the first image is received. Boyd further teaches the following:
presenting, the first prompt to a user in association with the first image when a specification of the first image is received. As Boyd shows in Fig. 13A, user input (first prompt) is presented along with generated images (first image). Boyd further shows in Fig. 14 that the two may further be displayed when an image is selected for generating more like the image.
It would have been obvious to one of ordinary skill in the art before the effective filing date of applicant’s claimed invention to have modified the input commands of Zhang with the presentation of input and selection of Boyd. One of ordinary skill would have been motivated to have made such modification because such visual feedback benefits users in ensuring that desired input has been selected.
Regarding claim 8, modified Zhang teaches the method of claim 1 as described above. While Zhang shows of displaying images both before and after processing (see Fig. 2, 204 and 208), Zhang does not explicitly teach of storing the prompt, first image, and image obtained in an associated manner. Boyd further teaches in paragraph [0073] that image assets created or enhanced using the machine-learned media asset generation pipeline can have metadata stored containing information about which tools/pipelines (and which versions) were used to create or enhance the asset. Under the modified system of Zhang, the metadata stored by Boyd may include the prompt and original image of Zhang, thus arriving at:
comprising storing the first prompt, the first image, and the image obtained by the processing in an associated manner in the memory, and presenting the first image and the image obtained by the processing to a user as images before and after the processing based on the first prompt.
It would have been obvious to one of ordinary skill in the art before the effective filing date of applicant’s claimed invention to have modified the asset generation of Zhang with the metadata storage of Boyd. One of ordinary skill would have been motivated to have made such modification because as Boyd teaches in paragraph [0073], storing such information may be beneficial in analyzing the success of how enhancements performed.
As per claim 9, Zhang teaches the following:
an information processing apparatus comprising: at least one memory that stores instructions; and at least one processor. See Fig. 1.
The remaining limitations are substantially similar to those of claim 1 and are rejected using the same reasoning.
As per claim 10, Zhang teaches the following:
a non-transitory computer-readable storage medium that stores instructions for providing an apparatus, wherein the instructions causes at least one processor. See Fig. 1.
The remaining limitations are substantially similar to those of claim 1 and are rejected using the same reasoning.
Claim(s) 3 and 4 is/are rejected under 35 U.S.C. 103 as being unpatentable over Zhang in view of Boyd as applied to claims 1 and 2, and further in view of Wong et al. (US 2021/0375320), hereinafter Wong.
Regarding claim 3, modified Zhang teaches the method of claim 2 as described above. However, as described above, Zhang in view of Boyd does nott explicitly teach of generating the second prompt after specification of the first prompt. In a similar field of endeavor, Wong teaches of editing media content (see abstract). Wong further teaches the following:
the second prompt is generated by reusing at least a part of the first prompt after receipt of a specification of the first image, and when the instruction for reusing the first prompt and a specification of a second image are received. As Wong teaches in paragraph [0033], a user may select to preview an effect for selection, whereby a video with the desired effect is displayed. Thus the user may view preview content (first image) and reuse the effect (prompt) on a target image (specification of a second image).
It would have been obvious to one of ordinary skill in the art before the effective filing date of applicant’s claimed invention to have modified the prompt selection of Zhang in view of Boyd with effect preview image of Wong. One of ordinary skill would have been motivated to have made such modification because as Wong teaches in paragraph [0033], such previews benefit users when the user has seen an effect previously but may not remember the title of the effect.
Regarding claim 4, modified Zhang teaches the method of claim 2 as described above. However, as described above, Zhang in view of Boyd does nott explicitly teach of generating the second prompt after specification of the first prompt. In a similar field of endeavor, Wong teaches of editing media content (see abstract). Wong further teaches the following:
the second prompt is generated by reusing at least a part of the first prompt after receipt of a specification of the second image, and when a specification of the first image and the instruction for reusing the first prompt are received. As Wong teaches in paragraph [0033], a user may select to preview an effect for selection, whereby a video with the desired effect is displayed. Thus the user may first select a target video to that is to have an effect applied (second image) and reuse the effect (prompt) of a selected image (first image).
It would have been obvious to one of ordinary skill in the art before the effective filing date of applicant’s claimed invention to have modified the prompt selection of Zhang in view of Boyd with effect preview image of Wong. One of ordinary skill would have been motivated to have made such modification because as Wong teaches in paragraph [0033], such previews benefit users when the user has seen an effect previously but may not remember the title of the effect.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
-Park et al. (US 12,380,569), see column 18, lines 4-29, where prompts may be re-used.
-Wu et al. (US 2016/0098851), applying an effect to an image.
-Xu et al. (US 2022/0399017), applying effects to images with natural language input.
-Soni (US 2024/0256218), using speech input to modify images.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to GREGORY A DISTEFANO whose telephone number is (571)270-1644. The examiner can normally be reached Monday - Friday: 9 am - 5 pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, William Bashore can be reached at 5712424088. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/GREGORY A. DISTEFANO/
Examiner
Art Unit 2174
/WILLIAM L BASHORE/ Supervisory Patent Examiner, Art Unit 2174