Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Arguments
Applicant’s arguments with respect to all claims have been considered but are moot in view of the new grounds of rejection (see below).
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1 – 12 and 14 – 20 are rejected under 35 U.S.C. 103 as being unpatentable over Feldman (US Pub. No. 2026/0094404 A1) in view of Abel et al. (US Pub. No. 2024/0201833 A1).
As to claims 1 and 14, Feldman shows a method (Figs. 6 and 7 and paras. 82 and 132) and associated apparatus (computing device 200, Fig. 2 and para. 36), comprising: obtaining an input image (625, for example, Fig. 6 and paras. 82 and 85) and a modification prompt (Fig. 6 and para. 85), wherein the input image depicts a first element (i.e. boot 627, for example, Fig. 6 and para. 85) and the modification prompt 630 describes a second element different from the first element (Fig. 6 and para. 85, note that the user inputs text so as to alter the selected item); generating, using an inversion model (i.e. diffusion model 208), an intermediate output based on the input image (paras. 89 and 122), wherein the intermediate output comprises image features representing the image (i.e. items other than the selected item, Fig. 6 and paras. 85, 86, 89 and 122); and generating, using an image generation model (i.e. inpainter module 206/diffusion module 208, for example, paras. 108 – 112), a synthetic image 651 based on the intermediate output and the modification prompt (Fig. 6 and paras. 85, 86 and 89), wherein the synthetic image replaces the first element from the input image with the second element from the modification prompt (Fig. 6 and paras. 85, 86 and 89).
Feldman does not show a text description describing an input content, generating predicted noise based on the input image and the text description, wherein the intermediate output comprises image features representing the input image corresponding to an intermediate timestep of a denoising process; or generating, using an image generation model, a synthetic image by denoising based on the intermediate output based on the modification prompt starting from the intermediate timestep.
Abel shows a text description describing an input content (i.e. 318/514, Fig. 5A and paras. 52 and 53), generating predicted noise based on the input image and the text description (Fig. 5B and paras. 65 and 66), wherein the intermediate output comprises image features representing the input image corresponding to an intermediate timestep of a denoising process (paras. 73 – 76); or generating, using an image generation model, a synthetic image by denoising based on the intermediate output based on the modification prompt starting from the intermediate timestep (paras. 73 – 76).
It would have been obvious to one of ordinary skill in the art at the time of filing to modify the teachings of Feldman with those of Abel because designing the system in this way allows the device to create custom images based on user preferences (para. 1).
As to claims 2 and 15, Feldman shows obtaining the text description of the input image, wherein the intermediate output is generated based on the text description (i.e. “change the boots…”, Fig. 6 and para. 85).
As to clams 3 and 16, Feldman shows that the modification prompt comprises an edit to a text description of the input image (i.e. “change the boots…”, Fig. 6 and para. 85).
As to claims 4 and 17, Feldman shows that generating the synthetic image comprises: iteratively alternating between generating successive intermediate outputs using the inversion model and the image generation model (Figs. 6 and 7 and para. 140, note that the process can be repeated/refreshed).
As to claims 5 ad 18, Feldman shows that generating a reconstructed image based on the intermediate output (Fig. 6 and para. 91), wherein the reconstructed image depicts the first element (i.e. the boot, Fig. 6 and paras. 86 and 91); and generating a subsequent intermediate output based on the reconstructed image, wherein the synthetic image is based on the subsequent intermediate output (Figs. 6 and 7 and para. 140, note that the process can be repeated/refreshed).
As to claims 6 and 19, Feldman shows that generating the synthetic image comprises: obtaining a noise input (para. 89); and denoising the noise input based on the intermediate output (para. 89).
As to claim 7, Feldman shows that the inversion model is trained using a training set including a training image and a training description of the training image (para. 93).
As to claim 20, Feldman shows that the inversion model comprises a diffusion model (paras. 89 and 122).
As to claim 8, Feldman shows a method for training a machine learning model (Figs. 3, 6 and 7 and paras. 64, 82 and 132) the method comprising: obtaining a training set including an input image 625 and a text description of the training image (Fig. 6 and paras. 82, 85 and 93); generating an intermediate output based on the input image and the text description (Fig. 6 and paras. 89, 93 and 122); generating, using an image generation model (i.e. inpainter module 206/diffusion module 208, for example, paras. 108 – 112), a reconstructed image based on the intermediate output and the text description (i.e. including a new version of the boot, Fig. 6 and paras. 86 and 91); and training, using the training set and the reconstructed image, the inversion model (i.e. diffusion model 208) to perform image inversion (paras. 89 and 122).
As to claim 9, Feldman shows that obtaining the training set comprises: generating the text description based on the input image (Fig. 6 and paras. 85 and 93).
As to claim10, Feldman show that training the inversion model comprises: computing a reconstruction loss based on the input image and the reconstructed image (paras. 93 and 97); and updating parameters of the inversion model based on the reconstruction loss (paras. 93 and 97).
As to claim 11, Feldman shows that training the inversion model comprises: generating a modified image based on the intermediate output and a modification prompt 630 (Fig. 6 and para. 85, note that the user inputs text so as to alter the selected item); computing a modification loss based on the modified image and a ground-truth modified image (paras. 69, 93 and 97); and updating parameters of the inversion model based on the modification loss (paras. 69, 93 and 97).
As to claim 12, Feldman shows initializing the inversion model using parameters from the image generation model (para. 68).
Claim 13 is rejected under 35 U.S.C. 103 as being unpatentable over Feldman and Abel as modified above in view of Atzmon et al. (US Pub. No. 2024/0249446 A1).
As to claim 13, Feldman does not show that an image generation model is frozen during the training of the inversion model.
Atzmon shows that an image generation model is frozen during the training of a model (para. 30).
It would have been obvious to one of ordinary skill in the art at the time of filing to modify the teachings of Feldman with those of Atzmon because designing the system in this way allows the device to reduce memory footprint (para. 32).
CONCLUSION
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to CARL ADAMS whose telephone number is (571)270-7448. The examiner can normally be reached Monday - Friday, 9AM - 5PM EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Ke Xiao can be reached at 571-272-7776. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/CARL ADAMS/Examiner, Art Unit 2627