Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Objections
Claim 19 is objected to because of the following informalities: the claim cannot depend from itself. Appropriate correction is required.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-3, 5-13 and 15-20 are rejected under 35 U.S.C. 103 as being unpatentable over Avrahami et al. in view of Zhang et al. and further in view of Nichol et al.
Regarding claim 1, Avrahami et al. discloses a computer system for image manipulation, the computer system comprising:
one or more processors (Examiner articulates that machine learning systems require specialized processors such as CPUs and GPUs to execute intensive logic and computations to run artificial intelligence models.);
a machine-learned image manipulation model configured to receive and process an input image and a natural language instruction to generate an edited image in accordance with the natural language instruction (performing local (region-based) edits in generic natural images, based on a natural language description along with a Region of Interest (ROI) mask, page 1, column 1, lines 3-5),
wherein the desired manipulation has been performed in the region of the input image to generate the edited image (given an input image and a mask, modifying the masked area according to a guiding text prompt, without affecting the unmasked regions, figure 1).
Avrahami et al. does not expressly disclose wherein the natural language instruction comprises a reference portion that refers to a region of the input image and a target portion that describes a desired manipulation to be performed in the region of the input image.
Zhang et al. teaches fill the semantic information in corrupted images according to the provided descriptive text, page 1, column 1, lines 6-7; the model fills the holes with the guidance of descriptive text, page 2, column 2, lines 51-52.
Avrahami et al. in view of Zhang et al. are analogous art because they are from the similar problem-solving area of image content manipulation. At the time of the invention, it would have been obvious to a person of ordinary skill in the art to add the task of Zhang et al. to the technique of Avrahami et al. in order to obtain generated images consistent with the guidance text. The motivation for doing so would be to improve the filling of corrupted images.
one or more non-transitory computer-readable media that collectively store instructions that, when executed by the one or more processors, cause the one or more processors to perform operations, the operations comprising:
Avrahami et al. in view of Zhang et al. does not expressly disclose obtaining the input image and the natural language instruction (GLIDE); processing the input image and the natural language instruction with the machine-learned image manipulation model to generate the edited image (GLIDE); and providing the edited image as an output (GLIDE).
Nichol et al. teaches given the masked input image I𝑚 and descriptive text T, the model outputs the target image I𝑔. The model uses the dual probabilistic structure and extend it to the multimodal condition. The overall structure of the model is shown in Figure 2. It’s composed of three components: Encoders for Image and Text, Dual multimodal Attention, and Inpainting Generation. The generated images are consistent with the guidance text, enabling the generation of various results by providing different descriptions.
Avrahami et al. in view of Zhang et al. and further in view of Nichol et al. are analogous art because they are from the similar problem-solving area of image content manipulation. At the time of the invention, it would have been obvious to a person of ordinary skill in the art to add the task of Nichol et al. to the technique of Avrahami et al./ Zhang et al. in order to obtain generated images consistent with the guidance text. The motivation for doing so would be to perform image inpainting, enabling powerful text-driven image editing.
Regarding claim 2, Nichol et al. discloses the computer system of claim 1, wherein:
the machine-learned image manipulation model comprises a machine-learned image segmentation model and a machine-learned inpainting model (see figure 2);
the machine-learned image segmentation model is configured to receive and process the input image and the reference portion of the natural language instruction to generate an image mask that identifies the region of the input image (see figure 2, "Text-conditional image inpainting examples. The green region is erased, and the model fills it in conditioned on the given prompt.", The green region represents the mask.; and
the machine-learned inpainting model is configured to receive and process the input image, the image mask, and the target portion of the natural language instruction to generate the edited image (see figure 2).
Regarding claim 3, Nichol et al. discloses the computer system of claim 2, wherein the machine-learned inpainting model comprises a conditional diffusion model (diffusion models for the problem of text- conditional image synthesis and compare two different guidance strategies: CLIP guidance and classifier-free guidance; train a 3.5 billion parameter text-conditional diffusion model at 64 × 64 resolution, and another 1.5 billion parameter text-conditional upsampling diffusion model to increase the resolution to 256 × 256, see section 4. Training)).
Regarding claim 5, Nichol et al. discloses the computer system of claim 2, wherein the machine-learned inpainting model is provided with conditional classifier-free guidance during generation of the edited image (Text-conditional image inpainting examples from GLIDE. The green region is erased, and the model fills it in conditioned on the given prompt. Our model is able to match the style and lighting of the surrounding context to produce a realistic completion, figure 2; and see section 5.1).
Regarding claim 6, Nichol et al. discloses the computer system of claim 5, wherein the conditional classifier-free guidance guides the generation of the edited image toward the target portion of the natural language instruction from the reference portion of the natural language instruction (Text-conditional image inpainting examples from GLIDE. The green region is erased, and the model fills it in conditioned on the given prompt. Our model is able to match the style and lighting of the surrounding context to produce a realistic completion, figure 2; and see section 5.1).
Regarding claim 7, Nichol et al. discloses the computer system of claim 5, wherein: the machine-learned inpainting model comprises a diffusion model (Gaussian diffusion model, section 2.1); and
the conditional classifier-free guidance perturbs additive noise of the diffusion model based on a probability associated with the reference portion of the natural language instruction (produce a Markov chain of latent variables x1, xT by progressively adding Gaussian noise to the sample, section 2.1).
Regarding claim 8, Nichol et al. discloses the computer system of claim 2, wherein one or both of the machine-learned image segmentation model and the machine-learned inpainting model were pretrained prior to inclusion within the machine-learned image manipulation model and did not undergo additional training or fine-tuning after inclusion in the machine-learned image manipulation model .
Regarding claim 9, Nichol et al. discloses the computer system claim 1, wherein processing the input image and the natural language instruction with the machine- learned image manipulation model to generate the edited image comprises:
processing, for a plurality of instances, the input image and the natural language instruction with the machine-learned image manipulation model to generate a plurality of candidate images; generating a respective semantic similarity score for at least a portion of each of the plurality of candidate images relative to at least the target portion of the natural language instruction; and selecting one of the candidate images to output as the edited image based at least in part the respective semantic similarity scores (ranking, see e.g. Fig. 5, "we generate samples at temperature 0.85 and select the best of 256 using CLIP reranking").
Regarding claim 10, Nichol et al. discloses the computer system of claim 9, wherein generating the respective semantic similarity score for at least the portion of each of the plurality of candidate images comprises generating the respective semantic similarity score for a respective context-aware bounding box generated for each of the plurality of candidate images, wherein the respective context-aware bounding box for each of the plurality of candidate images is defined by enlarging by an enlargement factor a bounding box that includes the region of the input image (ranking, see e.g. Fig. 5, "we generate samples at temperature 0.85 and select the best of 256 using CLIP reranking").
Claim 11, a computer-implemented method, is rejected for the same reason as claim 1.
Claim 12, a computer-implemented method, is rejected for the same reason as claim 2.
Claim 13, a computer-implemented method, is rejected for the same reason as claim 3.
Claim 15, a computer-implemented method, is rejected for the same reason as claim 5.
Claim 16, a computer-implemented method, is rejected for the same reason as claim 6.
Claim 17, a computer-implemented method, is rejected for the same reason as claim 7.
Claim 18, a computer-implemented method, is rejected for the same reason as claim 9.
Claim 19, a computer-implemented method, is rejected for the same reason as claim 10.
Claim 20, a non-transitory computer-readable media, is rejected for the same reason as claim 1.
Allowable Subject Matter
Claims 4 and 14 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to THOMAS J LETT whose telephone number is (571)272-7464. The examiner can normally be reached Mon-Fri 9-6 ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Tammy Goddard can be reached at (571) 272-7773. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/THOMAS J LETT/ Primary Examiner, Art Unit 2611