Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claim(s) 1-6 and 16-20 is/are rejected under 35 U.S.C. 102(a)(2) as being anticipated by Wilson (US 20250200843 A1).
Regarding claim 1, Wilson teaches a method comprising:
obtaining an input image, and input mask, and a masked image, wherein the input image depicts a scene, the input mask indicates an inpainting region of the input image (par. 0035: “In an image-to-image mode, an image is generated from a text prompt and an input image, and the generated image retains features of the input image while introducing new elements or styles consistent with the prompt. In an inpainting mode, the processing is similar to the image-to-image mode, but an image mask is used to determine which parts of the image are fixed to match the input image.”), and wherein the masked image depicts the input mask layered on top of the input image (par. 0057: “To create the composite images described above, a layered approach can be employed. FIG. 6 shows an example where a scene camera 602 is employed to generate a composite image from background layer 604, person layer 606, and foreground layer 608. The background layer can include the actual backgrounds from the received video feeds as well as the background environment generated by the generative image model.);
generating, using a generator network of an image generation model, a latent code by denoising a noise map based on the input image and the input mask, wherein the latent code includes synthesized content in the inpainting region (par. 0031: “In the latent space 110, a diffusion process 116 adds noise to obtain a noisy representation 118 (Z.sub.T). A denoising component 120 (E.sub.e) is trained to predict the noise in the compressed latent image Z.sub.T. The denoising component can include a series of denoising autoencoders implemented using UNet 2D convolutional layers.”); and
generating, using a decoder network of the image generation model, a synthetic image based on the latent code, the input mask, and the masked image (par. 0030: “An image 102 (X) in pixel space 104 (e.g., red, green, blue) is encoded by an encoder 106 (E) into a representation 108 (Z) in a latent space 110. A decoder 112 (D) is trained to decode the latent representation Z to produce a reconstructed image 114 (X˜) in the pixel space.”), wherein the decoder network takes the latent code, the input mask, and the masked image as input (), and wherein the synthetic image depicts the scene from the input image outside the inpainting region and includes the synthesized content within the inpainting region, and wherein the synthetic image comprises a seamless transition across a boundary of the inpainting region (par. 003: “Generative image model 100 can be employed for text to image generation, where an image is generated from a text prompt. In other cases, generative image model 100 can be employed for image-to-image mode, where an image is generated using an input image as well as a text prompt. Generative image model 100 can also be employed for inpainting, where parts of an image are masked and remain fixed while the rest of the image is generated by the model, in some cases conditioned on a text prompt.”).
Regarding claim 2, Wilson teaches the method of claim 1, further comprising:
selecting an inpainting mode, wherein the synthetic image is generated based on the inpainting mode (par. 0036: “The disclosed implementations employ the inpainting mode to create a new environment by merging users' video feed backgrounds into a unified environment, and employ the image-to-image mode to transform an existing image of an environment (i.e., an image prior) based on one or more prompts (e.g., relating to purpose of the teleconference) and composites users' videos within the transformed images.”).
Regarding claim 3, Wilson teaches the method of claim 1, further comprising:
obtaining an input prompt, wherein the synthesized content is based on the input prompt (par. 0035: “In a text-to-image mode, an image is generated from a given text prompt.”).
Regarding claim 4, Wilson teaches the method of claim 1, wherein generating the latent code comprises:
obtaining the noise map (par. 0032: “The denoising can involve conditioning 122 on other modalities, such as a semantic map 124, text 126, images 128, or other representations 130 which can be processed to obtain an encoded representation 132 (T.sub.e). For instance, text can be encoded using a text encoder (e.g., BERT, CLIP, etc.) to obtain the encoded representation. This encoded representation can be mapped to layers of the denoising component using cross-attention. The result is a text-conditioned latent diffusion model that can be employed to generate images conditioned on text inputs.”);
encoding the input image to obtain an input encoding (par. 0032, as above); and
denoising the noise map based on the input encoding (par. 0032, as above).
Regarding claim 5, Wilson teaches the method of claim 1, wherein:
the image generation model is trained for an inpainting task using a training set including a training latent code representing a seam artifact (par. 0030: “FIG. 1 illustrates an example generative image model 100. An image 102 (X) in pixel space 104 (e.g., red, green, blue) is encoded by an encoder 106 (E) into a representation 108 (Z) in a latent space 110. A decoder 112 (D) is trained to decode the latent representation Z to produce a reconstructed image 114 (X˜) in the pixel space.”).
Regarding claim 6, Wilson teaches the method of claim 1, further comprising:
generating the masked image based on the input image and the input mask (par. 0035: “In an inpainting mode, the processing is similar to the image-to-image mode, but an image mask is used to determine which parts of the image are fixed to match the input image. The rest of the image is generated in a way that it is consistent with the fixed parts of the image.”).
Claim 16 is substantially similar to claim 1, except that it teaches a system rather than a method. It is therefore rejected on similar grounds to claim 1.
Regarding claim 17, Wilson teaches the system of claim 16, wherein: the generator network comprises a latent diffusion model (par. 0032: “For instance, text can be encoded using a text encoder (e.g., BERT, CLIP, etc.) to obtain the encoded representation. This encoded representation can be mapped to layers of the denoising component using cross-attention. The result is a text-conditioned latent diffusion model that can be employed to generate images conditioned on text inputs.”).
Regarding claim 18, Wilson teaches the system of claim 16, wherein: the decoder network comprises a generative adversarial network (GAN) (par. 0026: “For instance, an image model can be implemented as a neural network, e.g., a generative image model such as Stable Diffusion or DALLE.”).
Claim 19 is substantially similar to claim 6, and differs only in that it depends from claim 16 rather than claim 1. It is therefore rejected on similar grounds to claim 6.
Regarding claim 20, Wilson teaches the system of claim 16, wherein:
the image generation model is trained to generate the synthetic image with the seamless transition based on a training latent code having a seam artifact (par. 0034: “In some cases, generative image model 100 can be implemented as a Stable Diffusion model (Rombach, et al., “High-Resolution Image Synthesis with Latent Diffusion Models,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022), which can be guided by a separate network, such as a ControlNet (Zhang, et al., “Adding Conditional Control to Text-to-Image Diffusion Models,” Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023). For instance, a ControlNet can guide the generative model to produce an image that preserves certain aspects of another image, e.g., the spatial layout and salient features of an image prior. A ControlNet can be implemented by locking the parameters of generative image model 100, cloning the model into another copy. The copy is connected to the original model with one or more zero convolutional layers which are then optimized with the parameters of the copy. For instance, the ControlNet can be trained to preserve edges, lines, boundaries, human poses, from an image, semantic segmentations, object depth, etc. The outputs of a ControlNet can be added to connections within the denoising layer. Thus, the generative image model can produce images that are conditioned not only on text, but also aspects of another image.”).
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 7, 9-10, and 12 is/are rejected under 35 U.S.C. 103 as being unpatentable over Koujan (US 20240127563 A1), and further in view of Casallas (US 11915362 B2).
Regarding claim 7, Koujan teaches a method of training an image generation model, the method comprising:
obtaining a training set including a training image (par. 0061: “Specifically, the machine learning technique is applied to a first set of the training data that includes a first training image of the plurality of training images depicting synthetically rendered whole bodies of persons to generate an estimated a stylized version of the whole body of the person depicted in the first training image.”);
encoding the training image to obtain a preliminary latent code (par. 0124: “In some examples, the body stylizing system 224 generates the training data by performing training operations. The body stylizing system 224 accesses a first set of latent code by first and second whole-body GANs.”); and
training, using the training set and the training latent code with the seam artifact, an image generation model to generate a synthetic image without the seam artifact (par. 0145: “The F1_prime and F2_prime are fused back smoothly onto the pair of original and stylized full- or whole-body images (I1 and 12). To do so, an optimization algorithm that searches for the best matching face images to faces in the full-body images is used. The optimization can start in a first iteration from the pair of latent vectors and searches for new latent vectors that given a pair of faces that when fused with I1 and I2 result in no seams or artefacts being visible.”).
Koujan fails to teach generating a training latent code representing the training image by adding a seam artifact to the preliminary latent code.
Casallas teaches generating a training latent code representing the training image by adding a seam artifact to the preliminary latent code (col. 7, lines 39-42: “As an example, model 210(1) may be trained to receive a 2D image depicting a view of a 3D model and generate output indicating predicted seams for the view of the 3D model depicted in the 2D image.”).
It would have been obvious to one familiar in the art prior to the effective filing date of the claimed invention to introduce the seam artifact addition of Cassallas to the stylizing method of Koujan, as both are in the same field of endeavor of training an artificial intelligence network for image enhancement. Adding a seam artifact to the preliminary latent code would prove obviously beneficial to one familiar in the art who wished to remove such artifacts from a final output image, as it would help train a GAN to identify such artifacts.
Regarding claim 9, Koujan and Casallas teach the method of claim 7. Koujan futher teaches wherein generating the training latent code comprises:
adding the seam artifact to an image to obtain an augmented image (par. 0145, as above in claim 7 rejection); and
encoding the augmented image to obtain the training latent code (par. 0144: “A synthetic dataset of paired full-body images is generated using the first and second whole body GANs 620 and 622 (as discussed above) for a given latent code/vector. Given a pair of original and stylized full or whole-body images (I1 and I2) generated by the first and second whole body GANs 620 and 622, the faces are cropped from this pair of images to provide cropped faces F1 and F2. These cropped faces are projected onto the latent space of the trained face-only GAN that is configured to generate stylized version of a synthesized face, such as using a StyleGAN encoder.”).
Regarding claim 10, Koujan and Casallas teach the method of claim 7. Koujan futher teaches wherein:
the seam artifact comprises a random noise distortion, color augmentation, erosion, dilation, blurring, or any combination thereof (par. 0078: “Depending on the specific request for modification, properties of the mentioned areas can be transformed in different ways. Such modifications may involve changing color of areas; removing at least some part of areas from the frames of the video stream; including one or more new objects into areas which are based on a request for modification; and modifying or distorting the elements of an area or object.”).
Regarding claim 12, Koujan and Casallas teach the method of claim 7. Koujan futher teaches wherein training the image generation model comprises:
computing a generative adversarial network (GAN) loss (par. 0124: “In some examples, the body stylizing system 224 generates the training data by performing training operations. The body stylizing system 224 accesses a first set of latent code by first and second whole-body GANs. The body stylizing system 224 renders, by the first whole body GAN, a first synthetic whole body of a person corresponding to the first set of latent code and renders, by the second whole-body GAN, a second synthetic whole body of the person corresponding to the first set of latent code. The body stylizing system 224 computes directional loss, by a directional loss model associated with the given style, based on the second synthetic whole body of the person. The body stylizing system 224 updates one or more weights of the second GAN based on the directional loss and repeats the operations for rendering of the second synthetic whole body of the person, the computing of the directional loss and the updating of the one or more weights until a stopping criterion is reached..”); and
updating parameters of the image generation model based on the GAN loss (as above).
Claim(s) 11 and 13-15 is/are rejected under 35 U.S.C. 103 as being unpatentable over Koujan (US 20240127563 A1) and Casallas (US 11915362 B2) as applied to claim 7 above, and further in view of Liba (US 20250037251 A1).
Regarding claim 11, Koujan and Casallas teach the method of claim 7, but fail to teach obtaining a mask, wherein the seam artifact is added at a boundary region of the mask.
Liba teaches obtaining a mask, wherein the seam artifact is added at a boundary region of the mask (par. 0063: “In a further example, similarity calculator 440 may use the grayscale version of inpainting mask 306 to weight the similarity between input image 302 and guide image 304. That is, a similarity between a portion of guide image 304 and pixels of input image 302 that are near the region to be inpainted may contribute to similarity metric 442 more than a similarity between the portion of guide image 304 and pixels of input image 302 that are further away from the region to be inpainted, since the inpainted image content should blend well with image content of input image 302 at the boundaries of the region to be inpainted, but may differ from image content of input image 302 that is not near the boundaries of the region to be inpainted.”).
It would have been obvious to one familiar in the art prior to the effective filing date of the claimed invention to combine the masking technique of Liba with the training method of Koujan and Casallas, as all are in the same field of endeavor of training artificial networks for image analysis/enhancement. Koujan and Casallas teach a segmentation mask like that explored in Liba, but fail to detail the process substantially; one familiar in the art would naturally seek out a similar invention to implement the segmentation mask taught by Kouhan, and Liba would be an obvious choice, given the similar field of endeavor.
Regarding claim 13, Koujan and Casallas teach the method of claim 7, but fail to teach wherein training the image generation model comprises:
computing a reconstruction loss; and
updating parameters of the image generation model based on the reconstruction loss.
Liba teaches wherein training the image generation model comprises:
computing a reconstruction loss (par. 0075: “Specifically, training system 500 may include inpainting system 300, perceptual loss model 510, perceptual loss function 516, discriminator model 520, adversarial loss function 524, and model parameter adjuster 528.”); and
updating parameters of the image generation model based on the reconstruction loss (par. 0075: “Training system 500 may be configured to determine updated model parameters 530 based on training input image 502, training guide image 504, and training inpainting mask 506.”).
It would have been obvious to one familiar in the art prior to the effective filing date of the claimed invention to combine the loss calculation of Liba with the training method of Koujan, as both are in the same field of endeavor of training artificial networks for image analysis/enhancement. Koujan teaches a similar technique for computing GAN loss (as above in claim 12 rejection), making it obvious to one familiar in the art to implement a similar technique for reconstruction loss.
Regarding claim 14, Koujan and Casallas teach the method of claim 7, but fail to teach wherein training the image generation model comprises:
computing a perceptual loss; and
updating parameters of the image generation model based on the perceptual loss.
Liba teaches wherein training the image generation model comprises:
computing a perceptual loss (par. 0075, as above in claim 13 rejection); and
updating parameters of the image generation model based on the perceptual loss (par. 0075, as above in claim 13 rejection).
It would have been obvious to one familiar in the art prior to the effective filing date of the claimed invention to combine the loss calculation of Liba with the training method of Koujan and Casallas, as all are in the same field of endeavor of training artificial networks for image analysis/enhancement. Koujan teaches a similar technique for computing GAN loss (as above in claim 12 rejection), making it obvious to one familiar in the art to implement a similar technique for perceptual loss.
Regarding claim 15, Koujan and Casallas teach the method of claim 7, but fail to teach wherein training the image generation model comprises:
freezing parameters of a generator network of the image generation model while training a decoder network of the image generation model.
Liba teaches wherein training the image generation model comprises:
freezing parameters of a generator network of the image generation model while training a decoder network of the image generation model (par. 0061: “One or more parameters of modulator 432, modulator 436, similarity calculator 440, and/or convolution 452 may be learnable during training of StyleGAN model 330B and/or inpainting system 300.”).
It would have been obvious to one familiar in the art prior to the effective filing date of the claimed invention to combine the parameter freezing of Liba with the training method of Koujan and Casallas, as all are in the same field of endeavor of training artificial networks for image analysis/enhancement. Koujan teaches decoder networks, but fails to explain the functionality of such; it is the opinion of the examiner that this failure is only due to such networks lying outside of the primary focus of the invention of Koujan. It is well-known in the art to freeze the parameters of a generator network while training a decoder network of the image generation model, as shown by Liba’s demonstration of the same.
Response to Arguments
Applicant’s arguments with respect to claim(s) 1-7 and 9-20 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to RYAN A BARHAM whose telephone number is (571)272-4338. The examiner can normally be reached Mon-Fri, 8:30am-5pm EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Xiao Wu can be reached at (571) 272-7761. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/RYAN ALLEN BARHAM/ Examiner, Art Unit 2613
/XIAO M WU/ Supervisory Patent Examiner, Art Unit 2613