Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejection – 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1, 9, 10, and 19 and 18-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Cragg (US 20240273670 A1) hereinafter referenced as Cragg, in view of the following: Rao (US 20250225627 A1) hereinafter referenced as Rao, Yang B (Paint by Example: Exemplar-based Image Editing with Diffusion Models) hereinafter referenced as Yang B, and Hong (Learning Subject-Aware Cropping by Outpainting Professional Photos) hereinafter referenced as Hong.
Regarding claim 1, Cragg teaches:
an image processing method, comprising:
“Systems and methods for image processing are provided.” (Cragg, Abstract);
obtaining an image to be processed and an image expansion text;
“Embodiments of the present disclosure obtain an image and a target dimension for expanding the image. The system generates a prompt based on the image using a prompt generation network.” (Cragg, Abstract);
Cragg teaches of obtaining an image to be expanding (reads on obtaining an image to be processed) and a prompt based on the image using a prompt generation network (reads on an image expansion text).
wherein the target model is obtained after a preset model to be trained is iteratively trained based on a preset training dataset, the training dataset comprises a plurality of training data pairs,
“initializing a diffusion model; obtaining training data including an input image, a prompt, and a ground-truth expanded image; and training the diffusion model to generate an expanded image that includes additional content in an outpainted region that is consistent with content of the input image and the prompt based on the training data.” (Cragg, ¶ 5);
PNG
media_image1.png
598
666
media_image1.png
Greyscale
(Cragg, Figure 15)
“FIG. 15 shows an example of a method for training a machine learning model according to aspects of the present disclosure.” (Cragg, ¶ 21); “the system at each stage n, starting with stage N, a reverse diffusion process is used to predict the image or image features at stage n−1. For example, the reverse diffusion process can predict the noise that was added by the forward diffusion process, and the predicted noise can be removed from the image to obtain the predicted image. In some cases, an original image is predicted at each stage of the training process. In some cases, the operations of this step refer to, or may be performed by, a training component as described with reference to FIG. 2.” (Cragg , ¶ 142);
Cragg teaches of a initializing and training diffusion model (reads on the target model is obtained after a preset model to be trained) which to be trained with training data including an input image, a prompt, and a ground-truth (reads on a preset training dataset, the training dataset comprises a plurality of training data pairs). Cragg trains the model through a repeated loop in which noise is added to a training image and the reverse diffusion process predicts the image. The predication is compared to the actual image. The system updates the parameter according using gradient descent (reads on iteratively training).
and the training data pair comprises an original image, an image expansion description text
“initializing a diffusion model; obtaining training data including an input image, a prompt, and a ground-truth expanded image” (Cragg, ¶ 5); “In some examples, an input image depicts a person taking a picture of a cityscape in San Francisco. The prompt of the input image is “a person taking a picture of a cityscape, award winning photo, optical illusion, anamorphic widescreen, photograph of San Francisco, built on a steep hill, detailed.” Additionally, a ground-truth expanded image is an expanded version of the input image. The ground-truth expanded image includes additional content in one or more outpainted regions compared to the input image.” (Cragg, ¶ 147);
Cragg teaches of the training data pair comprises an input image (reads on original image) and a prompt (reads on image expansion description text). For example, input image depicts a person taking a picture of a cityscape in San Francisco and the prompt is “a person taking a picture of a cityscape, award winning photo, optical illusion, anamorphic widescreen, photograph of San Francisco, built on a steep hill, detailed.” (reads on image expansion description text).
obtaining a predicted noise corresponding to the padded image, which is output by the target model, and performing an image expansion operation on the image to be processed based on the predicted noise. 1fds1fdsfds1djsklfjdsklfjds
“performing a reverse diffusion process using the diffusion model to obtain a plurality of predicted noise maps, wherein the training is based on the plurality of noise maps and the plurality of predicted noise maps.” (Cragg, ¶ 136); “the reverse diffusion process can predict the noise that was added by the forward diffusion process, and the predicted noise can be removed from the image to obtain the predicted image” (Cragg, ¶ 142);
Cragg teaches of a diffusion model predicts noise (reads on obtaining a predicted noise) in which the reverse process runs on the input map containing the image plus a pre-filled outpainted region (reads on corresponding to the padded image padded image). The predicted the image is the expanded image containing additional content in an outpainted region (reads on performing an image expansion operation on the image to be processed) which is done by removing the predicted noise (reads on based on the predicted noise).
Cragg fails to teach the following: performing a padding operation on the image to be processed based on a preset background to obtain a padded image; wherein a display size of the preset background is greater than a display size of the image to be processed; inputting the padded image and the image expansion text to a preset target model; a random mask, a masked image obtained by masking the original image based on the random mask, and a cropped image obtained through cropping based on the original image and the random mask; and the training data pair comprises an original image, text, a mask, a mask obtained by masking the original image based on the random mask, and a cropped image obtained through cropping based on the original image and the random mask;
But Rao does. Rao teaches the following:
performing a padding operation on the image to be processed based on a preset background
to obtain a padded image,
“Masking and fitting of the input image 201 can involve adding “blank” content adjacent to edges of the input image 201 in order to form border regions within the masked input image.” (Rao, ¶ 68);
Rao teaches of adding “blank” content adjacent to edges of the input image (reads on performing a padding operation on the image to be processed based. The blank content reads on the padding. The input image reads on the image to be processed based).
wherein a display size of the preset background is greater than a display size of the image to be processed;
“fitting the input image to a canvas (step 314). The canvas size can be based on the desired aspect ratio received in step 301. The masking and fitting operations allow the outpainting model 206 to handle images of different sizes, in addition to outpainting border regions of an image having a different aspect ratio from the display.” (Rao, ¶ 68); “FIG. 3A has a width of 1920 pixels and a height of 1080 pixels. Also assume that the image is to be fit for outpainting on a canvas that has a width of 960 pixels and a height of 480 pixels. In some cases, the height may be fixed to 480 pixels, and the width may be determined as (1920/1080)×480≈850 pixels. After resizing, the image can be placed centrally on the canvas so that the outpainted regions defined by masking of the input image total (960−850)=110 pixels, meaning there is a 55 pixel by 480 pixel outpainting region on each side of the input image.” (Rao, ¶ 81)
Rao teaches of a canvas size can be based on the desired aspect ratio. For example, a width of 1920 pixels and a height of 1080 pixels and the image is to be fit for outpainting on a canvas that has a width of 960 pixels and a height of 480 pixels. After resizing, the image can be placed centrally on the canvas so that the outpainted regions defined by masking of the input image total (960−850)=110 pixels, meaning there is a 55 pixel by 480 pixel outpainting region on each side of the input image. (The canvas reads on a display size of the preset background and the fitted image reads on the image to be processed 960x480 canvas > fitted image of 850x480 reads on wherein a display size of the preset background is greater than a display size of the image to be processed).
inputting the padded image and the image expansion text to a preset target model,
“Masking and fitting of the input image 201 can involve adding “blank” content adjacent to edges of the input image 201 in order to form border regions within the masked input image.” (Rao, ¶ 68); “the outpaint model is run once (step 316) using the masked input image and the contextualized prompt to produce an initial outpainted image 317.” (Rao, ¶ 70); “In particular embodiments, this approach may primarily be used on the denoising U-Net 410 of the outpainting model 206,” (Rao ¶ 11);
PNG
media_image2.png
423
959
media_image2.png
Greyscale
(Rao, Figure 4);
Rao teaches of inputting image embedding 407 comprising of the masked input image with of “blank” content adjacent to edges of the input image (reads on inputting the padded image) and text embedding 412 (reads on image expansion text) into U-Net 410 of the outpainting model (reads on a preset target model).
Rao is analogous art with respect to Cragg because they are from the same field of endeavor, namely image generation via outpainting using diffusion models with image and text inputs. Before the effective filling date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to modify Cragg with the feature of Rao to incorporate “blank” content adjacent to edges of the input image and a canvas size can be based on the desired aspect ratio. A person of ordinary skill in the art would do such in order to improve the quality of the outpainted image.
Cragg in view of Rao fail to teach the following: a random mask, a masked image obtained by masking the original image based on the random mask, and a cropped image obtained through cropping based on the original image and the random mask; and the training data pair comprises an original image, text, a mask, a mask obtained by masking the original image based on the random mask, and a cropped image obtained through cropping based on the original image and the random mask;
But Hong does. Hong teaches:
the training data pair comprises an original image, text, a mask, a mask obtained by masking the original image based on the random mask, and a cropped image obtained through cropping based on the original image and the random mask; and
PNG
media_image3.png
412
1430
media_image3.png
Greyscale
(Hong, Figure 2);
Hong teaches of training data comprises an original image (A), text (B), a random mask (C), a mask obtained by masking the original image based on the mask (D), and a cropped image obtained through cropping based on the original image and the mask (E).
Hong is analogous art with respect to Cragg in view Rao because they are from the same field of endeavor, namely image diffusion models. Before the effective filling date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to modify Cragg in view of Rao with the feature of Hong to incorporate training data comprises an original image, text, a random mask, a mask obtained by masking the original image based on the mask, and a cropped image obtained through cropping based on the original image and the mask. A person of ordinary skill in the art would do such in order to improve training performance.
Claim 10 is rejected using the same rationale or bases as applied to claim 1 and the mentioned structure.
Additionally, claim 10 recites the following structure:
An electronic device, comprising a processor and a memory, wherein the memory stores computer-executable instructions; and the processor executes the computer-executable instructions stored in the memory to cause the processor to
“An apparatus and method for image processing are described. One or more embodiments of the apparatus and method include a processor; and a memory including instructions executable by the processor to” (Cragg, ¶ 6);
Claim 19 is rejected using the same rationale or bases as applied to claim 1 and the mentioned structure.
Additionally, claim 19 recites the following structure:
non-transitory computer-readable storage medium storing computer-executable instructions that, when executed by a processor, cause the processor to:
“In another embodiment, a non-transitory machine readable medium includes instructions that when executed cause at least one processor of an electronic device to perform the method of the first embodiment.” (Cragg, ¶ 4);
Regarding claim 9, Cragg in view of Rao and Hong teaches method according to claim 1 wherein performing the image expansion operation on the image to be processed based on the predicted noise, and additionally teaches the following. Rao teaches
performing a denoising operation on the padded image based on the predicted noise, and
“The masking and fitting operations allow the outpainting model 206 to handle images of different sizes, in addition to outpainting border regions of an image having a different aspect ratio from the display. Masking and fitting of the input image 201 can involve adding “blank” content adjacent to edges of the input image 201 in order to form border regions within the masked input image.” (Rao, ¶ 68); “(i) a variational autoencoder (VAE) in which the VAE encoder compresses an image from a pixel space to a smaller dimensional latent space and iteratively applies Gaussian noise to the compressed latent representation during forward diffusion and (ii) a U-Net that denoises the output from the forward diffusion backwards to obtain a latent representation (where a VAE decoder converts the latent representation back into pixel space).” (Rao, ¶ 76); “The denoise U-Net 410 can iteratively denoise the noisy latent image representation 409 based on the personalization features 411 and the text embedding 412 in order to produce a denoised latent representation 413.” (Rao, ¶ 78);
Rao teaches performing a denoising operation on the noisy latent image representation which is generated from a “blank” content adjacent to edges of the input image (reads on padded image). A variational autoencoder compresses the image from a pixel space to a smaller dimensional latent space and iteratively applies Gaussian noise (reads on predicted noise). The U-Net denoises the Gaussian noise from the forward diffusion backwards to produce a denoised latent representation (reads on performing a denoising operation).
determining the denoised padded image as a target image after image expansion.
“The initial outpainted image 317 is used to calculate an image score (step 321), such as by using one or more metrics determined as described below in connection with FIGS. 9 and 10. A determination (step 322) is made as to whether the image score is greater than a threshold (such as 0.5 in the example process 300). If so, the outpainted image 317 is output as the final outpainted image 323. If not, the outpainted image 317 is provided as an input image to prompt extraction (step 311), and the outpainting model is rerun.” (Rao, ¶ 72);
Rao teaches determining outpainted image (reads on determining the denoised padded image. The outpointed image is the padded image which has been denoised by the system) as the final outpainted image (reads on a target image after image expansion).
Before the effective filling date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to modify Cragg in view of Rao and Hong with the feature of Rao to incorporate
performing a denoising operation on the noisy latent image representation which is generated from adding “blank” content adjacent to edges of the input image based predicted noisy latent image representation; and determining outpainted image as the final outpainted image. A person of ordinary skill in the art would do such in order to improve the quality of the outpainted image.
Claim 18 is rejected using the same rationale or bases as applied to claim 9.
Claim(s) 2-4, 8, 11-13, 17, and 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Cragg in view of the following: Rao, Hong, and Yang B (Paint by Example: Exemplar-based Image Editing with Diffusion Models) hereinafter referenced as Yang B.
Regarding claim 2, Cragg in view of Rao and Hong teaches the method of claim 1. Additionally, Cragg teaches the following:
before inputting the padded image and the image expansion text to the preset target model, obtaining an original dataset, wherein the original dataset comprises original data groups, and the original data group comprises an original image, an image expansion description text;
“initializing a diffusion model; obtaining training data including an input image, a prompt, and a ground-truth expanded image; and training the diffusion model to generate an expanded image that includes additional content in an outpainted region that is consistent with content of the input image and the prompt based on the training data.” (Cragg, ¶ 5);
Cragg teaches initializing a diffusion model (reads on before inputting the padded image and the image expansion text to the preset target model) by obtaining a training data (reads on obtaining an original dataset), wherein the training data comprises an input image, a prompt (reads on an image expansion description text), and a ground-truth expanded image (reads on original image). The training data is made up of smaller data in (reads on an original dataset, wherein the original dataset comprises original data groups. The data before training so therefore it is original).
iteratively training a preset model to be trained through the training dataset to obtain the
target model.
PNG
media_image4.png
598
668
media_image4.png
Greyscale
(Cragg, Figure 15)
“FIG. 15 shows an example of a method for training a machine learning model according to aspects of the present disclosure.” (Cragg, ¶ 21); “the system at each stage n, starting with stage N, a reverse diffusion process is used to predict the image or image features at stage n−1. For example, the reverse diffusion process can predict the noise that was added by the forward diffusion process, and the predicted noise can be removed from the image to obtain the predicted image. In some cases, an original image is predicted at each stage of the training process. In some cases, the operations of this step refer to, or may be performed by, a training component as described with reference to FIG. 2.” (¶ 142);
Cragg teaches of initializing and training a diffusion model 1505 and (reads on training a preset model to be trained) through the training dataset 1515 to obtain a trained model (reads on be trained through the training dataset to obtain the target model). Cragg trains the model through a repeated loop in which noise is added to a training image and the reverse diffusion process predicts the image. The predication is compared to the actual image. The system updates the parameter according using gradient descent (reads on iteratively training).
Hong additionally teaches:
the original data group comprises an original image, an image expansion description text;
PNG
media_image5.png
473
680
media_image5.png
Greyscale
(Hong, Figure 2);
Hong teaches of a dataset (reads on original dataset) comprises an original image (A), text (B) and a mask (C);
performing data processing on a plurality of original data groups in the original dataset to
obtain a training dataset; and
“Our first goal is to construct a dataset of image pairs, one casually-framed and one expertly-framed, to supervise our cropping model… For each stock image, we apply the following operations: 1. Pre-processing and filtering. We filter for images that include an identifiable subject (e.g., person in portraiture; Fig. 2a). This is done with metadata tags first and then with an object detector (Ultralytics 2023). We also discard the image if it contains too many possible subjects (e.g., > 5 ). For simplicity, if there are multiple possible subjects, we select the largest one as the dominant subject (by bounding box area).” (Hong Section 3.1);
Hong teaches of pre-processing the stock images (performing data processing on a plurality of original data groups) to construct a dataset of image pairs (reads on obtain a training dataset).
Before the effective filling date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to modify Cragg in view of Rao and Hong with the feature of Hong to incorporate a dataset comprises an original image, text and a mask; and pre-processing the stock images to construct a dataset of image pairs. A person of ordinary skill in the art would do such in order to improve to improve training performance.
Cragg in view of Rao and Hong fail to teach training data including a random mask:
But Yang B does. Yang B teaches:
training data including a random mask
“we generate an arbitrarily shaped mask based on the bounding box and use it in training.” (Yang B, Section 3.2.2);
Yang B teaches of using an arbitrarily shaped mask in training, therefore the training data including a random mask (reads on training data including a random mask).
Yang B is analogous art with respect to Cragg in view Rao because they are from the same field of endeavor, namely image generation via diffusion models. Before the effective filling date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to modify Cragg in view of Rao and Hong with the feature of Yang B to incorporate a random mask in training data. A person of ordinary skill in the art would do such in order to improve performance and enables controllable editing on in-the-wild images with high fidelity.
Claim 11 is rejected using the same rationale or bases as applied to claim 2.
Claim 20 is rejected using the same rationale or bases as applied to claim 2.
Regarding claim 3, Cragg in view of Rao, Hong, and Yang B teaches the method of claim 2, wherein performing data processing on the plurality of original data groups in the original dataset to obtain the training dataset. Additionally, Cragg teaches the following:
determining the original image and the image expansion description text as a training data pair;
“initializing a diffusion model; obtaining training data including an input image, a prompt, and a ground-truth expanded image” (Cragg, ¶ 5); “In some examples, an input image depicts a person taking a picture of a cityscape in San Francisco. The prompt of the input image is “a person taking a picture of a cityscape, award winning photo, optical illusion, anamorphic widescreen, photograph of San Francisco, built on a steep hill, detailed.” Additionally, a ground-truth expanded image is an expanded version of the input image. The ground-truth expanded image includes additional content in one or more outpainted regions compared to the input image.” (Cragg, ¶ 147);
Cragg teaches of obtaining training data (reads on determining as a training data pair) comprising of an input image (reads on original image) and a prompt (reads on image expansion description text). For example, input image depicts a person taking a picture of a cityscape in San Francisco and the prompt is “a person taking a picture of a cityscape, award winning photo, optical illusion, anamorphic widescreen, photograph of San Francisco, built on a steep hill, detailed.” (reads on image expansion description text).
Yang B. further teaches:
determining a target region in the original image that matches the random mask, and performing a masking operation on a region in the original image other than the target region to obtain a masked image;
PNG
media_image6.png
490
380
media_image6.png
Greyscale
(Yang B, Figure 4);
“we generate an arbitrarily shaped mask based on the bounding box and use it in training.” (Yang B, Section 3.2.2);
Yang B teaches of determine bounding boxes of the object in the image (reads on determining a target region in the original image) in which the original image matches the arbitrarily shaped random mask See Figure 4. Yang B teaches a masking operation on a region in the original image other than the target region to obtain a masked image. In Figure 4, the original image shows a kitten in a field. The yellow bounding box around the kitten represents the target region (reads on determining a target region in the original image that matches the random mask). To obtain the masked image, a masking operation is performed on the background, the region other than the target region, leaving the kitten clear while the background is altered or greyed out. The only kitten is greyed out and the background is in color (reads on performing a masking operation on a region in the original image other than the target region to obtain a masked image. The kitten is in the target zone. The background is a region in the original image other than the target region).
performing a cropping operation on the target region to obtain a cropped image;
PNG
media_image7.png
291
641
media_image7.png
Greyscale
(Yang B, Figure 4);
Yang B teaches of cropping an image on the target region, the yellow box to obtain a cropped image, xr.
training data group comprises of a random mask
“we generate an arbitrarily shaped mask based on the bounding box and use it in training.” (Yang B, Section 3.2.2);
Yang B teaches of using an arbitrarily shaped mask in training, therefore the training data including a random mask (reads on training data group comprises of a random mask).
Before the effective filling date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to modify Cragg in view of Rao, Hong, and Yang B with the feature of Yang B to incorporate determining a bounding box in the original image that matches an arbitrarily shaped random mask; performing a masking operation on a region in the original image other than the target region to obtain a masked image; and performing a cropping operation on the target region to obtain a cropped image; and using an arbitrarily shaped mask in training. A person of ordinary skill in the art would do such in order to improve performance and enables controllable editing on in-the-wild images with high fidelity.
Hong continues and teaches:
determining the original image, the mask, the description text, the cropped image, and the masked image as a training data pair; and
PNG
media_image8.png
412
1435
media_image8.png
Greyscale
(Hong, Figure 2);
Hong teaches of training data comprises an original image (A), description text (B), a mask (C), a mask obtained by masking the original image based on the mask (D), and a cropped image obtained through cropping based on the original image and the mask (E).
constructing the training dataset based on a plurality of training data pairs corresponding to the plurality of original data groups.
PNG
media_image9.png
338
334
media_image9.png
Greyscale
(Hong, Figure 1)
“Our proposed method, GenCrop, addresses this challenge by combining a readily available dataset of stock images with powerful, pre-trained image generation models to synthesize the required inputs. Specifically, we use text-to-image diffusion to “out-paint” (i.e., outward pixel generate or outward inpaint) stock images and generate plausible uncropped-and-cropped pairs (Fig. 1). By scaling this automatic process, we can generate a large and diverse set of images to train our subject-aware cropping model.” (Hong, Section 1);
Hong teaches generating a large and diverse set of images (reads on constructing the training dataset) based on dataset of stock images (a plurality of training data pairs corresponding to the plurality of original data groups. The dataset of stock images comprises of different kinds of professional images. So, a plurality of training data pairs, the data of stock images, corresponding to the plurality of original data groups, different kinds of professional images).
Before the effective filling date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to modify Cragg in view of Rao, Hong, and Yang B with the feature of Hong to incorporate training data comprises an original image, text, a mask, a mask obtained by masking the original image based on the mask, and a cropped image obtained through cropping based on the original image and the mask; and generating a large and diverse set of based on dataset of stock images. A person of ordinary skill in the art would do such in order to improve to improve training performance.
Claim 12 is rejected using the same rationale or bases as applied to claim 3.
Regarding claim 4, Cragg in view of Rao, Hong, and Yang B teaches the method of claim 2, wherein iteratively training the preset model to be trained through the training dataset to obtain the target model and additionally teaches the following. Cragg teaches
determining a text feature vector corresponding to the image expansion description text,
“The prompt is then fed to an image generation network (e.g., a diffusion model) to generate an expanded image that includes an outpainted region.” (Cragg, ¶ 3); “The text prompt 335 can be encoded using a text encoder 340 (e.g., a multi-modal encoder) to obtain guidance features 345 in guidance space 350.” (Cragg, ¶ 340);
Cragg teaches of the text prompt (reads on the image expansion description text) is encoded to obtain guidance features (reads on text feature vector) in guidance space. The text vector is guidance features inside the guidance space after the text prompt goes through the text encoder (reads on determining a text feature vector corresponding to the image expansion description text).
Input image is a cropped image
“In some cases, the input image is a cropped image out of the ground-truth expanded image.” (Cragg, ¶ 147);
performing an iterative training operation on the model to be trained based on a plurality of feature data groups until the model to be trained satisfies a preset convergence condition, so as to obtain the trained target model.
PNG
media_image1.png
598
666
media_image1.png
Greyscale
(Cragg, Figure 15)
“FIG. 15 shows an example of a method for training a machine learning model according to aspects of the present disclosure.” (Cragg, ¶ 21); “the system at each stage n, starting with stage N, a reverse diffusion process is used to predict the image or image features at stage n−1. For example, the reverse diffusion process can predict the noise that was added by the forward diffusion process, and the predicted noise can be removed from the image to obtain the predicted image. In some cases, an original image is predicted at each stage of the training process. In some cases, the operations of this step refer to, or may be performed by, a training component as described with reference to FIG. 2.” (¶ 142);
Cragg trains the model through a repeated loop in which noise is added to a training image and the reverse diffusion process predicts the image. The predication is compared to the actual image. Training is done based on the input image, a prompt, and a ground truth (reads on model to be trained based on a plurality of feature data groups). The system updates the parameter according using gradient descent (reads on performing an iterative training operation). Training is done until N stages are reached (reads on until the model to be trained satisfies a preset convergence condition, so as to obtain the trained target model. A preset number of iteration stages is a preset convergence condition.)
Rao continues and teaches the following:
and determining an image feature vector corresponding to the input image;
“The image encoder model 402 processes the input image 201 and extracts an image feature vector, such as by using a series of convolutional layers alternated with maximum pooling layers followed by a series of fully-connected (FC) layers.” (Rao, ¶ 75);
Rao teaches extracting image feature vector (reads on reads on extracting the image feature vector) using a series of convolutional layers alternated with maximum pooling layers followed by a series of fully-connected (FC) layers (reads on based on a preset multi-modal pre-trained neural network and/or a preset large vision model)
determining a first latent space vector corresponding to the original image, and
“a text encoder may process the enhanced text prompt 205 for the input image 201 and produce a latent representation that is provided as context to the outpainting model 206 in order to ensure that the outpainted canvas is in line with the content of the input image” (Rao, ¶ 77);
Rao teaches of determining a latent representation (reads on determining a first latent space vector) ensuring that the outpainted canvas is in line with the content of the input image (reads on corresponding to the original image)
determining a second latent space vector corresponding to the masked image;
“The masked input image 401 is processed using a convolution model 404 to produce a feature map 405 for the masked input image 401.” (Rao, ¶ 75);
Rao teaches of the masked input image (reads on mask image) is processed using a convolution model to produce a feature map (reads on a second latent space vector. A feature map can be an intermediate grid of values evaluated using a secondary set of compressed, learned directional coordinates in model).
determining the text feature vector, the image feature vector, the first latent space vector, the second latent space vector, and the mask as a feature data group; and
PNG
media_image10.png
453
1020
media_image10.png
Greyscale
“The input image 201 is processed using an encoder model 402 to produce an image representation 403 in a latent space…The masked input image 401 is processed using a convolution model 404 to produce a feature map 405 for the masked input image 401.” (Rao, ¶ 75); “A denoise U-Net 410 is run a number of times (T) on the noisy latent image representation 409, personalization features 411 from the personalization block 207, and text embedding 412 based on the enhanced text prompt 205… In some cases, a text encoder may process the enhanced text prompt 205 for the input image 201 and produce a latent representation that is provided as context to the outpainting model 206 in order to ensure that the outpainted canvas is in line with the content of the input image” (Rao, ¶ 77);
Rao teaches a text embedding (reads on text feature vector), an image representation 403 (reads on image feature vector), a latent representation based on the input image (reads on the first latent space vector), a feature map 404/405 (reads on the second latent space vector), and a masked input image 401 (reads on mask image). Rao further teaches processing a noisy latent image representation 409, personalization features 411, and the text embedding 412 through a denoise U-Net 410 (reads on processing a feature data group).
Before the effective filling date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to modify Cragg in view of Rao, Hong, and Yang B with the feature of Rao to incorporate a latent representation which ensures that the outpainted canvas is in line with the content of the input image; processing the masked input image using a convolution model to produce a feature map; including text embedding, image representation, a latent representation based on the input image, a feature map, and masked input image as training input for the denoise U-Net. A person of ordinary skill in the art would do such in order to improve the quality of the outpainted image.
Yang B also teaches:
a random mask
“we generate an arbitrarily shaped mask based on the bounding box and use it in training.” (Yang B, Section 3.2.2);
Yang B teaches of arbitrarily shaped mask (reads on a random mask).
Before the effective filling date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to modify Cragg in view of Rao, Hong, and Yang B with the feature of Yang B to incorporate a random mask in the feature data group. A person of ordinary skill in the art would do such in order to improve performance and enables controllable editing on in-the-wild images with high fidelity.
Claim 13 is rejected using the same rationale or bases as applied to claim 4.
Regarding claim 8, Cragg in view of Rao, Hong, and Yang B teaches method according to claim 4 wherein determining the image feature vector corresponding to the cropped image, and additionally teaches the following. Cragg teaches
Input image is a cropped image
“In some cases, the input image is a cropped image out of the ground-truth expanded image.” (Cragg, ¶ 147);
Rao further teaches:
extracting the image feature vector corresponding to the input image based on a preset multi-modal pre-trained neural network and/or a preset large vision model.
“The image encoder model 402 processes the input image 201 and extracts an image feature vector, such as by using a series of convolutional layers alternated with maximum pooling layers followed by a series of fully-connected (FC) layers.” (Rao, ¶ 75);
Rao teaches extracting image feature vector (reads on reads on determining the image feature vector) using a series of convolutional layers alternated with maximum pooling layers followed by a series of fully-connected (FC) layers (reads on based on a preset multi-modal pre-trained neural network and/or a preset large vision model)
Before the effective filling date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to modify Cragg in view of Rao, Hong, and Yang B with the feature of Rao to incorporate extracting image feature using a series of convolutional layers alternated with maximum pooling layers followed by a series of fully-connected (FC) layers. A person of ordinary skill in the art would do such in order to improve the quality of the outpainted image.
Claim 17 is rejected using the same rationale or bases as applied to claim 8.
Claim(s) 5-6 and 6-15 is/are rejected under 35 U.S.C. 103 as being unpatentable over Cragg in view of the following: Rao, Hong, Yang B, Rombach (High-Resolution Image Synthesis with Latent Diffusion Models), hereinafter referenced as Rombach, and Pan (Locate, Assign, Refine: Taming Customized Promptable Image Inpainting).
Regarding claim 5, Cragg in view of Rao, Hong, and Yang B teaches the method according to claim 4 of claim additionally teaches the following:
determining the text feature vector corresponding to the image expansion description text
“The prompt is then fed to an image generation network (e.g., a diffusion model) to generate an expanded image that includes an outpainted region.” (Cragg, ¶ 3); “The text prompt 335 can be encoded using a text encoder 340 (e.g., a multi-modal encoder) to obtain guidance features 345 in guidance space 350.” (Cragg, ¶ 340);
Cragg teaches the text prompt is used as training data and (reads on the image expansion description text) is encoded to obtain guidance features (reads on determining text feature vector) in guidance space (reads on determining a text feature vector corresponding to the image expansion description text).
Input image is a cropped image
“In some cases, the input image is a cropped image out of the ground-truth expanded image.” (Cragg, ¶ 147);
Rao further teaches:
determining an image feature vector corresponding to the input image;
“The image encoder model 402 processes the input image 201 and extracts an image feature vector, such as by using a series of convolutional layers alternated with maximum pooling layers followed by a series of fully-connected (FC) layers.” (Rao, ¶ 75);
Rao teaches extracting image feature vector (reads on reads on determining the image feature vector) using a series of convolutional layers alternated with maximum pooling layers followed by a series of fully-connected (FC) layers (reads on based on a preset multi-modal pre-trained neural network and/or a preset large vision model) and determining an image feature vector corresponding to the input image;
“The image encoder model 402 processes the input image 201 and extracts an image feature vector, such as by using a series of convolutional layers alternated with maximum pooling layers followed by a series of fully-connected (FC) layers.” (Rao, ¶ 75);
Rao teaches extracting image feature vector (reads on reads on determining the image feature vector) using a series of convolutional layers alternated with maximum pooling layers followed by a series of fully-connected (FC) layers (reads on based on a preset multi-modal pre-trained neural network and/or a preset large vision model)
Before the effective filling date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to modify Cragg in view of Rao, Hong, and Yang B with the feature of Rao to incorporate determining an image feature vector corresponding to the input imag. A person of ordinary skill in the art would do such in order to improve the quality of the outpainted image.
Yang B further teaches the following:
training data group comprises of a random mask
“we generate an arbitrarily shaped mask based on the bounding box and use it in training.” (Yang B, Section 3.2.2);
Yang B teaches of using an arbitrarily shaped mask in training, therefore the training data including a random mask (reads on training data group comprises of a random mask).
Before the effective filling date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to modify Cragg in view of Rao, Hong, and Yang B with the feature of Yang B to incorporate a random mask in the training data group. A person of ordinary skill in the art would do such in order to improve performance and enables controllable editing on in-the-wild images with high fidelity.
However, Cragg in view of Rao, Hong, and Yang B the following: after determining training data, calculating cross-attention information based on the text feature vector and the image feature vector; and determining the cross-attention information, the first latent space vector, the second latent space vector, and the random mask as a training data group.
Rombach does. Rombach teaches the following:
after determining text vector and image vector, calculating cross-attention information based on the text feature vector and the image feature vector; and
PNG
media_image11.png
518
1149
media_image11.png
Greyscale
“Figure 3. We condition LDMs either via concatenation or by a more general cross-attention mechanism”
(Rombach, Figure 3)
“This can be implemented with a conditional denoising autoencoder E0(zt, t, y) and paves the way to controlling the synthesis process through inputs y such as text [68], semantic maps [33, 61] or other image-to-image translation tasks [34]…intermediate representation τθ(y) ∈ RM× dτ” (Rombach, Section 3.3);
Rombach teaches the latent diffusion models is conditioned with cross-attention mechanism using encoded y inputs which include text, semantic maps or other image-to-image translation tasks (reads on text vector and image vector). This implies that the y inputs training data need to be determined prior to the conditioning of the LDM (reads after determining training data, calculating cross-attention information). Additionally, the intermediate representation of text and image (reads on based on the text feature vector and the image feature vector) is processed through cross-attention, the product of this operation is cross-attention information based on the text feature vector and the image feature vector.
Rombach BASE is analogous art with respect to Cragg in view of Rao and Hong because they are from the same field of endeavor, namely image generation using diffusion models. Before the effective filling date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to modify Cragg in view of Rao and Hong with the feature of Rombach to incorporate processing intermediate representation of text and image through cross-attention to generate cross-attention information based on the text feature vector and the image feature vector. A person of ordinary skill in the art would do such in order to improve diffusion model training given computational resources while retaining their quality and flexibility.
However, Cragg in view of Rao, Hong, Yang B, and Rombach fail to teach after determining text vector and image vector, determining the cross-attention information, the first latent space vector, the second latent space vector, and the mask as a training data group.
after determining text vector and image vector, determining the cross-attention information, the first latent space vector, the second latent space vector, and the mask as a training data group
“the source image is first encoded into latent space and concatenated with the noise input, along with the mask prompt. This stage compels the model to seamlessly inpaint the masked region while keeping the background unaltered. Then, a decoupled cross-attention mechanism is designed to effectively guide the diffusion process under the joint control of the text prompt and the image prompt, ensuring that the guidance process conforms to the semantics of the local textual description and the coarse-grained subject reference.” (Pan, Section 1);
PNG
media_image12.png
804
1146
media_image12.png
Greyscale
PNG
media_image13.png
785
919
media_image13.png
Greyscale
PNG
media_image14.png
878
1552
media_image14.png
Greyscale
Text prompt is encoded to generate a text vector to be processed into the text cross attention. The image prompt is encoded to generate an image vector to be processed into the image cross attention. This is denoted in the green and red boxes in Figure 3. Equation 2 defines z~ a training tensor which concatenates the following into one grouping: the noisy image z (reads on the first latent space), m* is the mask, z_s the blanked-out image (reads on the second latent space). Equation 12 then defines the training rule for the model wherein t is the noisy step, x_obj is the cropped image, s is the text prompt, and training tensor z~ previously defined by Equation 2. The text and cropped image get processed through the cross-attention. All the terms are determined and then feed into the model for training. Therefore, only after determining text vector and image vector, does the the model takes in two latent spaces, a mask, and cross-attention information as a training data group (reads on after determining text vector and image vector, determining the cross-attention information, the first latent space vector, the second latent space vector, and the mask as a training data group).
Pan is analogous art with respect to Cragg in view of Rao, Hong, Yang B, and Rombach because they are from the same field of endeavor, namely image generation using diffusion models. Before the effective filling date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to modify Cragg in view of Rao, Hong, Yang B, and Rombach with the feature of Pan to incorporate after determining text vector and image vector, determining the cross-attention information, the first latent space vector, the second latent space vector, and the mask as a training data group. A person of ordinary skill in the art would do such in order to improve diffusion model for seamless image generation.
Claim 15 is rejected using the same rationale or bases as applied to claim 5.
Allowable Subject Matter
Claim 6-7 and 15-16 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to DUNE NGUYEN whose telephone number is (571)272-8919. The examiner can normally be reached M-TH 7:00AM - 5:00PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Devona E Faulk can be reached at (571) 272-7515. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/DUNE NGOC NGUYEN/Examiner, Art Unit 2618
/DEVONA E FAULK/Supervisory Patent Examiner, Art Unit 2618