DETAILED ACTION
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-3, 8-10, 13, and 17-20 are rejected under 35 U.S.C. 103 as being unpatentable over Feng et al. (Pub. No. US 2021/0034981) in view of Shi et al. (Pub. No. US 20240355022).
Regarding claims 1, 17, and 20, Feng teaches a method performed by one or more computers, the method comprising: training (perform optimization and adjustment on the decoding RNN 202) an image representation neural network (image caption model) that is configured to receive an input image (image sample) and to process the input image (image sample) to generate a representation of the input image (sentence used for describing the image sample) as a set of text tokens (plurality of words) from a vocabulary of text tokens (word list) [Para. 79 “The method for training an image caption model provided in the embodiments of this application is described below with reference to FIG. 3. FIG. 3 is a flowchart of the method for training an image caption model according to the embodiments of this application”. Para. 82 “Decode the image eigenvector by using the decoding RNN, to obtain a sentence used for describing the image sample”; Para. 84-85; and para. 72 “reconstruct an image according to the sentence obtained through decoding, determine a similarity degree between the reconstructed image and the image sample 211, and perform optimization and adjustment on the decoding RNN 202 according to the similarity degree.”], the training comprising: obtaining a set (training set) of one or more training images (image sample) [Para. 40 and 80];
for each training image in the set: processing the training image using the image representation neural network image caption model) to generate a training representation (sentence obtained through decoding) of the training image (image sample) as a set of text tokens (plurality of words) from the vocabulary of text tokens (word list) [Para. 40, 8p, 82, and 84-85]; and
Feng teaches processing a text input (sentence obtained through decoding) comprising the set of text tokens (plurality of words) in the training representation (sentence obtained through decoding) of the training image [Para. 84-84 and 73].
However, Feng doesn’t explicitly teach using a text-conditioned image generation neural network to generate an output that defines an output image.
Shi teaches processing a text input (subject injected textual embedding) comprising the set of text tokens (textual token) in the training representation (compact textual embedding) of the training image (training image) using a text-conditioned image generation neural network (diffusion model) to generate an output (noise prediction) that defines an output image [Para. 74 “The subject encoder 410 can learn the general concept of the input images 115 by converting the images 115 to a textual token. The subject encoder 410 can map the subject of the images 115 to the compact textual embedding.”; Para. 81 “the embedding of identifier, {circumflex over (V)}, 445 can be replaced with the concept feature vector, f.sub.c, generated by the subject encoder 410, to obtain the subject injected textual embedding c. This embedding can be the condition in the cross-attention layers 462 of a U-Net 460 in a text-to-image diffusion model”; Para. 166 “a noise prediction can be generated based on the subject embedding using the image generation model. The noise prediction can be generated using a diffusion model.”; Para. 132, 136, 159, 160, and 164].
It would have been obvious to one of ordinary skill in the art before the effective filing date to modify Feng’s image caption training method by augmenting it’s image feature space reconstruction stage with Shi’s processing of text input (subject injected textual embedding) through a text conditioned image generation neural network (diffusion model) to produce an output (noise prediction) defining an output image. This modification improves Feng by providing image generative feedback from the generated textual representation rather than limiting reconstruction feedback to comparison of image feature vectors.
Feng teaches training the image representation neural network (image caption model) for each training image (image sample) in the set (training set), a difference (similarity degree) between (i) a ground truth output (image eigenvector of the image sample) corresponding to the training image and (ii) the output (image eigenvector obtained through mapping) of an image generation neural network generated by processing the text input (sentence obtained through decoding) comprising the set of text tokens (plurality of words) in the training representation (sentence obtained through decoding) of the training image [Para. 72, 73, 154, 162, and 163].
However, Fend doesn’t explicitly teach an objective function including a first term that measures the difference and having a text-conditioned image generation neural network.
Shi teaches training (trained) the image representation neural (subject encoder) on an objective function (loss function) that includes a first term denoising loss) that measures, for each training image (training image) in the set, a difference (comparison) between (i) a ground truth output (ground truth noise map) corresponding to the training image (training image) and (ii) the output (noise prediction) of the text-conditioned image generation neural network (diffusion model) generated by processing the text input (final embedding) comprising the set of text tokens (textual token) in the training representation (compact textual embedding) of the training image [Para. 74 “The subject encoder 410 can learn the general concept of the input images 115 by converting the images 115 to a textual token.”; Para. 132 “The subject of the input training images can be learned by the subject image encoder 410 in FIG. 4, and mapped to a compact textual embedding”; Para. 136 “This final embedding can be the condition in the cross-attention layers of the text-to-image diffusion model”; Para. 139 “a single denoising step can be used for training, where the training loss can be calculated, as a loss function for an image at a single noise level”; Para. 138 “The model can be optimized with only the denoising loss of the diffusion model”; para. 148, 149, 166 and 167].
It would have been obvious to one of ordinary skill in the art before the effective filing date to modify Feng’s image caption model training loop by incorporating Shi’s teaching of conditioning a text-conditioned image generation neural network (diffusion model) with an image-derived text input (final embedding) comprising a set of text tokens (textual token), generating an output (noise prediction), comparing the output (noise prediction) with a ground truth output (ground truth noise map) using an objective function (loss function) having a first term (denoising loss), and training the image representation neural network (subject encoder) based on that result. This modification improves Feng by replacing its feature space only reconstruction award with differentiable text-conditioned image generation feedback, thereby encouraging the image-caption model to produce text representations that retina image information needed for reconstruction.
Regarding claims 2 and 18, Feng doesn’t explicitly teach the claim limitations.
Shi teaches wherein the text-conditional image generation neural network (pre-trained text to image model) has been pre-trained on a text-conditional image generation task (trained on text-image pairs) and wherein training the image representation neural network on the objective function (denoising loss) comprises training the image representation neural network (subject encoder 410) while holding the text-conditional image generation neural network fixed (frozen in the full training procedure) [Para. 23, 55, 84 and 132].
It would have been obvious to one of ordinary skill in the art before the effective filing date to modify Feng’s image caption training method by using Shi’s frozen pre-trained text-to-image model as the text-conditioned image generation network while training the image-to-text network based on the reconstruction objective. This modification improves Feng by preserving the pretrained generators previously learned image synthesis and language conditioning capabilities while supplying reliable image supervision.
Regarding claims 3 and 19, Feng doesn’t explicitly teach the claim limitations.
Shi teaches wherein the text-conditional image generation neural network is a text-conditional diffusion neural network and wherein the output of the text-conditional diffusion neural network (diffusion model) that defines the output image is a denoising output (prediction) and the ground truth output is a ground truth denoising output (ground truth noise map) corresponding to the training image [Para. 66-68, 136-138, 148, 149, 164-167].
It would have been obvious to one of ordinary skill in the art before the effective filing date to modify Feng’s image caption model training loop by incorporating Shi’s teaching of conditioning a text-conditioned image generation neural network (diffusion model) with an image-derived text input (final embedding) comprising a set of text tokens (textual token), generating an output (noise prediction), comparing the output (noise prediction) with a ground truth output (ground truth noise map) using an objective function (loss function) having a first term (denoising loss), and training the image representation neural network (subject encoder) based on that result. This modification improves Feng by replacing its feature space only reconstruction award with differentiable text-conditioned image generation feedback, thereby encouraging the image-caption model to produce text representations that retina image information needed for reconstruction.
Regarding claim 6, Feng doesn’t explicitly teach the claim limitations.
Shi teaches wherein the text-conditional diffusion neural network (diffusion model) comprises: a text encoder neural network (text encoder) configured to process the text input (text prompt) to generate an encoded representation of the text input (text embedding); and an image diffusion neural network (U-net) configured to generate the output image over a plurality of sampling steps (iterative denoising process) conditioned on the encoded representation of the text input (text embedding) [Para. 60, 66-68, 91, 94, and 104].
It would have been obvious to one of ordinary skill in the art before the effective filing date to modify Feng’s image caption training method by using Shi’s frozen pre-trained text-to-image model as the text-conditioned image generation network while training the image-to-text network based on the reconstruction objective. This modification improves Feng by preserving the pretrained generators previously learned image synthesis and language conditioning capabilities while supplying reliable image supervision.
Regarding claim 8, Feng teaches wherein the image representation neural network (image caption model) comprises: an image backbone neural network (encoding CNN 201) that is configured to process the input image (image sample 211) to generate a feature representation of the input image (image eigenvector 214); and an encoder neural network (decoding RNN 202) configured to process the feature representation of the input image (image eigenvector 214) to generate the representation of the input image (sentence 213) [Para. 37-39, 51-55 and 79-82].
Regarding claim 9. Feng doesn’t explicitly teach the claim limitations.
Shi teaches wherein the image backbone neural network (pre-trained multimodal encoder as the backbone) has been pre-trained on an image representation learning task and wherein training the image representation neural network on the objective function comprises training the encoder neural network (fully connected layer) while holding the image backbone neural network fixed (only the fully connected layer of the image encoders are updated) [Para. 60 and 132].
It would have been obvious to one of ordinary skill in the art before the effective filing date to modify Feng’s image caption training method by using Shi’s frozen pre-trained text-to-image model as the text-conditioned image generation network while training the image-to-text network based on the reconstruction objective. This modification improves Feng by preserving the pretrained generators previously learned image synthesis and language conditioning capabilities while supplying reliable image supervision.
Regarding claim 10 Feng doesn’t explicitly teach the claim limitations.
Shi teaches wherein training the image representation neural network (subject image encoder 410) on the objective function comprises training the encoder neural network and the image backbone neural network (backbone 415) [Para. 76 and 82].
It would have been obvious to one of ordinary skill in the art before the effective filing date to modify Feng’s image caption training method by using Shi’s frozen pre-trained text-to-image model as the text-conditioned image generation network while training the image-to-text network based on the reconstruction objective. This modification improves Feng by preserving the pretrained generators previously learned image synthesis and language conditioning capabilities while supplying reliable image supervision.
Regarding claim 13 Feng doesn’t explicitly teach the claim limitations.
Shi teaches after training the image representation neural network: receiving a query input (input image) for a downstream task (output image generation) that comprises a query image; processing the query image using the image representation neural network (subject encoder) to generate a representation of the query image as a set of text tokens (textual token); and providing the representation of the query image (subject embedding) as input to a downstream neural network (image generation model) configured to perform the downstream task (output-image generation) [Para. 74, 109, 155-160].
It would have been obvious to one of ordinary skill in the art before the effective filing date to modify Feng’s image caption training method by using Shi’s frozen pre-trained text-to-image model as the text-conditioned image generation network while training the image-to-text network based on the reconstruction objective. This modification improves Feng by preserving the pretrained generators previously learned image synthesis and language conditioning capabilities while supplying reliable image supervision.
Claims 4, 5, and 7 are rejected under 35 U.S.C. 103 as being unpatentable over Feng et al. (Pub. No. US 2021/0034981) in view of Shi et al. (Pub. No. US 20240355022) further in view of Wei (“ELITE: Encoding visual concepts into textual embeddings for customized text-to-image generation”).
Regarding claim 4, Feng in view of Shi doesn’t explicitly teach the claim limitation.
However, Wei teaches wherein processing a text input comprising the set of text tokens in the training representation of the training image using a text-conditioned image generation neural network to generate an output that defines an output image comprises: sampling a noise level [section 3.1 eq. (1)]; generating a noisy image from the training image by applying noise to the training image in accordance with the noise level [section 3.1 eq. (1)]; and processing the noisy image and the text input (text condition) using the text-conditional diffusion neural network (conditional diffusion model) to generate the denoising output, wherein the denoising output defines an estimate of the training image given the noisy image and the text input [section 3.1 para. 1].
It would have been obvious to one of ordinary skill in the art before the effective filing date to modify Feng’s image caption reconstruction training method, as augmented by Shi, diffusion-based image generation feedback, by incorporating Wei’s sampled time step training procedure in which Gaussin noised is applied to the encoded training image. This modification improves Feng by training the reconstruction path across different noise levels.
Regarding claim 5, Feng in view of Shi further in view of Wei teaches all claim limitations as stated above. Furthermore, Wei teaches wherein the denoising output is one of: (i) the estimate of the training image, (ii) an estimate of the noise applied to the training image to generate the noisy image, or (iii) an estimate of a v-prediction output generated from the noise and the training image [Section 3.1 , eq. 1 and related description].
Regarding claim 7, Feng in view of Shi further in view of Wei teaches all claim limitations as stated above. Furthermore, Shi teaches wherein the vocabulary of text tokens is an input vocabulary of the text encoder neural network [fig. 4, 6 and related description].
Claims 11 and 12 are rejected under 35 U.S.C. 103 as being unpatentable over Feng et al. (Pub. No. US 2021/0034981) in view of Shi et al. (Pub. No. US 20240355022) further in view of K. Shi (“Non-autoaggressive sequence-to-sequence vision-language models”).
Regarding claim 11 Feng in view of Shi doesn’t explicitly teach the claim limitation.
However, K. Shi teaches wherein the encoder neural network (NARVL decoder) has a respective learned query (query token) corresponding to each text token in the representation, wherein the encoder neural network comprises a sequence of self-attention layer blocks and an output layer block, and wherein processing the feature representation of the input image to generate the representation of the input image [Section 3.1] comprises:
processing the learned queries (learnable query tokens) through the sequence of self-attention layer blocks, wherein each self-attention layer block is configured to update the learned queries conditioned on the feature representation of the input image (outputs of the encoder) [section 3.1]; and
after processing the learned queries through the sequence of self-attention layer blocks, processing each learned query using the output layer block to generate the corresponding text token in the representation [section 3.2.1 and 3.2.3].
It would have been obvious to one of ordinary skill in the art before the effective filing date to modify Feng’s decoding RNN image-to-text component by incorporating K. Shi’s learned query parallel transformer arrangement, in which learned query tokens self-attend, are conditioned on the encoded image features, and are projected to corresponding vocabulary tokens. This modification improves Feng by generating multiple representation tokens in parallel rather than recurrently, thereby reducing token generation latency while retaining image conditioned token prediction.
Regarding claim 12, Feng in view of Shi further view K. Shi teach all claim limitations as stated above. Furthermore, K. Shi teaches the output layer block is a linear neural network layer [Section 3.2.1.-3.2.3].
Claims 14-16 are rejected under 35 U.S.C. 103 as being unpatentable over Feng et al. (Pub. No. US 2021/0034981) in view of Shi et al. (Pub. No. US 20240355022) further in view of Nguyen et al. (Pub. No. US 2025/0005293).
Regarding claim 14, Feng in view of Shi doesn’t explicitly teach the claim limitations.
Nguyen teaches wherein the query input comprises the query image (digital images) and text, wherein the downstream neural network is a large language model, and wherein providing the representation of the query image (output sequence of tokens) as input to the downstream neural network configured to perform the downstream task comprises providing the representation of the query image and the text from the query input as input to the large language model [Para. 19-22].
It would have been obvious to one of ordinary skill in the art before the effective filing date to modify Feng’s post training image representation pipeline, as augmented by Shi, by incorporating Nguyen’s assembly of the image derived token representation and accompanying query text into a prompt processed by a large language model. This modification improves Feng by enabling it’s generated image representation to supply visual context to a text processing LLM, there by allowing LLM to respond to text queries concerning image content without directly ingesting raw image pixels.
Regarding claim 15 Feng in view of Shi in view of Nguyen teaches all claim limitations. Furthermore, Nguyen teaches wherein the downstream task is a multi-modal dialogue task [fig. 4, 6 and corresponding description].
Regarding claim 16 Feng in view of Shi in view of Nguyen teaches all claim limitations. Furthermore, Nguyen teaches wherein the downstream task is a zero-shot task or a multi-modal few-shot learning task [fig. 2, 6 and corresponding description].
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SOLOMON G BEZUAYEHU whose telephone number is (571)270-7452. The examiner can normally be reached on Monday-Friday 10 AM-7 PM..
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, O’Neal Mistry can be reached on 313-446-4912. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-0101 (IN USA OR CANADA) or 571-272-1000.
/SOLOMON G BEZUAYEHU/ Primary Examiner, Art Unit 2666