DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
This office action is responsive to the above identified application filed January 25, 2024.
Claims 1-20 are pending, all examined and rejected.
Drawings
The drawings are objected to under 37 CFR 1.83(a) because they contain details and features which are blurry and hard to read (See at least Fig. 6B and Fig. 9).
Corrected drawing sheets in compliance with 37 CFR 1.121(d) are required in reply to the Office action to avoid abandonment of the application. Any amended replacement drawing sheet should include all of the figures appearing on the immediate prior version of the sheet, even if only one figure is being amended. The figure or figure number of an amended drawing should not be labeled as “amended.” If a drawing figure is to be canceled, the appropriate figure must be removed from the replacement sheet, and where necessary, the remaining figures must be renumbered and appropriate changes made to the brief description of the several views of the drawings for consistency. Additional replacement sheets may be necessary to show the renumbering of the remaining figures. Each drawing sheet submitted after the filing date of an application must be labeled in the top margin as either “Replacement Sheet” or “New Sheet” pursuant to 37 CFR 1.121(d). If the changes are not accepted by the examiner, the applicant will be notified and informed of any required corrective action in the next Office action. The objection to the drawings will not be held in abeyance.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claims 1-20 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Zhou, Yufan, et al. "Shifted diffusion for text-to-image generation." 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, 2023. [Yufan]
Regarding claim 1, Yufan teaches a method for training a neural network model to transform a text description into nontextual data, comprising:
receiving, via a communication interface, a dataset comprising a first subset of training data without labels and a second subset of training data with labels; (See page 2, column 1, paragraph 3, “Our method is general and can be applied in different settings of text-to-image generation, e.g., it naturally enables semi-supervised and language-free text-to-image generation;”. See Page 3, column 1, paragraph 3, “In addition, our framework naturally enables semi-supervised training, i.e., the training dataset can be a mixture of image-text pairs and pure images that are not captioned”, - using semi-supervised training means the model is using a labeled and unlabeled training data)
training the neural network model according to first loss computed using the first subset of training data without labels. (See page 3, column 1, paragraph 2, “In this setting, the image-text pairs will be used to train the prior model, and all the pure images will be used to train the decoder. See page 4, column 1, equation (4)
PNG
media_image1.png
104
543
media_image1.png
Greyscale
.
See page 4, column 2, algorithm 1.
PNG
media_image2.png
566
558
media_image2.png
Greyscale
”
– pure images/unlabeled data are being used to train the neural network mode/decoder)
retraining the trained neural network model according to a second loss computed using the second subset of training data with labels. (See page 4, column 2, paragraph 2, “Specifically, we propose to update
PNG
media_image3.png
24
63
media_image3.png
Greyscale
during training with the following loss function”. See page 4, column 2, equation (7) and algorithm 1
PNG
media_image4.png
120
577
media_image4.png
Greyscale
. See page 4, column 1, paragraph 3 “Let z0 and y be the ground truth image embedding and its associated text caption, respectively. For each pair of (z0, y), we select its corresponding Gaussian
PNG
media_image5.png
24
31
media_image5.png
Greyscale
by the top-1 cosine similarity as
PNG
media_image6.png
59
298
media_image6.png
Greyscale
” – The second loss
PNG
media_image7.png
29
31
media_image7.png
Greyscale
is computed using
PNG
media_image8.png
22
21
media_image8.png
Greyscale
, which uses pairs of text-image/labeled data and updating
PNG
media_image9.png
26
66
media_image9.png
Greyscale
/model using the second loss.)
and according to a constraint that a third loss computing using the second subset of training data but without labels is no greater than the first loss. (See page 4, column 2, algorithm 1.- consider
PNG
media_image10.png
23
27
media_image10.png
Greyscale
is the first loss, updating
PNG
media_image11.png
16
16
media_image11.png
Greyscale
/model by gradient descent w.r.t
PNG
media_image10.png
23
27
media_image10.png
Greyscale
means minimizing
PNG
media_image10.png
23
27
media_image10.png
Greyscale
. After minimizing
PNG
media_image10.png
23
27
media_image10.png
Greyscale
, minimized
PNG
media_image10.png
23
27
media_image10.png
Greyscale
(third loss) is no greater than the first loss.)
deploying the retrained neural network model on one or more hardware processors to generate the non-textual data according to the text description. (See page 6, column 2, paragraph 1, “Specifically, we first generate four images for each prompt, then ten random human laborers are asked to judge which model is better (or comparable) in terms of image-text alignment and fidelity. The results are shown in Figure 6, where our model is shown to perform better in the two-evaluation metrics consistently.” – After training, the model is deployed and used to generate images based on text prompts.)
Regarding claim 2,
As discussed with regard to claim 1, Yufan teaches all of the limitations.
Yufan further teaches the neural network model comprise a diffusion model. (See Figure. 2 “Illustration of our framework. The trainable modules are colored in green, and the frozen modules are colored in blue. ting across different downstream datasets”. See page 3, column 1, paragraph 1, “The decoder can be either a diffusion model or a generative adversarial network (GAN). Note that if one chooses the decoder as a hierarchical diffusion model and makes it conditioned on both image embedding and text, our final structure will be similar to DALL-E 2” – the framework/model includes a decoder, which is a diffusion model.)
Regarding claim 3,
As discussed with regard to claim 2, Yufan teaches all the limitations.
Yufan further teaches the neural network model comprises:
Adding a noise term to a training data sample from the first subset to generate a noised sample; (See page 5, column 1, paragraph 2 Decoder “Specifically, three diffusion models are trained on 64, 256, and 1024 resolutions, respectively. All the 900M images are used to train the decoder. During training, each image is processed by three different pre-trained CLIP models: ViT-B/16, ViT-B/32, and RN-101. The outputs from these three models will be concatenated into a single 1536-dimensional embedding, which is then projected into eight vectors and fed into the decoder.”. See page 5, column 2, paragraph 1. “Our shifted diffusion model is a decoder only transformer whose input is a sequence consisting of encoded text from T5 [18], CLIP text embedding, an embedding representing diffusion timestep, an embedding representing the index of corresponding Gaussian, a noised CLIP image embedding and a final embedding which will be used to predict the target CLIP image embedding.”) –CLIP models add noise terms to CLIP images to generate vector of noised CLIP images. These vectors are used to fed to the decoder in the algorithm (1), which contains
PNG
media_image12.png
31
26
media_image12.png
Greyscale
.)
Iteratively predicting, by the diffusion model, a first predicted noise term from the noised sample. (See page 5, column 2, paragraph 1. “CLIP text embedding, an embedding representing diffusion timestep, an embedding representing the index of corresponding Gaussian, a noised CLIP image embedding and a final embedding which will be used to predict the target CLIP image embedding.” – predicted image is obtained by using vector of noised images generated by the shifted diffusion model implemented/optimized using algorithm (1).)
Regarding claim 4,
As discussed with regard to claim 3, Yufan teaches all the limitations.
Yufan further teaches the first loss is computed based on a difference between the first predicted noise term and the added noise term. (See page 4, column 1, equation (4).
PNG
media_image13.png
98
541
media_image13.png
Greyscale
- the loss function is computed based on
PNG
media_image14.png
21
21
media_image14.png
Greyscale
,
PNG
media_image15.png
20
49
media_image15.png
Greyscale
and
PNG
media_image16.png
18
22
media_image16.png
Greyscale
, which are the target image, first sample and initial image. The added noise term
PNG
media_image17.png
28
91
media_image17.png
Greyscale
and
PNG
media_image15.png
20
49
media_image15.png
Greyscale
are used in computing the loss.)
Regarding claim 5,
As discussed with regard to claim 2, Yufan teaches all the limitations.
Yufan further teaches retraining the neural network model comprises:
Adding a noise term to a training data sample from the second subset to generate a noised sample; (See page 5, column 1, paragraph 2 Decoder “Specifically, three diffusion models are trained on 64, 256, and 1024 resolutions, respectively. All the 900M images are used to train the decoder. During training, each image is processed by three different pre-trained CLIP models: ViT-B/16, ViT-B/32, and RN-101. The outputs from these three models will be concatenated into a single 1536-dimensional embedding, which is then projected into eight vectors and fed into the decoder.”. See page 5, column 2, paragraph 1. “Our shifted diffusion model is a decoder only transformer whose input is a sequence consisting of encoded text from T5 [18], CLIP text embedding, an embedding representing diffusion timestep, an embedding representing the index of corresponding Gaussian, a noised CLIP image embedding and a final embedding which will be used to predict the target CLIP image embedding.”– the CLIP models add noise terms to CLIP images to generate vector of noised CLIP images, are used to fed to the decoder in the algorithm (1), which contains
PNG
media_image18.png
33
30
media_image18.png
Greyscale
. )
Iteratively predicting by the diffusion model, a second predicted noise term from the noised sample; and (See page 5, column 2, paragraph 1. “CLIP text embedding, an embedding representing diffusion timestep, an embedding representing the index of corresponding Gaussian, a noised CLIP image embedding and a final embedding which will be used to predict the target CLIP image embedding.” – predicted image is obtained by using vector of noised images generated by the shifted diffusion model implemented/optimized using algorithm (1).)
Iteratively predicting by the diffusion model, a third predicted noise term from the noised sample conditioned on a text label associated with the training data sample. (See page 4, column 2, equation (7)
PNG
media_image19.png
114
570
media_image19.png
Greyscale
. See page 5, column 2, paragraph 1, “We train two variants on datasets with different scales: one prior is trained on the full 900M image-text pairs; the other one is trained on CC15M‡, which is a subset of our full dataset.” See page 4, column 1 “Let z0 and y be the ground truth image embedding and its associated text caption, respectively. For each pair of (z0, y), we select its corresponding Gaussian
PNG
media_image20.png
21
25
media_image20.png
Greyscale
by the top-1 cosine similarity as
PNG
media_image21.png
46
365
media_image21.png
Greyscale
”, and equation (7) – The shift diffusion, which is a decoder only transformer uses pair of image-text label
PNG
media_image20.png
21
25
media_image20.png
Greyscale
to predict the sample, and the images are noised CLIP images.)
Regarding claim 6,
As discussed with regard to claim 5, Yufan teaches all the limitations.
Yufan further teaches the second loss is computed based on a difference between the second predicted noise term and the added noise term; and (See page 4, column 2, equation (7) – the loss function is computed based on
PNG
media_image22.png
21
22
media_image22.png
Greyscale
and
PNG
media_image23.png
29
112
media_image23.png
Greyscale
, which are the predicted term and the added noise term)
The third loss is computed based on a difference between the third predicted noise term and the added noise term. (See page 4, column 1, equation 4
PNG
media_image24.png
102
537
media_image24.png
Greyscale
. The loss is computed using
PNG
media_image25.png
24
23
media_image25.png
Greyscale
and
PNG
media_image26.png
24
21
media_image26.png
Greyscale
, which is the ground truth and the previous predicted image, and the third loss is computed after minimizing the first loss using gradient descent. In some cases, the loss is computed based on the ground truth and the previous predicted image.)
Regarding claim 7,
As discussed with regard to claim 6, Yufan teaches all the limitations.
Yufan further teaches retraining the neural network model comprise:
Updating at a training iteration, parameters of the diffusion model such that the third loss conditioned on the parameters of the diffusion model is no greater than the first loss conditioned on the parameters of the diffusion model. (See page 4, column 2, algorithm 1. – iteratively train the model while minimizing
PNG
media_image27.png
23
19
media_image27.png
Greyscale
/parameter using gradient descent so the parameter of the third loss is no greater than the first loss.
Regarding claim 8,
As discussed with regard to claim 1, Yufan teaches all the limitations.
Yufan further teaches the non-textual data comprises any of:
Biological structure data; time-series data; video-motion data. (See page 2, Figure. 1 “We propose Corgi, a novel diffusion model designed for flexible text-to-image generation which can “bridge the gap”. – broadest reasonable interpretation of biological structure data includes the corgi images.)
With regard to claims 9 and 16,
These claims are similar in scope to claim 1 and are rejected under a similar rationale.
With regard to claims 10 and 17,
These claims are similar in scope to claim 2 and are rejected under a similar rationale.
With regard to claims 11 and 18,
These claims are similar in scope to claim 3 and are rejected under a similar rationale.
With regard to claims 12 and 19,
These claims are similar in scope to claim 4 and are rejected under a similar rationale.
With regard to claims 13 and 20,
These claims are similar in scope to claim 5 and are rejected under a similar rationale.
With regard to claim 14,
The claim is similar in scope to claim 6 and is rejected under a similar rationale.
With regard to claim 15,
The claim is similar in scope to claim 7 and is rejected under a similar rationale
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure:
US 12675917 B2
Identity-preserving image generation using diffusion models
Kansy; Manuel Jakob et al.
US 12651399 B2
Training data sampling for neural networks
Lorraine; Jonathan Peter et al.
US 12555275 B2
Personalized text-to-image diffusion model
Aberman; Kfir et al.
US 12462348 B2
Multimodal diffusion models
Ham; Cusuh et al.
US 20230095092 A1
DENOISING DIFFUSION GENERATIVE ADVERSARIAL NETWORKS
Xiao; Zhisheng et al.
Li, Yan, et al. "Generative time series forecasting with diffusion, denoise, and disentanglement." Advances in Neural Information Processing Systems 35 (2022): 23009-23022.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to MICHAEL HOANG LE whose telephone number is (571)270-7292. The examiner can normally be reached Monday-Friday 8:00 am - 5 pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Matthew Ell can be reached at 5712703264. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/MICHAEL HOANG LE/Examiner, Art Unit 2141
/MATTHEW ELL/Supervisory Patent Examiner, Art Unit 2141