Prosecution Insights
Last updated: October 01, 2026
Application No. 18/423,081

SYSTEMS AND METHODS FOR CONTROLLABLE DATA GENERATION FROM TEXT

Non-Final OA §102
Filed
Jan 25, 2024
Priority
Aug 25, 2023 — provisional 63/578,906
Examiner
LE, MICHAEL HOANG
Art Unit
Tech Center
Assignee
Salesforce Inc.
OA Round
1 (Non-Final)
Grant Probability
Favorable
1-2
OA Rounds

Office Action

§102
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . This office action is responsive to the above identified application filed January 25, 2024. Claims 1-20 are pending, all examined and rejected. Drawings The drawings are objected to under 37 CFR 1.83(a) because they contain details and features which are blurry and hard to read (See at least Fig. 6B and Fig. 9). Corrected drawing sheets in compliance with 37 CFR 1.121(d) are required in reply to the Office action to avoid abandonment of the application. Any amended replacement drawing sheet should include all of the figures appearing on the immediate prior version of the sheet, even if only one figure is being amended. The figure or figure number of an amended drawing should not be labeled as “amended.” If a drawing figure is to be canceled, the appropriate figure must be removed from the replacement sheet, and where necessary, the remaining figures must be renumbered and appropriate changes made to the brief description of the several views of the drawings for consistency. Additional replacement sheets may be necessary to show the renumbering of the remaining figures. Each drawing sheet submitted after the filing date of an application must be labeled in the top margin as either “Replacement Sheet” or “New Sheet” pursuant to 37 CFR 1.121(d). If the changes are not accepted by the examiner, the applicant will be notified and informed of any required corrective action in the next Office action. The objection to the drawings will not be held in abeyance. Claim Rejections - 35 USC § 102 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. Claims 1-20 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Zhou, Yufan, et al. "Shifted diffusion for text-to-image generation." 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, 2023. [Yufan] Regarding claim 1, Yufan teaches a method for training a neural network model to transform a text description into nontextual data, comprising: receiving, via a communication interface, a dataset comprising a first subset of training data without labels and a second subset of training data with labels; (See page 2, column 1, paragraph 3, “Our method is general and can be applied in different settings of text-to-image generation, e.g., it naturally enables semi-supervised and language-free text-to-image generation;”. See Page 3, column 1, paragraph 3, “In addition, our framework naturally enables semi-supervised training, i.e., the training dataset can be a mixture of image-text pairs and pure images that are not captioned”, - using semi-supervised training means the model is using a labeled and unlabeled training data) training the neural network model according to first loss computed using the first subset of training data without labels. (See page 3, column 1, paragraph 2, “In this setting, the image-text pairs will be used to train the prior model, and all the pure images will be used to train the decoder. See page 4, column 1, equation (4) PNG media_image1.png 104 543 media_image1.png Greyscale . See page 4, column 2, algorithm 1. PNG media_image2.png 566 558 media_image2.png Greyscale ” – pure images/unlabeled data are being used to train the neural network mode/decoder) retraining the trained neural network model according to a second loss computed using the second subset of training data with labels. (See page 4, column 2, paragraph 2, “Specifically, we propose to update PNG media_image3.png 24 63 media_image3.png Greyscale during training with the following loss function”. See page 4, column 2, equation (7) and algorithm 1 PNG media_image4.png 120 577 media_image4.png Greyscale . See page 4, column 1, paragraph 3 “Let z0 and y be the ground truth image embedding and its associated text caption, respectively. For each pair of (z0, y), we select its corresponding Gaussian PNG media_image5.png 24 31 media_image5.png Greyscale by the top-1 cosine similarity as PNG media_image6.png 59 298 media_image6.png Greyscale ” – The second loss PNG media_image7.png 29 31 media_image7.png Greyscale is computed using PNG media_image8.png 22 21 media_image8.png Greyscale , which uses pairs of text-image/labeled data and updating PNG media_image9.png 26 66 media_image9.png Greyscale /model using the second loss.) and according to a constraint that a third loss computing using the second subset of training data but without labels is no greater than the first loss. (See page 4, column 2, algorithm 1.- consider PNG media_image10.png 23 27 media_image10.png Greyscale is the first loss, updating PNG media_image11.png 16 16 media_image11.png Greyscale /model by gradient descent w.r.t PNG media_image10.png 23 27 media_image10.png Greyscale means minimizing PNG media_image10.png 23 27 media_image10.png Greyscale . After minimizing PNG media_image10.png 23 27 media_image10.png Greyscale , minimized PNG media_image10.png 23 27 media_image10.png Greyscale (third loss) is no greater than the first loss.) deploying the retrained neural network model on one or more hardware processors to generate the non-textual data according to the text description. (See page 6, column 2, paragraph 1, “Specifically, we first generate four images for each prompt, then ten random human laborers are asked to judge which model is better (or comparable) in terms of image-text alignment and fidelity. The results are shown in Figure 6, where our model is shown to perform better in the two-evaluation metrics consistently.” – After training, the model is deployed and used to generate images based on text prompts.) Regarding claim 2, As discussed with regard to claim 1, Yufan teaches all of the limitations. Yufan further teaches the neural network model comprise a diffusion model. (See Figure. 2 “Illustration of our framework. The trainable modules are colored in green, and the frozen modules are colored in blue. ting across different downstream datasets”. See page 3, column 1, paragraph 1, “The decoder can be either a diffusion model or a generative adversarial network (GAN). Note that if one chooses the decoder as a hierarchical diffusion model and makes it conditioned on both image embedding and text, our final structure will be similar to DALL-E 2” – the framework/model includes a decoder, which is a diffusion model.) Regarding claim 3, As discussed with regard to claim 2, Yufan teaches all the limitations. Yufan further teaches the neural network model comprises: Adding a noise term to a training data sample from the first subset to generate a noised sample; (See page 5, column 1, paragraph 2 Decoder “Specifically, three diffusion models are trained on 64, 256, and 1024 resolutions, respectively. All the 900M images are used to train the decoder. During training, each image is processed by three different pre-trained CLIP models: ViT-B/16, ViT-B/32, and RN-101. The outputs from these three models will be concatenated into a single 1536-dimensional embedding, which is then projected into eight vectors and fed into the decoder.”. See page 5, column 2, paragraph 1. “Our shifted diffusion model is a decoder only transformer whose input is a sequence consisting of encoded text from T5 [18], CLIP text embedding, an embedding representing diffusion timestep, an embedding representing the index of corresponding Gaussian, a noised CLIP image embedding and a final embedding which will be used to predict the target CLIP image embedding.”) –CLIP models add noise terms to CLIP images to generate vector of noised CLIP images. These vectors are used to fed to the decoder in the algorithm (1), which contains PNG media_image12.png 31 26 media_image12.png Greyscale .) Iteratively predicting, by the diffusion model, a first predicted noise term from the noised sample. (See page 5, column 2, paragraph 1. “CLIP text embedding, an embedding representing diffusion timestep, an embedding representing the index of corresponding Gaussian, a noised CLIP image embedding and a final embedding which will be used to predict the target CLIP image embedding.” – predicted image is obtained by using vector of noised images generated by the shifted diffusion model implemented/optimized using algorithm (1).) Regarding claim 4, As discussed with regard to claim 3, Yufan teaches all the limitations. Yufan further teaches the first loss is computed based on a difference between the first predicted noise term and the added noise term. (See page 4, column 1, equation (4). PNG media_image13.png 98 541 media_image13.png Greyscale - the loss function is computed based on PNG media_image14.png 21 21 media_image14.png Greyscale , PNG media_image15.png 20 49 media_image15.png Greyscale and PNG media_image16.png 18 22 media_image16.png Greyscale , which are the target image, first sample and initial image. The added noise term PNG media_image17.png 28 91 media_image17.png Greyscale and PNG media_image15.png 20 49 media_image15.png Greyscale are used in computing the loss.) Regarding claim 5, As discussed with regard to claim 2, Yufan teaches all the limitations. Yufan further teaches retraining the neural network model comprises: Adding a noise term to a training data sample from the second subset to generate a noised sample; (See page 5, column 1, paragraph 2 Decoder “Specifically, three diffusion models are trained on 64, 256, and 1024 resolutions, respectively. All the 900M images are used to train the decoder. During training, each image is processed by three different pre-trained CLIP models: ViT-B/16, ViT-B/32, and RN-101. The outputs from these three models will be concatenated into a single 1536-dimensional embedding, which is then projected into eight vectors and fed into the decoder.”. See page 5, column 2, paragraph 1. “Our shifted diffusion model is a decoder only transformer whose input is a sequence consisting of encoded text from T5 [18], CLIP text embedding, an embedding representing diffusion timestep, an embedding representing the index of corresponding Gaussian, a noised CLIP image embedding and a final embedding which will be used to predict the target CLIP image embedding.”– the CLIP models add noise terms to CLIP images to generate vector of noised CLIP images, are used to fed to the decoder in the algorithm (1), which contains PNG media_image18.png 33 30 media_image18.png Greyscale . ) Iteratively predicting by the diffusion model, a second predicted noise term from the noised sample; and (See page 5, column 2, paragraph 1. “CLIP text embedding, an embedding representing diffusion timestep, an embedding representing the index of corresponding Gaussian, a noised CLIP image embedding and a final embedding which will be used to predict the target CLIP image embedding.” – predicted image is obtained by using vector of noised images generated by the shifted diffusion model implemented/optimized using algorithm (1).) Iteratively predicting by the diffusion model, a third predicted noise term from the noised sample conditioned on a text label associated with the training data sample. (See page 4, column 2, equation (7) PNG media_image19.png 114 570 media_image19.png Greyscale . See page 5, column 2, paragraph 1, “We train two variants on datasets with different scales: one prior is trained on the full 900M image-text pairs; the other one is trained on CC15M‡, which is a subset of our full dataset.” See page 4, column 1 “Let z0 and y be the ground truth image embedding and its associated text caption, respectively. For each pair of (z0, y), we select its corresponding Gaussian PNG media_image20.png 21 25 media_image20.png Greyscale by the top-1 cosine similarity as PNG media_image21.png 46 365 media_image21.png Greyscale ”, and equation (7) – The shift diffusion, which is a decoder only transformer uses pair of image-text label PNG media_image20.png 21 25 media_image20.png Greyscale to predict the sample, and the images are noised CLIP images.) Regarding claim 6, As discussed with regard to claim 5, Yufan teaches all the limitations. Yufan further teaches the second loss is computed based on a difference between the second predicted noise term and the added noise term; and (See page 4, column 2, equation (7) – the loss function is computed based on PNG media_image22.png 21 22 media_image22.png Greyscale and PNG media_image23.png 29 112 media_image23.png Greyscale , which are the predicted term and the added noise term) The third loss is computed based on a difference between the third predicted noise term and the added noise term. (See page 4, column 1, equation 4 PNG media_image24.png 102 537 media_image24.png Greyscale . The loss is computed using PNG media_image25.png 24 23 media_image25.png Greyscale and PNG media_image26.png 24 21 media_image26.png Greyscale , which is the ground truth and the previous predicted image, and the third loss is computed after minimizing the first loss using gradient descent. In some cases, the loss is computed based on the ground truth and the previous predicted image.) Regarding claim 7, As discussed with regard to claim 6, Yufan teaches all the limitations. Yufan further teaches retraining the neural network model comprise: Updating at a training iteration, parameters of the diffusion model such that the third loss conditioned on the parameters of the diffusion model is no greater than the first loss conditioned on the parameters of the diffusion model. (See page 4, column 2, algorithm 1. – iteratively train the model while minimizing PNG media_image27.png 23 19 media_image27.png Greyscale /parameter using gradient descent so the parameter of the third loss is no greater than the first loss. Regarding claim 8, As discussed with regard to claim 1, Yufan teaches all the limitations. Yufan further teaches the non-textual data comprises any of: Biological structure data; time-series data; video-motion data. (See page 2, Figure. 1 “We propose Corgi, a novel diffusion model designed for flexible text-to-image generation which can “bridge the gap”. – broadest reasonable interpretation of biological structure data includes the corgi images.) With regard to claims 9 and 16, These claims are similar in scope to claim 1 and are rejected under a similar rationale. With regard to claims 10 and 17, These claims are similar in scope to claim 2 and are rejected under a similar rationale. With regard to claims 11 and 18, These claims are similar in scope to claim 3 and are rejected under a similar rationale. With regard to claims 12 and 19, These claims are similar in scope to claim 4 and are rejected under a similar rationale. With regard to claims 13 and 20, These claims are similar in scope to claim 5 and are rejected under a similar rationale. With regard to claim 14, The claim is similar in scope to claim 6 and is rejected under a similar rationale. With regard to claim 15, The claim is similar in scope to claim 7 and is rejected under a similar rationale Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure: US 12675917 B2 Identity-preserving image generation using diffusion models Kansy; Manuel Jakob et al. US 12651399 B2 Training data sampling for neural networks Lorraine; Jonathan Peter et al. US 12555275 B2 Personalized text-to-image diffusion model Aberman; Kfir et al. US 12462348 B2 Multimodal diffusion models Ham; Cusuh et al. US 20230095092 A1 DENOISING DIFFUSION GENERATIVE ADVERSARIAL NETWORKS Xiao; Zhisheng et al. Li, Yan, et al. "Generative time series forecasting with diffusion, denoise, and disentanglement." Advances in Neural Information Processing Systems 35 (2022): 23009-23022. Any inquiry concerning this communication or earlier communications from the examiner should be directed to MICHAEL HOANG LE whose telephone number is (571)270-7292. The examiner can normally be reached Monday-Friday 8:00 am - 5 pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Matthew Ell can be reached at 5712703264. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /MICHAEL HOANG LE/Examiner, Art Unit 2141 /MATTHEW ELL/Supervisory Patent Examiner, Art Unit 2141
Read full office action

Prosecution Timeline

Jan 25, 2024
Application Filed
Aug 26, 2026
Non-Final Rejection mailed — §102 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month