Prosecution Insights
Last updated: October 04, 2026
Application No. 19/043,064

CONTENT SYNTHESIS USING LATENT ADVERSARIAL DIFFUSION DISTILLATION

Non-Final OA §103
Filed
Jan 31, 2025
Priority
Mar 19, 2024 — provisional 63/567,137
Examiner
LE, MICHAEL
Art Unit
Tech Center
Assignee
Stability AI Ltd.
OA Round
1 (Non-Final)
66%
Grant Probability
Favorable
1-2
OA Rounds
1y 7m
Est. Remaining
87%
With Interview

Examiner Intelligence

Grants 66% — above average
66%
Career Allowance Rate
594 granted / 903 resolved
+5.8% vs TC avg
Strong +22% interview lift
Without
With
+21.6%
Interview Lift
resolved cases with interview
Typical timeline
3y 3m
Avg Prosecution
38 currently pending
Career history
952
Total Applications
across all art units

Statute-Specific Performance

§101
11.8%
-28.2% vs TC avg
§103
54.7%
+14.7% vs TC avg
§102
13.9%
-26.1% vs TC avg
§112
15.1%
-24.9% vs TC avg
Black line = Tech Center average estimate • Based on career data from 903 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status 1. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. Information Disclosure Statement 2. The information disclosure statements (IDS) submitted on the following dates are in compliance with the provisions of 37 CFR 1.97 and are being considered by the Examiner: 01/31/2025. Claim Rejections - 35 USC § 103 3. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention. 4. Claims 1-3, 7-9, 12, 14 and 16-17 are rejected under 35 U.S.C. 103 as being unpatentable over Nagano et al., (“Nagano”) [US-2023/0298243-A1] in view of Li et al. (“Li”) [US-2024/0161369-A1] Regarding claim 1, Nagano discloses a system (Nagano- ¶0019, at least discloses a system is provided for generating a digital avatar in accordance with a first machine learning (ML) model and a second machine learning (ML) model; Fig. 5B and ¶0143, at least disclose an exemplary system 565 in which the various architecture and/or functionality) comprising: one or more storage media storing instructions (Nagano- ¶0019, at least discloses The system comprises: a memory storing a two-dimensional (2D) input image; Fig. 5B and ¶0145-0146, at least disclose The system 565 also includes a main memory 540. Control logic (software) and data are stored in the main memory 540 which may take the form of a variety of computer-readable media. The computer-readable media may be any available media that may be accessed by the system 565 […] the main memory 540 may store computer-readable instructions (e.g., that represent a program(s) and/or a program element(s), such as an operating system); and one or more processors configured to execute the instructions (Nagano- Fig. 5B and ¶0143, at least disclose the CPU(s) 530 may be directly connected to the main memory 540. Further, the CPU(s) 530 may be directly connected to the parallel processing system 525; ¶0148, at least discloses Computer programs, when executed, enable the system 565 to perform various functions. The CPU(s) 530 may be configured to execute at least some of the computer-readable instructions to control one or more components of the system 565 to perform one or more of the methods) to cause the system to: receive a first representation of an image in a first latent space of a first machine learning model (Nagano- ¶0045, at least discloses The first ML model is referred to as EG3D: an Efficient Geometry-aware 3D Generative Adversarial Network. EG3D is trained to process a single 2D image to synthesize high-resolution multi-view consistent images in real-time along with corresponding high-quality 3D geometry; Fig. 1 and ¶0055, at least disclose In a first step, a GAN inversion optimization 104 is performed for the first ML model 110. The GAN inversion optimization 104 is an iterative process that determines a latent code [first latent space] that represents the input to the first ML model 110 [first machine learning model] that will produce an output that most closely matches the input image 102 [first representation of an image]; Fig. 2A and ¶0057, at least disclose As shown in FIG. 2A, the first ML model 110 processes a latent code w1 202 via a generator 204 network to produce the 3D data 112 [first representation of an image] and color data 114. The generator 204 network represents the layers of the first ML model 110, and the latent code w1 202 represents an input to the first ML model 110. The latent code w1 202 is a vector or tensor suitable as an input to the first ML model 110; Fig. 3 and ¶0079, at least disclose At 324, a GAN inversion optimization algorithm is performed for the first ML model in order to determine a latent code for the first ML model. The GAN inversion optimization algorithm optimizes, over a number of iterations, the latent code to cause the first ML model, when processing the optimized latent code, to generate unstructured 3D data and color data that best represents the object or person in the input image(s). The end result of the GAN inversion optimization algorithm is an optimal latent code within the latent space of the first ML model that can be used to generate a sample of the output data of the first ML model.); generate, by a second machine learning model based at least in part on the first representation (Nagano- ¶0012, at least discloses The GAN inversion optimization algorithm further comprises, for each iteration in a number of iterations: computing a candidate latent code by adding noise to the current latent code for the second ML model, applying the candidate latent code to the second ML model to generate 3D data and color data corresponding to the candidate latent code for the second ML model, and computing a loss function based on a comparison of at least one of the 3D data and color data corresponding to the candidate latent code for the second ML model with at least one of the 3D data and color data corresponding to the latent code for the first ML model), a second representation of the image in a second latent space of the second machine learning model (Nagano- ¶0046, at least discloses The latent vector for a particular style can also be generated by analyzing a second input image to extract the style representation from the second input image. StyleGAN2 can be modified to output 3D geometry data rather than a 2D image, thereby generating a mesh or surfaces that represent an object that can be rendered to generate the synthesized image; Fig. 1 and ¶0055, at least disclose In a second step, a GAN inversion optimization 106 is performed for the second ML model 120. The GAN inversion optimization 106 is an iterative process that determines a latent code that represents the input to the second ML model 120 that will produce an output that most closely matches the result of the GAN inversion optimization 104 [generate a second representation of the image]; Fig. 2B and ¶0063-0064, at least disclose As shown in FIG. 2B, the second ML model 120 processes a latent code w2 206 via a generator 208 network to produce the 3D data 122 and color data 124 […] The generator 208 network represents the layers of the second ML model 120, and the latent code w2 206 represents an input to the second ML model 120 […] another GAN inversion optimization is performed to determine a latent code for the second ML model 120. The GAN inversion optimization 106 of the second ML model 120 is a technique for finding the latent code w2 206 that generates an output of the second ML model 120 that most closely matches the optimized output of the first ML model 110; Fig. 2C and ¶0068-0069, at least disclose a different approach for generating a 3D digital avatar from an input image […] The encoder 214 network is trained to infer the latent code w2 206 for the generator 208 of the second ML model 120 directly based on the input image 102 […] The output of the first ML model 110 can then be used in the GAN inversion optimization 106 for the second ML model 120, to generate high-quality digital avatars based on a randomized input for the first ML model 110; Fig. 3 and ¶0080-0082, at least disclose At 326, a GAN inversion optimization algorithm is performed for the second ML model in order to determine a latent code for the second ML model. The GAN inversion optimization algorithm optimizes, over a number of iterations, the latent code to cause the second ML model, when processing the optimized latent code, to generate 3D data and color data that best matches the 3D data and color data generated by the first ML model using the latent code determined during step 304. At 328, the optimized latent code from step 306 is processed by the second ML model to generate 3D data and color data for the digital avatar. At 330, the 3D data and color data is processed by a renderer to generate an image of the digital avatar. The renderer may use any technique known in the art to produce an image of the digital avatar from the mesh and texture map(s) included in the 3D data and color data); and adjust, without generating an output image corresponding to the image, a set of weights of the machine learning model (Nagano- ¶0045-0046, at least disclose The first ML model is referred to as EG3D: an Efficient Geometry-aware 3D Generative Adversarial Network. EG3D is trained to process a single 2D image to synthesize high-resolution multi-view consistent images in real-time along with corresponding high-quality 3D geometry […] The latent vector for a particular style can also be generated by analyzing a second input image to extract the style representation from the second input image. StyleGAN2 can be modified to output 3D geometry data rather than a 2D image, thereby generating a mesh or surfaces that represent an object that can be rendered to generate the synthesized image [second representation]; ¶0166, at least discloses During training, data flows through the DNN in a forward propagation phase until a prediction is produced that indicates a label corresponding to the input. If the neural network does not correctly label the input, then errors between the correct label and the predicted label are analyzed, and the weights are adjusted for each feature during a backward propagation phase until the DNN correctly labels the input and other inputs in a training dataset; ). Nagano does not explicitly disclose update a set of weights of the second machine learning model based at least in part on the first representation and the second representation. However, Li discloses update a set of weights of the second machine learning model (Li- ¶0055, at least discloses the neural network based subject-driven image generation module 430 and one or more of its submodules 431-434 may be trained by iteratively updating the underlying parameters (e.g., weights 451, 452, etc., bias parameters and/or coefficients in the activation functions 461, 462 associated with neurons) of the neural network based on a loss function; ¶0059, at least discloses the training process transforms the neural network into an “updated” trained neural network with updated parameters such as weights, activation functions, and biases). It would have been obvious to one of ordinary in the art before the effective filing date of the claimed invention to have modified Nagano to incorporate the teachings of Li, and apply updating the underlying parameters such as weights into Nagano’s teachings in order to update, without generating an output image corresponding to the image, a set of weights of the second machine learning model based at least in part on the first representation and the second representation. The suggestion/motivation would have been in order to substitute one known element for another to obtain predictable results and to provide a subject-driven image generation model that generates accurate images portraying renditions of a given subject using one or more subject images. Regarding claim 2, Nagano in view of Li, discloses the system of claim 1, and further discloses wherein the execution of the instructions further causes the system to: generate a first noisy representation of the first representation by adding first noise to the first representation (Nagano- ¶0011, at least discloses The GAN inversion optimization algorithm further comprises, for each iteration in a number of iterations: computing a candidate latent code by adding noise to the current latent code, applying the candidate latent code to the first ML model to generate 3D data and color data corresponding to the candidate latent code, generating a candidate output image from the 3D data and color data corresponding to the candidate latent code, and computing a loss function based on a comparison of the candidate output image with the 2D input image; ¶0021, at least discloses The GAN inversion optimization algorithm further comprises, for each iteration in a number of iterations: computing a candidate latent code by adding noise to the current latent code, applying the candidate latent code to the first ML model to generate 3D data and color data corresponding to the candidate latent code, generating a candidate output image from the 3D data and color data corresponding to the candidate latent code, and computing a loss function based on a comparison of the candidate output image with the 2D input image; Li- ¶0079, at least discloses This process of incrementally adding noise to latent image representations effectively generates training data that is used in training the diffusion denoising model 612 […] Denoising model ε θ 612 may be a neural network based model, which has parameters that may be learned. Input to denoising model ε θ 612 may include a noisy latent representation (e.g., noised latent representation z T 606 t), and conditioning input 610 such as a text prompt describing desired content of an output image, e.g., “a hand holding a globe.”); and generate the second representation based at least in part on the first noisy representation (Nagano- ¶0012, at least discloses The GAN inversion optimization algorithm further comprises, for each iteration in a number of iterations: computing a candidate latent code by adding noise to the current latent code for the second ML model, applying the candidate latent code to the second ML model to generate 3D data and color data corresponding to the candidate latent code for the second ML model, and computing a loss function based on a comparison of at least one of the 3D data and color data corresponding to the candidate latent code for the second ML model with at least one of the 3D data and color data corresponding to the latent code for the first ML model; Li- ¶0079, at least discloses This process of incrementally adding noise to latent image representations effectively generates training data that is used in training the diffusion denoising model 612 […] Denoising model ε θ 612 may be a neural network based model, which has parameters that may be learned. Input to denoising model ε θ 612 may include a noisy latent representation (e.g., noised latent representation z T 606 t), and conditioning input 610 such as a text prompt describing desired content of an output image, e.g., “a hand holding a globe.”). It would have been obvious to one of ordinary in the art before the effective filing date of the claimed invention to have modified Nagano to incorporate the teachings of Li, and apply noisy latent representation into Nagano’s teachings in order to generate a first noisy representation of the first representation by adding first noise to the first representation; and generate the second representation based at least in part on the first noisy representation. Doing so would obtain predictable results and provide a subject-driven image generation model that generates accurate images portraying renditions of a given subject using one or more subject images. Regarding claim 3, Nagano in view of Li, discloses the system of claim 2, and further discloses wherein the execution of the instructions further causes the system to: generate a second noisy latent representation of the second representation by adding second noise to the first representation (Nagano- ¶0011, at least discloses The GAN inversion optimization algorithm further comprises, for each iteration in a number of iterations: computing a candidate latent code by adding noise to the current latent code, applying the candidate latent code to the first ML model to generate 3D data and color data corresponding to the candidate latent code, generating a candidate output image from the 3D data and color data corresponding to the candidate latent code, and computing a loss function based on a comparison of the candidate output image with the 2D input image; ¶0021, at least discloses The GAN inversion optimization algorithm further comprises, for each iteration in a number of iterations: computing a candidate latent code by adding noise to the current latent code, applying the candidate latent code to the first ML model to generate 3D data and color data corresponding to the candidate latent code, generating a candidate output image from the 3D data and color data corresponding to the candidate latent code, and computing a loss function based on a comparison of the candidate output image with the 2D input image; Li- ¶0079, at least discloses This process of incrementally adding noise to latent image representations effectively generates training data that is used in training the diffusion denoising model 612 […] Denoising model ε θ 612 may be a neural network based model, which has parameters that may be learned. Input to denoising model ε θ 612 may include a noisy latent representation (e.g., noised latent representation z T 606 t), and conditioning input 610 such as a text prompt describing desired content of an output image, e.g., “a hand holding a globe.”); and generate the second representation based at least in part on the second noisy representation (Nagano- ¶0012, at least discloses The GAN inversion optimization algorithm further comprises, for each iteration in a number of iterations: computing a candidate latent code by adding noise to the current latent code for the second ML model, applying the candidate latent code to the second ML model to generate 3D data and color data corresponding to the candidate latent code for the second ML model, and computing a loss function based on a comparison of at least one of the 3D data and color data corresponding to the candidate latent code for the second ML model with at least one of the 3D data and color data corresponding to the latent code for the first ML model; Li- ¶0079, at least discloses This process of incrementally adding noise to latent image representations effectively generates training data that is used in training the diffusion denoising model 612 […] Denoising model ε θ 612 may be a neural network based model, which has parameters that may be learned. Input to denoising model ε θ 612 may include a noisy latent representation (e.g., noised latent representation z T 606 t), and conditioning input 610 such as a text prompt describing desired content of an output image, e.g., “a hand holding a globe.”). Regarding claim 7, Nagano in view of Li, discloses the system of claim 1, and further discloses wherein execution of the instructions for updating the set of weights causes the system to: adjusting a weight included in the set of weights based at least in part on an adjustment value included in a weight adjustment signal received from a loss comparison system that was generated based at least in part on the second representation (Li- ¶0034, at least discloses The output image may be compared to the ground-truth subject image (e.g., modified image 204) by loss computation 206. The loss computed by loss computation 206 may be used to update parameters of subject-driven image model 130 via backpropagation 208; ¶0055, at least discloses the neural network based subject-driven image generation module 430 and one or more of its submodules 431-434 may be trained by iteratively updating the underlying parameters (e.g., weights 451, 452, etc., bias parameters and/or coefficients in the activation functions 461, 462 associated with neurons) of the neural network based on a loss function […] The output generated by the output layer 443 is compared to the expected output (e.g., a “ground-truth” such as the corresponding subject image with a replace background) from the training data, to compute a loss function that measures the discrepancy between the predicted output and the expected output. For example, the loss function may be a cross entropy loss. Given the loss, the negative gradient of the loss function is computed with respect to each weight of each layer individually; ¶0059, at least discloses the training process transforms the neural network into an “updated” trained neural network with updated parameters such as weights, activation functions, and biases). Regarding claim 8, Nagano in view of Li, discloses the system of claim 1, and further discloses wherein execution of the instructions for generating the second representation of the image causes the system (see Claim 1 rejection for detailed analysis) to: receive a prompt describing a desired characteristic of the image (Li- ¶0019, at least discloses Information about the subject may be provided to the base image generation model by generating an input prompt which includes a subject representation based on one or more input subject images; ¶0024, at least discloses The subject representation is combined with a text prompt and provided to a generic image generation model which generates an image of the subject based on the text prompt; ¶0026, at least discloses Subject-driven image model 130 takes an input subject image 102, a subject text 112, and a text prompt 118, and based on those inputs generates an output image 124); generate, using an encoding model, a prompt encoding based on the prompt (Li- ¶0024, at least discloses At inference, given a subject image and a text description of the subject, the multimodal encoder generates a multimodal subject representation. The subject representation is combined with a text prompt and provided to a generic image generation model which generates an image of the subject based on the text prompt; ¶0028, at least discloses Subject embedding 116 and text prompt 118 maybe combined, and input to text encoder 120 to generate the prompt for image model 122. Image model 122 may then generate an output image 124 based on the prompt. Subject text 112 may also be combined with subject embedding 116 and text prompt 118. In some embodiments, subject embedding 116, text prompt 118, and subject text 112 may be combined by the use of a prompt template. The prompt template may be, for example, “[text prompt], the [subject text] is [subject embedding]”. For example, if the text prompt is “a backpack at the grand canyon” and the subject text is “backpack”, then the combined prompt would be “a backpack at the grand canyon, the backpack is” concatenated with the subject embedding 116); and generate, using at least one transformer block of the second machine learning model, the second representation based at least in part on the first representation and the prompt encoding (Li- ¶0027, at least discloses Input subject image 102 may be encoded by an image encoder 104 into an image feature vector. Image encoder 104 may be a pretrained image encoder which extracts generic image features. Subject text 112 may be encoded by text encoder 106 into a text feature vector. The image feature vector and text feature vector may be input to multimodal encoder 108. Multimodal encoder 108 may be a query transformer (“Q-Former”) […] Multimodal encoder 108 may also take queries 110 as an input. Queries 110 may be randomly initialized vectors which may be tuned as part of the training process. Multimodal encoder 108 generates a vector representation of the subject (e.g., subject embedding) by using the subject text 112 to attend to the most relevant portions (i.e., the subject) of input subject image 102). It would have been obvious to one of ordinary in the art before the effective filing date of the claimed invention to have modified Nagano to incorporate the teachings of Li, and apply the query transformer into Nagano’s teachings in order to generate a first noisy representation of the first representation by adding first noise to the first representation; and generate the second representation based at least in part on the first noisy representation. The same motivation that was utilized in the rejection of claim 1 applies equally to this claim. Regarding claim 9, Nagano in view of Li, discloses the system of claim 1, and further discloses wherein the first machine learning model generates the first representation (see Claim 1 rejection for detailed analysis) using a first number of steps and the second machine learning model generates the second representation (see Claim 1 rejection for detailed analysis) using a second number of steps which is less than the first number of steps (Nagano- Figs. 1, 2A-2B and ¶0057-0059, at least disclose the two-step process for generating a 3D digital avatar from an input image […] As shown in FIG. 2A, the first ML model 110 processes a latent code w1 202 via a generator 204 network to produce the 3D data 112 and color data 114. The generator 204 network represents the layers of the first ML model 110, and the latent code w1 202 represents an input to the first ML model 110. The latent code w1 202 is a vector or tensor suitable as an input to the first ML model 110. In an embodiment, the latent code w1 202 is a 512 element vector, where each element may take the form of a scalar value within a set range (e.g., 0 to 1). By varying the values of each of the 512 elements in the latent code w1 202, the output produced by the generator 204 will also vary. In the first step of the two-step approach described above, the objective is to discover the latent code w1 202 that produces an output that most closely resembles the input image 102; ¶0063-0065, at least disclose As shown in FIG. 2B, the second ML model 120 processes a latent code w2 206 via a generator 208 network to produce the 3D data 122 and color data 124 [Wingdings font/0xE0] suggests a second number of steps which is less than the first number of steps). The method of claim 12 is similar in scope to the functions performed by the system of claim 1 and therefore claim 1 is rejected under the same rationale. Regarding claim 14, Nagano in view of Li, discloses the computer-implemented method of claim 12, and further discloses wherein the first machine learning model includes frozen weight (Li- ¶0058-0059, at least disclose all or a portion of parameters of one or more neural-network model being used together may be frozen, such that the “frozen” parameters are not updated during that training phase […] the training process transforms the neural network into an “updated” trained neural network with updated parameters such as weights, activation functions, and biases. The trained neural network thus improves neural network technology in image generation). It would have been obvious to one of ordinary in the art before the effective filing date of the claimed invention to have modified Nagano to incorporate the teachings of Li, and apply frozen weight into Nagano’s teachings in order the first machine learning model includes frozen weight. Doing so would obtain predictable results and provide a subject-driven image generation model that generates accurate images portraying renditions of a given subject using one or more subject images. Regarding claim 16, Nagano in view of Li, discloses the computer-implemented method of claim 12, and discloses the method further comprising: receiving a prompt describing a desired characteristic of a second image to generate (Li- ¶0019, at least discloses Information about the subject may be provided to the base image generation model by generating an input prompt which includes a subject representation based on one or more input subject images; ¶0024, at least discloses The subject representation is combined with a text prompt and provided to a generic image generation model which generates an image of the subject based on the text prompt; ¶0026, at least discloses Subject-driven image model 130 takes an input subject image 102, a subject text 112, and a text prompt 118, and based on those inputs generates an output image 124); generating, using an encoding model, a prompt encoding based at least in part on the prompt (Li- ¶0024, at least discloses At inference, given a subject image and a text description of the subject, the multimodal encoder generates a multimodal subject representation. The subject representation is combined with a text prompt and provided to a generic image generation model which generates an image of the subject based on the text prompt; ¶0028, at least discloses Subject embedding 116 and text prompt 118 maybe combined, and input to text encoder 120 to generate the prompt for image model 122. Image model 122 may then generate an output image 124 based on the prompt. Subject text 112 may also be combined with subject embedding 116 and text prompt 118. In some embodiments, subject embedding 116, text prompt 118, and subject text 112 may be combined by the use of a prompt template. The prompt template may be, for example, “[text prompt], the [subject text] is [subject embedding]”. For example, if the text prompt is “a backpack at the grand canyon” and the subject text is “backpack”, then the combined prompt would be “a backpack at the grand canyon, the backpack is” concatenated with the subject embedding 116); generating, using at least one transformer block of the second machine learning model, a third representation based at least in part on the prompt encoding (Li- ¶0027, at least discloses Input subject image 102 may be encoded by an image encoder 104 into an image feature vector. Image encoder 104 may be a pretrained image encoder which extracts generic image features. Subject text 112 may be encoded by text encoder 106 into a text feature vector. The image feature vector and text feature vector may be input to multimodal encoder 108. Multimodal encoder 108 may be a query transformer (“Q-Former”) […] Multimodal encoder 108 may also take queries 110 as an input. Queries 110 may be randomly initialized vectors which may be tuned as part of the training process. Multimodal encoder 108 generates a vector representation of the subject (e.g., subject embedding) by using the subject text 112 to attend to the most relevant portions (i.e., the subject) of input subject image 102); and generating, using a decoding model, the second image based at least in part on the third representation (Nagano- ¶0197, at least discloses The client device 604 may receive the encoded display data via the communication interface 621 and the decoder 622 may decode the encoded display data to generate the display data. The client device 604 may then display the display data via the display 624; Li- ¶0080, at least discloses The latent image representation produced using denoising model ε θ 612 may be decoded using decoder 614 to provide an output 616 which is the denoised image). It would have been obvious to one of ordinary in the art before the effective filing date of the claimed invention to have modified Nagano to incorporate the teachings of Li, and apply decoding into Nagano’s teachings for generating, using at least one transformer block of the second machine learning model, a third representation based at least in part on the prompt encoding; and generating, using a decoding model, the second image based at least in part on the third representation. Doing so would obtain predictable results and provide a subject-driven image generation model that generates accurate images portraying renditions of a given subject using one or more subject images. Regarding claim 17, Nagano in view of Li, discloses one or more non-transitory computer-readable storage media storing instructions that, upon execution executable by one or more processors of a system (Nagano- ¶0019, at least discloses The system comprises: a memory storing a two-dimensional (2D) input image; Fig. 5B and ¶0143, ¶0145-0146, at least disclose the CPU(s) 530 may be directly connected to the main memory 540. Further, the CPU(s) 530 may be directly connected to the parallel processing system 525 […] The system 565 also includes a main memory 540. Control logic (software) and data are stored in the main memory 540 which may take the form of a variety of computer-readable media. The computer-readable media may be any available media that may be accessed by the system 565 […] the main memory 540 may store computer-readable instructions (e.g., that represent a program(s) and/or a program element(s), such as an operating system; ¶0148, at least discloses Computer programs, when executed, enable the system 565 to perform various functions. The CPU(s) 530 may be configured to execute at least some of the computer-readable instructions to control one or more components of the system 565 to perform one or more of the methods), cause the system to perform operations comprising the functions of claim 1. 5. Claims 4-5 are rejected under 35 U.S.C. 103 as being unpatentable over Nagano in view of Li, further in view over “HIVE: Harnessing Human Feedback for Instructional Visual Editing” by Zhang et al. (“Zhang”) Regarding claim 4, Nagano in view of Li, discloses the system of claim 3, and does not explicitly disclose, but Zhang discloses wherein the first noise and the second noise are the same noise (Zhang- page 15, 4th paragraph, at least discloses As a result, the reverse diffusion process can be viewed as a black box function defined by Eg and noises E := (zr, ... ,z1,xT), which we can view as a shared parameter network with noises). It would have been obvious to one of ordinary in the art before the effective filing date of the claimed invention to have modified Nagano/Li to incorporate the teachings of Zhang, and apply sharing parameter network with noises into Nagano/Li’s teachings in order the first noise and the second noise are the same noise. Doing so would introduce scalable diffusion model fine-tuning methods that can incorporate human preferences based on the estimated reward. Regarding claim 5, Nagano in view of Li and Zhang, discloses the system of claim 4, and further discloses wherein the first noise is selected from a Gaussian distribution (Zhang- page 14, last paragraph, at least discloses I is a Gaussian distribution, whose parameters are defined by score function Eg and stepsize of noise scalings.). 6. Claim 6 is rejected under 35 U.S.C. 103 as being unpatentable over Nagano in view of Li, further in view of Martinez et al. (“Martinez”) [US-11,568,576-B1] Regarding claim 6, Nagano in view of Li, discloses the system of claim 1, and does not explicitly disclose, but Martinez discloses wherein the set of weights is updated based at least on using at least one of an adversarial loss comparison or a distillation loss comparison (Martinez- col 4, lines 34-46, at least discloses The parameters (e.g., weights and/or biases) of the machine learning model may be updated to minimize (or maximize) the cost. For example, the machine learning model may use a gradient descent (or ascent) algorithm to incrementally adjust the weights to cause the most rapid decrease (or increase) to the output of the loss function; col 13, lines 33-43, at least discloses the adversarial loss from the per-class discriminators 316 may be fed back to generator 304 and may be used to update parameters (e.g., weights) of the generator 304 to improve the quality of the different objects drawn by the generator 304. Additionally, the loss may be used to update parameters of discriminators 316). It would have been obvious to one of ordinary in the art before the effective filing date of the claimed invention to have modified Nagano/Li to incorporate the teachings of Martinez, and apply the adversarial loss into Nagano/Li’s teachings in order the set of weights is updated based at least on using at least one of an adversarial loss comparison or a distillation loss comparison. Doing so would generate millions of synthetic images and/or videos of any desired scene. 7. Claim 13 is rejected under 35 U.S.C. 103 as being unpatentable over Nagano in view of Li, further in view over “Scalable Diffusion Models with Transformers” by Peebles et al. (“Peebles”) Regarding claim 13, Nagano in view of Li, discloses the computer-implemented method of claim 12, and does not explicitly disclose, but Peebles discloses wherein the first machine learning model is a first diffusion transformer model and the second machine learning model is a second diffusion transformer model (Peebles- Fig. 3 and page 4175, section 3.2. Diffusion Transformer Design Space, at least discloses We introduce Diffusion Transformers (DiTs), a new architecture for diffusion models. We aim to be as faithful to the standard transformer architecture as possible to retain its scaling properties. Since our focus is training DDPMs of images (specifically, spatial representations of images), DiT is based on the Vision Transformer (ViT) architecture which operates on sequences of patches; page 4175, section DiT block design, right column, at least discloses Following patchify, the input tokens are processed by a sequence of transformer blocks. In addition to noised image inputs, diffusion models sometimes process additional conditional information such as noise timesteps t, class labels c, natural language, etc. We explore four variants of transformer blocks that process conditional inputs differently. The designs introduce small, but important, modifications to the standard ViT block design. The designs of all blocks are shown in Figure 3). It would have been obvious to one of ordinary in the art before the effective filing date of the claimed invention to have modified Nagano/Li to incorporate the teachings of Peebles, and apply the Diffusion Transformers into Nagano/Li’s teachings in order the first machine learning model is a first diffusion transformer model and the second machine learning model is a second diffusion transformer model. Doing so would demystify the significance of architectural choices in diffusion models and offer empirical baselines for future generative modeling research. 8. Claims 10-11, 15 and 19-20 are rejected under 35 U.S.C. 103 as being unpatentable over Nagano in view of Li, further in view over “Adversarial Diffusion Distillation” by Sauer et al. (“Sauer”) Regarding claim 10, Nagano in view of Li, discloses the system of claim 1, and further discloses wherein execution of the instructions for updating the set of weights (see Claim 1 rejection for detailed analysis) causes the system to: generate a second noisy representation by adding noise to the second representation (Nagano- ¶0012, at least discloses The GAN inversion optimization algorithm further comprises, for each iteration in a number of iterations: computing a candidate latent code by adding noise to the current latent code for the second ML model, applying the candidate latent code to the second ML model to generate 3D data and color data corresponding to the candidate latent code for the second ML model, and computing a loss function based on a comparison of at least one of the 3D data and color data corresponding to the candidate latent code for the second ML model with at least one of the 3D data and color data corresponding to the latent code for the first ML model; ¶0021, at least discloses The GAN inversion optimization algorithm further comprises, for each iteration in a number of iterations: computing a candidate latent code by adding noise to the current latent code, applying the candidate latent code to the first ML model to generate 3D data and color data corresponding to the candidate latent code, generating a candidate output image from the 3D data and color data corresponding to the candidate latent code, and computing a loss function based on a comparison of the candidate output image with the 2D input image; Li- ¶0079, at least discloses This process of incrementally adding noise to latent image representations effectively generates training data that is used in training the diffusion denoising model 612 […] Denoising model ε θ 612 may be a neural network based model, which has parameters that may be learned. Input to denoising model ε θ 612 may include a noisy latent representation (e.g., noised latent representation z T 606 t), and conditioning input 610 such as a text prompt describing desired content of an output image, e.g., “a hand holding a globe.”); generate a third representation by inputting the second noisy representation to the first machine learning model (Nagano- ¶0012, at least discloses The GAN inversion optimization algorithm further comprises, for each iteration in a number of iterations: computing a candidate latent code by adding noise to the current latent code for the second ML model, applying the candidate latent code to the second ML model to generate 3D data and color data corresponding to the candidate latent code for the second ML model, and computing a loss function based on a comparison of at least one of the 3D data and color data corresponding to the candidate latent code for the second ML model with at least one of the 3D data and color data corresponding to the latent code for the first ML model; Li- ¶0079, at least discloses This process of incrementally adding noise to latent image representations effectively generates training data that is used in training the diffusion denoising model 612 […] Denoising model ε θ 612 may be a neural network based model, which has parameters that may be learned. Input to denoising model ε θ 612 may include a noisy latent representation (e.g., noised latent representation z T 606 t), and conditioning input 610 such as a text prompt describing desired content of an output image, e.g., “a hand holding a globe.”) . The prior does not explicitly disclose, but Sauer discloses determine a distillation loss using at least the representation (Sauer- page 2, left column, 1st paragraph, at least discloses We propose Adversarial Diffusion Distillation (ADD), a general approach that reduces the number of inference steps of a pre-trained diffusion model to 1-4 sampling steps while maintaining high sampling fidelity and potentially further improving the overall performance of the model. To this end, we introduce a combination of two training objectives: (i) an adversarial loss and (ii) a distillation loss that corresponds to score distillation sampling (SDS); page 4, left column, 1st paragraph, at least discloses we utilize the gradient of a pretrained diffusion model via a score distillation objective to improve text alignment and sample quality). It would have been obvious to one of ordinary in the art before the effective filing date of the claimed invention to have modified Nagano/Li to incorporate the teachings of Sauer, and apply the distillation loss into Nagano/Li’s teachings in order to generate a second noisy representation by adding noise to the second representation; generate a third representation by inputting the second noisy representation to the first machine learning model; and determine a distillation loss using at least the third representation. Doing so would ensure high image fidelity. Regarding claim 11, Nagano in view of Li and Sauer, discloses the system of claim 10, and further discloses wherein execution of the instructions for updating the set of weights (see Claim 1 rejection for detailed analysis) further causes the system to: generate a fourth representation by inputting the second noisy representation to the first machine learning model (Nagano- ¶0012, at least discloses The GAN inversion optimization algorithm further comprises, for each iteration in a number of iterations: computing a candidate latent code by adding noise to the current latent code for the second ML model, applying the candidate latent code to the second ML model to generate 3D data and color data corresponding to the candidate latent code for the second ML model, and computing a loss function based on a comparison of at least one of the 3D data and color data corresponding to the candidate latent code for the second ML model with at least one of the 3D data and color data corresponding to the latent code for the first ML model; Li- ¶0079, at least discloses This process of incrementally adding noise to latent image representations effectively generates training data that is used in training the diffusion denoising model 612 […] Denoising model ε θ 612 may be a neural network based model, which has parameters that may be learned. Input to denoising model ε θ 612 may include a noisy latent representation (e.g., noised latent representation z T 606 t), and conditioning input 610 such as a text prompt describing desired content of an output image, e.g., “a hand holding a globe.”); and determine the distillation loss using at least the fourth representation (Sauer- page 2, left column, 1st paragraph, at least discloses We propose Adversarial Diffusion Distillation (ADD), a general approach that reduces the number of inference steps of a pre-trained diffusion model to 1-4 sampling steps while maintaining high sampling fidelity and potentially further improving the overall performance of the model. To this end, we introduce a combination of two training objectives: (i) an adversarial loss and (ii) a distillation loss that corresponds to score distillation sampling (SDS); page 4, left column, 1st paragraph, at least discloses we utilize the gradient of a pretrained diffusion model via a score distillation objective to improve text alignment and sample quality). Regarding claim 15, Nagano in view of Li, discloses the computer-implemented method of claim 12, and further discloses wherein the set of weights is updated (see Claim 1 rejection for detailed analysis), and does not explicitly disclose, but Sauer discloses based at least on using an adversarial loss comparison and a distillation loss comparison (Sauer- page 2, left column, 1st paragraph, at least discloses We propose Adversarial Diffusion Distillation (ADD), a general approach that reduces the number of inference steps of a pre-trained diffusion model to 1-4 sampling steps while maintaining high sampling fidelity and potentially further improving the overall performance of the model. To this end, we introduce a combination of two training objectives: (i) an adversarial loss and (ii) a distillation loss that corresponds to score distillation sampling (SDS); page 4, left column, 1st paragraph, at least discloses we utilize the gradient of a pretrained diffusion model via a score distillation objective to improve text alignment and sample quality. Furthermore, instead of training from scratch, we initialize our model with pretrained diffusion model weights; pretraining the generator network is known to significantly improve training with an adversarial loss; page 4, left column, section 3.1. Training Procedure, at least discloses The ADD-student is initialized from a pretrained UNet-DM with weights 0, a discriminator with trainable weights¢, and a DM teacher with frozen weights 'lj; page 6, right column, section Loss terms, at least discloses We find that both losses are essential. The distillation loss on its own is not effective, but when combined with the adversarial loss, there is a noticeable improvement in results. Different weighting schedules lead to different behaviours, the exponential schedule tends to yield more diverse samples, as indicated by lower FID, SDS and NFSD schedules improve quality and text alignment. While we use the exponential schedule as the default setting in all other ablations, we opt for the NFSD weighting for training our final model. Choosing an optimal weighting function presents an opportunity for improvement. Alternatively, scheduling the distillation weights over training, as explored in the 3D generative modeling literature [::3] could be considered). It would have been obvious to one of ordinary in the art before the effective filing date of the claimed invention to have modified Nagano/Li to incorporate the teachings of Sauer, and apply the combination of two training objectives: an adversarial loss and a distillation loss into Nagano/Li’s teachings in order the set of weights is updated based at least on using an adversarial loss comparison and a distillation loss comparison. Doing so would ensure high image fidelity. Regarding claim 19, Nagano in view of Li, discloses the non-transitory computer-readable storage medium of claim 17, and further discloses wherein instructions for updating the set of weights (see Claim 1 rejection for detailed analysis) cause the system to: generate a second noisy representation by adding noise to the second representation (see Claim 10 rejection for detailed analysis); generate a third representation by inputting the second noisy representation to the first machine learning model (see Claim 10 rejection for detailed analysis); and determine a distillation loss using at least the third representation (see Claim 10 rejection for detailed analysis). Regarding claim 20, Nagano in view of Li and Sauer, discloses the non-transitory computer-readable storage medium of claim 10, and further discloses wherein instructions for updating the set of weights (see Claim 1 rejection for detailed analysis) further cause the system to: generate a fourth representation by inputting the second noisy representation to the first machine learning model (Nagano- ¶0011, at least discloses The GAN inversion optimization algorithm further comprises, for each iteration in a number of iterations: computing a candidate latent code by adding noise to the current latent code, applying the candidate latent code to the first ML model to generate 3D data and color data corresponding to the candidate latent code, generating a candidate output image from the 3D data and color data corresponding to the candidate latent code, and computing a loss function based on a comparison of the candidate output image with the 2D input image; ¶0021, at least discloses The GAN inversion optimization algorithm further comprises, for each iteration in a number of iterations: computing a candidate latent code by adding noise to the current latent code, applying the candidate latent code to the first ML model to generate 3D data and color data corresponding to the candidate latent code, generating a candidate output image from the 3D data and color data corresponding to the candidate latent code, and computing a loss function based on a comparison of the candidate output image with the 2D input image); and determine the distillation loss using at least the fourth representation (Sauer- page 2, left column, 1st paragraph, at least discloses We propose Adversarial Diffusion Distillation (ADD), a general approach that reduces the number of inference steps of a pre-trained diffusion model to 1-4 sampling steps while maintaining high sampling fidelity and potentially further improving the overall performance of the model. To this end, we introduce a combination of two training objectives: (i) an adversarial loss and (ii) a distillation loss that corresponds to score distillation sampling (SDS); page 4, left column, 1st paragraph, at least discloses we utilize the gradient of a pretrained diffusion model via a score distillation objective to improve text alignment and sample quality). 9. Claim 18 is rejected under 35 U.S.C. 103 as being unpatentable over Nagano in view of Li, further in view over “TRACT: Denoising Diffusion Models with Transitive Closure Time-Distillation” by Berthelot et al. (“Berthelot”) Regarding claim 18, Nagano in view of Li, discloses the non-transitory computer-readable storage medium of claim 17, and further discloses wherein the first machine learning model generates the first representation using a first number of steps and the second machine learning model generates the second representation using a second number of steps which is less than the first number of steps (Nagano- Figs. 1, 2A-2B and ¶0057-0059, at least disclose the two-step process for generating a 3D digital avatar from an input image […] As shown in FIG. 2A, the first ML model 110 processes a latent code w1 202 via a generator 204 network to produce the 3D data 112 and color data 114. The generator 204 network represents the layers of the first ML model 110, and the latent code w1 202 represents an input to the first ML model 110. The latent code w1 202 is a vector or tensor suitable as an input to the first ML model 110. In an embodiment, the latent code w1 202 is a 512 element vector, where each element may take the form of a scalar value within a set range (e.g., 0 to 1). By varying the values of each of the 512 elements in the latent code w1 202, the output produced by the generator 204 will also vary. In the first step of the two-step approach described above, the objective is to discover the latent code w1 202 that produces an output that most closely resembles the input image 102; ¶0063-0065, at least disclose As shown in FIG. 2B, the second ML model 120 processes a latent code w2 206 via a generator 208 network to produce the 3D data 122 and color data 124 [Wingdings font/0xE0] suggests a second number of steps which is less than the first number of steps). The prior art does not explicitly disclose, but Berthelot discloses using a first number of sampling steps and the second machine learning model generates the second representation using a second number of sampling steps which is less than the first number of steps. However, Berthelot discloses using a first number of sampling steps and using a second number of sampling steps (Berthelot- Figs. 2, 3, 4 shows varying number of sampling steps; page 16, section A.6 Generated samples, at least discloses We present random samples from our distilled models with varying sampling steps. As shown in Figure 2, 3 and 4, the deterministic mapping between noises and samples are mostly preserved in these distilled models. We can see that there is a slight degrade of image quality for the I-step student compared with students distilled to more sampling steps). It would have been obvious to one of ordinary in the art before the effective filing date of the claimed invention to have modified Nagano/Li to incorporate the teachings of Berthelot, and apply the varying number of sampling steps into Nagano/Li’s teachings in order the first machine learning model generates the first representation using a first number of sampling steps and the second machine learning model generates the second representation using a second number of sampling steps which is less than the first number of steps. Doing so would provide proficiency for generative sampling. Conclusion 10. The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. They are as recited in the attached PTO-892 form. 11. Any inquiry concerning this communication or earlier communications from the examiner should be directed to MICHAEL LE whose telephone number is (571)272-5330. The examiner can normally be reached 9am-5pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kent Chang can be reached at (571) 272-7667. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /MICHAEL LE/Primary Examiner, Art Unit 2614
Read full office action

Prosecution Timeline

Jan 31, 2025
Application Filed
Sep 08, 2026
Non-Final Rejection mailed — §103
Sep 28, 2026
Interview Requested

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12739358
PROJECTION SYSTEM AND METHOD WITH BLENDED COLOR GAMUT
2y 0m to grant Granted Sep 15, 2026
Patent 12721389
SYSTEM AND METHOD FOR CONTROLLING SELECTIVE REVEALING OBJECT
2y 2m to grant Granted Sep 01, 2026
Patent 12718470
IMAGE BLENDING USING ONE OR MORE NEURAL NETWORKS
4y 8m to grant Granted Aug 25, 2026
Patent 12718430
METHODS AND SYSTEMS RELATING TO DIGITAL MARK OPACITY, BLENDING AND CANVAS TEXTURE
3y 2m to grant Granted Aug 25, 2026
Patent 12718466
SIMPLIFIED LOW-PRECISION RAY INTERSECTION THROUGH ACCELERATED HIERARCHY STRUCTURE PRECOMPUTATION
2y 11m to grant Granted Aug 25, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
66%
Grant Probability
87%
With Interview (+21.6%)
3y 3m (~1y 7m remaining)
Median Time to Grant
Low
PTA Risk
Based on 903 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month