Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Status of Claims
This communication is in response to the Application Filed on 11/14/2024.
Claims 1-20 are pending in this application.
Drawings
The drawing(s) filed on 11/14/2024 are accepted by the Examiner.
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 11/14/2024 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Positive Statement regarding 35 U.S.C. 101: Claims 1-20 are determined to be eligible under 35 U.S.C. 101. For example, claim 8, like claims 1 and 15, recites, “generating an intermediate output based on the condition input by performing a first diffusion process for a first number of timesteps based on the adherence parameter; and generating a synthetic image based on the intermediate output.” It is given the weight of the description in the specification paragraph [0023], which states the practical purpose of the invention is to allow an image generation model to generate a new image while keeping some characteristics of a reference image. The data gathered is important to the solution presented by the claimed invention. The generation of image data is not an insignificant post-solution step. It is a meaningful step that generates an intermediate representation for a generative model to process as guidance to another generative model, which accomplishes the overall goal in the disclosure. Because the claims are not well-understood, routine, conventional, or insignificant extra-solution data gathering, they are not directed to an abstract idea. Therefore, the claims are determined to be eligible under 35 U.S.C. 101.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or non-obviousness.
Claims 1-5, 7-12, and 14-19 are rejected under 35 U.S.C. 103 as being unpatentable over Yu et al. (US 2024/0386623 A1, hereinafter, “Yu”) in view of Zhang et al. (Adding Conditional Control to Text-to-Image Diffusion Models, 2023, hereinafter, “Zhang”).
Regarding claim 1, Yu teaches a method (See Yu, ¶ [0082], FIG. 9 is an example logic flow diagram illustrating a method of controllable image generation) comprising:
obtaining a condition input (See Yu, ¶ [0023], based on a user input (e.g., a text prompt in conditioning input 110)) and [an adherence parameter], wherein the condition input indicates an image attribute (See Yu, ¶ [0023], based on a user input (e.g., a text prompt in conditioning input 110). Examiner considers a text prompt an image attribute as it describes the image) and the [adherence parameter indicates a level of the image attribute];
generating, using a first image generation model (See Yu, ¶ [0086], a first neural network based image model (e.g., diffusion model 214)), an intermediate output (See Yu, ¶ [0086], a first latent representation) based on the condition input (See Yu, ¶ [0023], based on a user input (e.g., a text prompt in conditioning input 110)) and [the adherence parameter], wherein the intermediate output represents the image attribute (See Yu, ¶ [0086], At step 903, the system generates, by a first neural network based image model (e.g., diffusion model 214), a first latent representation based on the task-specific feature map); and
generating, using a second image generation model (See Yu, ¶ [0086], a second neural network based image model (e.g., diffusion model 212)), a synthetic image (See Yu, ¶ [0086], second latent representation) based on the intermediate output (See Yu, ¶ [0088], At step 905, the system modifies a second latent representation of a second neural network based image model (e.g., diffusion model 212) based on the first latent representation and the task embedding), wherein the synthetic image (See Yu, ¶ [0086], second latent representation) depicts the image attribute [at the level indicated by the adherence parameter].
However, Yu does not disclose an adherence parameter and [wherein the condition input indicates an image attribute] at the level indicated by the adherence parameter.
Zhang teaches an adherence parameter (See Zhang, Pg. 3817 Section: Classifier-free guidance resolution weighting, line 6, a user-specified weight) and [wherein the condition input indicates an image attribute] at the level indicated by the adherence parameter (See Zhang, Pg. 3817 Section: Classifier-free guidance resolution weighting, line 6, a user-specified weight. Examiner considers the user-specified weight to vary depending on what the user selects, and thus the level indicated can vary).
Thus, it would have been obvious to one of ordinary skills in the art before the effective filing date of the claimed invention to modify Yu’s reference to include an adherence parameter and [wherein the condition input indicates an image attribute] at the level indicated by the adherence parameter based on the method of Zhang’s reference. The suggestion/motivation would have been to personalize content in the generated image by finetuning the image diffusion model as suggested by Zhang at Pg. 3815, Section: Controlling Image Diffusion Models, lines 12-14.
Further, one skilled in the art could have combined the elements as described above by known method with no change in their respective functions, and the combination would have yielded nothing more than predictable results.
Therefore, it would have been obvious to combine Zhang with Yu to obtain the invention as specified in claim 1.
Regarding claim 2, in which claim 1 is incorporated, Yu discloses wherein obtaining the condition input (See Yu, ¶ [0023], based on a user input (e.g., a text prompt in conditioning input 110)) comprises:
obtaining a 3D model of an object (See Yu, ¶ [0037], a normal surface is an image which represents 3D shapes by using red/green/blue (RGB) values of the image to correspond to X, Y, and Z axis in 3D space); and
generating (See Yu, ¶ [0036], applies a convolutional kernel to the visual condition 204 to produce a feature map) a depth map based on the 3D model (See Yu, ¶ [0037], Another example of a visual condition and corresponding task instruction is a depth map… A depth map may be an image which indicates depth with luminance values. Examiner considers the depth map to include a z-axis, which is also a part of the 3D space of an image as noted above. The visual condition is used to generate a feature map, which can be a depth map that is generated), wherein the condition input comprises the depth map (See Yu, ¶ [0018], The input conditioning image may take on a variety of different forms such as a sketch, a relief map, etc. Examiner considers the depth map to be one of a variety of different forms an input conditioning image could take) and the image attribute comprises a shape of the object (See Yu, ¶ [0037], a normal surface is an image which represents 3D shapes).
Regarding claim 3, in which claim 1 is incorporated, Yu discloses wherein generating the intermediate output (See Yu, ¶ [0086], a first latent representation) comprises:
generating a plurality of layer-specific features based on the condition input (See Yu, Fig. 1, elements 102 and 106a to 106t; Fig. 2, elements 202 and 212; ¶ [0025], This process is repeated T times (e.g., 50 iterations) until it results in a noised latent representation zT 106t. Examiner considers zT as a plurality of layer-specific features); and
providing the plurality of layer-specific features to a plurality of corresponding layers of the first image generation model, respectively (See Yu, Fig. 1; ¶ [0027], As illustrated, denoising model εθ 112 is iteratively used to reverse the process of noising latents (i.e., perform reverse diffusion) from zT 118t to z0 118a. Denoising model εθ 112 may be a neural network based model, which has parameters that may be learned. Input to denoising model ε 112 may include a noisy latent representation (e.g., noised latent representation zT 106t)).
Regarding claim 4, in which claim 1 is incorporated, Yu discloses wherein generating the intermediate output (See Yu, ¶ [0086], a first neural network based image model (e.g., diffusion model 214)) comprises:
determining a timestep threshold (See Yu, ¶ [0025], This process is repeated T times (e.g., 50 iterations) until it results in a noised latent representation zT 106t. Examiner considers the 50 iterations to be a timestep threshold) [based on the adherence parameter]; and
performing, using the first image generation model (See Yu, ¶ [0086], a first neural network based image model (e.g., diffusion model 214)), a first diffusion process (See Yu, ¶ [0086], a first neural network based image model (e.g., diffusion model 214)) before the timestep threshold (See Yu, Fig. 1; ¶ [0025], This process is repeated T times (e.g., 50 iterations)). examiner considers before the timestep threshold to be from 0 to 50 iterations in the forward process).
However, Yu does not disclose an adherence parameter.
Zhang teaches an adherence parameter (See Zhang, Pg. 3817 Section: Classifier-free guidance resolution weighting, line 6, a user-specified weight).
Thus, it would have been obvious to one of ordinary skills in the art before the effective filing date of the claimed invention to modify Yu’s reference to include an adherence parameter based on the method of Zhang’s reference. The suggestion/motivation would have been to personalize content in the generated image by finetuning the image diffusion model as suggested by Zhang at Pg. 3815, Section: Controlling Image Diffusion Models, lines 12-14.
Further, one skilled in the art could have combined the elements as described above by known method with no change in their respective functions, and the combination would have yielded nothing more than predictable results.
Therefore, it would have been obvious to combine Zhang with Yu to obtain the invention as specified in claim 4.
Regarding claim 5, in which claim 4 is incorporated, Yu discloses wherein generating the synthetic image (See Yu, ¶ [0086], second latent representation) comprises:
performing, using the second image generation model (See Yu, ¶ [0086], a second neural network based image model (e.g., diffusion model 212)), a second diffusion process (See Yu, ¶ [0088], a second neural network based image model (e.g., diffusion model 212)) after the timestep threshold (See Yu, ¶ [0025], This process is repeated T times (e.g., 50 iterations)); ¶ [0027], reverse the process of noising latents (i.e., perform reverse diffusion) from zT 118t to z0 118a). Examiner considers after the timestep threshold to be once 50 iterations has been reached back down to 0 iterations in the reverse process).
Regarding claim 7, in which claim 1 is incorporated, Yu discloses obtaining a text prompt, wherein the synthetic image is generated based on the text prompt (See Yu, ¶ [0023], a denoising diffusion model is trained to generate an image (e.g., output 116) based on a user input (e.g., a text prompt in conditioning input 110)).
Regarding claim 8, Yu discloses a non-transitory computer readable medium storing code for image processing (See Yu, ¶ [0050], non-transitory, tangible, machine readable media that includes executable code), the code comprising instructions that, when executed by at least one processor (See Yu, ¶ [0047], processor 710), cause the at least one processor to perform operations (See Yu, ¶ [0047], Operation of computing device 700 is controlled by processor 710) comprising:
obtaining a condition input (See Yu, ¶ [0023], based on a user input (e.g., a text prompt in conditioning input 110)) and [an adherence parameter;]
generating an intermediate output (See Yu, ¶ [0086], a first latent representation) based on the condition input (See Yu, ¶ [0023], based on a user input (e.g., a text prompt in conditioning input 110)) by performing a first diffusion process (See Yu, ¶ [0086], a first neural network based image model (e.g., diffusion model 214)) for a first number of timesteps (See Yu, Fig. 1; ¶ [0025], This process is repeated T times (e.g., 50 iterations). Examiner considers 50 iterations to be a first number of timesteps) [based on the adherence parameter;] and
generating a synthetic image (See Yu, ¶ [0086], second latent representation) based on the intermediate output (See Yu, ¶ [0088], At step 905, the system modifies a second latent representation of a second neural network based image model (e.g., diffusion model 212) based on the first latent representation and the task embedding. Examiner considers the first latent representation to be the intermediate output).
However, Yu does not disclose an adherence parameter and [performing a first diffusion process for a first number of timesteps] based on the adherence parameter.
Zhang teaches an adherence parameter (See Zhang, Pg. 3817 Section: Classifier-free guidance resolution weighting, line 6, a user-specified weight) and [performing a first diffusion process for a first number of timesteps] based on the adherence parameter (See Zhang, Pg. 3817 Section: Classifier-free guidance resolution weighting, line 6, a user-specified weight. Examiner considers the user-specified weight to vary depending on what the user selects, and thus the level indicated can vary.
Thus, it would have been obvious to one of ordinary skills in the art before the effective filing date of the claimed invention to modify Yu’s reference to include an adherence parameter and [performing a first diffusion process for a first number of timesteps] based on the adherence parameter based on the method of Zhang’s reference. The suggestion/motivation would have been to personalize content in the generated image by finetuning the image diffusion model as suggested by Zhang at Pg. 3815, Section: Controlling Image Diffusion Models, lines 12-14.
Further, one skilled in the art could have combined the elements as described above by known method with no change in their respective functions, and the combination would have yielded nothing more than predictable results.
Therefore, it would have been obvious to combine Zhang with Yu to obtain the invention as specified in claim 8.
Regarding claim 9, in which claim 8 is incorporated, Yu discloses wherein obtaining the condition input (See Yu, ¶ [0023], based on a user input (e.g., a text prompt in conditioning input 110)) comprises:
obtaining a 3D model of an object (See Yu, ¶ [0037], a normal surface is an image which represents 3D shapes by using red/green/blue (RGB) values of the image to correspond to X, Y, and Z axis in 3D space); and
generating (See Yu, ¶ [0036], applies a convolutional kernel to the visual condition 204 to produce a feature map) a depth map based on the 3D model (See Yu, ¶ [0037], Another example of a visual condition and corresponding task instruction is a depth map… A depth map may be an image which indicates depth with luminance values. Examiner considers the depth map to include a z-axis, which is also a part of the 3D space of an image as noted above. The visual condition is used to generate a feature map, which can be a depth map that is generated), wherein the condition input comprises the depth map (See Yu, ¶ [0018], The input conditioning image may take on a variety of different forms such as a sketch, a relief map, etc. Examiner considers the depth map to be one of a variety of different forms an input conditioning image could take) and the image attribute comprises a shape of the object (See Yu, ¶ [0037], a normal surface is an image which represents 3D shapes).
Regarding claim 10, in which claim 8 is incorporated, Yu discloses wherein the first diffusion process (See Yu, ¶ [0086], a first neural network based image model (e.g., diffusion model 214)) comprises:
generating a plurality of layer-specific features based on the condition input (See Yu, Fig. 1, elements 102 and 106a to 106t; Fig. 2, elements 202 and 212; ¶ [0025], This process is repeated T times (e.g., 50 iterations) until it results in a noised latent representation zT 106t. Examiner considers zT as a plurality of layer-specific features); and
providing the plurality of layer-specific features to a plurality of corresponding layers of the first image generation model, respectively (See Yu, Fig. 1; ¶ [0027], As illustrated, denoising model εθ 112 is iteratively used to reverse the process of noising latents (i.e., perform reverse diffusion) from zT 118t to z0 118a. Denoising model εθ 112 may be a neural network based model, which has parameters that may be learned. Input to denoising model ε 112 may include a noisy latent representation (e.g., noised latent representation zT 106t)).
Regarding claim 11, in which claim 8 is incorporated, Yu discloses wherein the first diffusion process (See Yu, ¶ [0086], a first neural network based image model (e.g., diffusion model 214)) comprises:
determining a timestep threshold (See Yu, ¶ [0025], This process is repeated T times (e.g., 50 iterations) until it results in a noised latent representation zT 106t. Examiner considers the 50 iterations to be a timestep threshold) [based on the adherence parameter;] and
performing, using a first image generation model (See Yu, ¶ [0086], a first neural network based image model (e.g., diffusion model 214)), the first diffusion process (See Yu, ¶ [0086], a first neural network based image model (e.g., diffusion model 214)) for the first number of timesteps until the timestep threshold is met (See Yu, Fig. 1; ¶ [0025], This process is repeated T times (e.g., 50 iterations). Examiner considers 50 iterations to be a first number of timesteps).
However, Yu does not disclose [determining a timestep threshold] based on the adherence parameter.
Zhang teaches [determining a timestep threshold] based on the adherence parameter (See Zhang, Pg. 3817 Section: Classifier-free guidance resolution weighting, line 6, a user-specified weight).
Thus, it would have been obvious to one of ordinary skills in the art before the effective filing date of the claimed invention to modify Yu’s reference to [determining a timestep threshold] based on the adherence parameter based on the method of Zhang’s reference. The suggestion/motivation would have been to personalize content in the generated image by finetuning the image diffusion model as suggested by Zhang at Pg. 3815, Section: Controlling Image Diffusion Models, lines 12-14.
Further, one skilled in the art could have combined the elements as described above by known method with no change in their respective functions, and the combination would have yielded nothing more than predictable results.
Therefore, it would have been obvious to combine Zhang with Yu to obtain the invention as specified in claim 11.
Regarding claim 12, in which claim 11 is incorporated, Yu discloses wherein generating the synthetic image (See Yu, ¶ [0086], second latent representation) comprises:
performing, using a second image generation model (See Yu, ¶ [0088], a second neural network based image model (e.g., diffusion model 212)), a second diffusion process (See Yu, ¶ [0088], a second neural network based image model (e.g., diffusion model 212)) for a second number of timesteps after the timestep threshold (See Yu, ¶ [0025], This process is repeated T times (e.g., 50 iterations) until it results in a noised latent representation zT 106t. Examiner considers the 50 iterations to be a timestep threshold; ¶ [0027], reverse the process of noising latents (i.e., perform reverse diffusion) from zT 118t to z0 118a). Examiner considers a second number of timesteps to be once 50 iterations has been reached back down to 0 iterations in the reverse process).
Regarding claim 14, in which claim 8 is incorporated, Yu discloses the code further comprising instructions executable by the processor to perform operations comprising:
obtaining a text prompt, wherein the synthetic image is generated based on the text prompt (See Yu, ¶ [0023], a denoising diffusion model is trained to generate an image (e.g., output 116) based on a user input (e.g., a text prompt in conditioning input 110)).
Regarding claim 15, Yu discloses a memory component (See Yu, ¶ [0047], memory 720); and
a processing device coupled to the memory component (See Yu, ¶ [0047], computing device 700 includes a processor 710 coupled to memory 720), the processing device configured to perform operations comprising:
obtaining a condition input (See Yu, ¶ [0023], based on a user input (e.g., a text prompt in conditioning input 110)) and [an adherence parameter], wherein the condition input indicates an image attribute (See Yu, ¶ [0023], based on a user input (e.g., a text prompt in conditioning input 110). Examiner considers a text prompt an image attribute as it describes the image) and the [adherence parameter indicates a level of the image attribute];
generating, using a first image generation model (See Yu, ¶ [0086], a first neural network based image model (e.g., diffusion model 214)), an intermediate output (See Yu, ¶ [0086], a first latent representation) based on the condition input (See Yu, ¶ [0023], based on a user input (e.g., a text prompt in conditioning input 110)) and [the adherence parameter], wherein the intermediate output represents the image attribute (See Yu, ¶ [0086], At step 903, the system generates, by a first neural network based image model (e.g., diffusion model 214), a first latent representation based on the task-specific feature map); and
generating, using a second image generation model (See Yu, ¶ [0086], a second neural network based image model (e.g., diffusion model 212)), a synthetic image (See Yu, ¶ [0086], second latent representation) based on the intermediate output (See Yu, ¶ [0088], At step 905, the system modifies a second latent representation of a second neural network based image model (e.g., diffusion model 212) based on the first latent representation and the task embedding), wherein the synthetic image (See Yu, ¶ [0086], second latent representation) depicts the image attribute [at the level indicated by the adherence parameter].
However, Yu does not disclose an adherence parameter and [wherein the condition input indicates an image attribute] at the level indicated by the adherence parameter.
Zhang teaches an adherence parameter (See Zhang, Pg. 3817 Section: Classifier-free guidance resolution weighting, line 6, a user-specified weight) and [wherein the condition input indicates an image attribute] at the level indicated by the adherence parameter (See Zhang, Pg. 3817 Section: Classifier-free guidance resolution weighting, line 6, a user-specified weight. Examiner considers the user-specified weight to vary depending on what the user selects, and thus the level indicated can vary).
Thus, it would have been obvious to one of ordinary skills in the art before the effective filing date of the claimed invention to modify Yu’s reference to include an adherence parameter and [wherein the condition input indicates an image attribute] at the level indicated by the adherence parameter based on the method of Zhang’s reference. The suggestion/motivation would have been to personalize content in the generated image by finetuning the image diffusion model as suggested by Zhang at Pg. 3815, Section: Controlling Image Diffusion Models, lines 12-14.
Further, one skilled in the art could have combined the elements as described above by known method with no change in their respective functions, and the combination would have yielded nothing more than predictable results.
Therefore, it would have been obvious to combine Zhang with Yu to obtain the invention as specified in claim 15.
Regarding claim 16, in which claim 15 is incorporated, Yu discloses wherein obtaining the condition input (See Yu, ¶ [0023], based on a user input (e.g., a text prompt in conditioning input 110)) comprises:
obtaining a 3D model of an object (See Yu, ¶ [0037], a normal surface is an image which represents 3D shapes by using red/green/blue (RGB) values of the image to correspond to X, Y, and Z axis in 3D space); and
generating (See Yu, ¶ [0036], applies a convolutional kernel to the visual condition 204 to produce a feature map) a depth map based on the 3D model (See Yu, ¶ [0037], Another example of a visual condition and corresponding task instruction is a depth map… A depth map may be an image which indicates depth with luminance values. Examiner considers the depth map to include a z-axis, which is also a part of the 3D space of an image as noted above. The visual condition is used to generate a feature map, which can be a depth map that is generated), wherein the condition input comprises the depth map (See Yu, ¶ [0018], The input conditioning image may take on a variety of different forms such as a sketch, a relief map, etc. Examiner considers the depth map to be one of a variety of different forms an input conditioning image could take) and the image attribute comprises a shape of the object (See Yu, ¶ [0037], a normal surface is an image which represents 3D shapes).
Regarding claim 17, in which claim 15 is incorporated, Yu discloses wherein generating the intermediate output (See Yu, ¶ [0086], a first latent representation) comprises:
generating a plurality of layer-specific features based on the condition input (See Yu, Fig. 1, elements 102 and 106a to 106t; Fig. 2, elements 202 and 212; ¶ [0025], This process is repeated T times (e.g., 50 iterations) until it results in a noised latent representation zT 106t. Examiner considers zT as a plurality of layer-specific features); and
providing the plurality of layer-specific features to a plurality of corresponding layers of the first image generation model, respectively (See Yu, Fig. 1; ¶ [0027], As illustrated, denoising model εθ 112 is iteratively used to reverse the process of noising latents (i.e., perform reverse diffusion) from zT 118t to z0 118a. Denoising model εθ 112 may be a neural network based model, which has parameters that may be learned. Input to denoising model ε 112 may include a noisy latent representation (e.g., noised latent representation zT 106t)).
Regarding claim 18, in which claim 15 is incorporated, Yu discloses wherein generating the intermediate output (See Yu, ¶ [0086], a first latent representation) comprises:
determining a timestep threshold (See Yu, ¶ [0025], This process is repeated T times (e.g., 50 iterations) until it results in a noised latent representation zT 106t. Examiner considers the 50 iterations to be a timestep threshold) [based on the adherence parameter]; and
performing, using the first image generation model (See Yu, ¶ [0086], a first neural network based image model (e.g., diffusion model 214)), a first diffusion process (See Yu, ¶ [0086], a first neural network based image model (e.g., diffusion model 214)) before the timestep threshold (See Yu, Fig. 1; ¶ [0025], This process is repeated T times (e.g., 50 iterations)). examiner considers before the timestep threshold to be from 0 to 50 iterations in the forward process).
However, Yu does not disclose an adherence parameter.
Zhang teaches an adherence parameter (See Zhang, Pg. 3817 Section: Classifier-free guidance resolution weighting, line 6, a user-specified weight).
Thus, it would have been obvious to one of ordinary skills in the art before the effective filing date of the claimed invention to modify Yu’s reference to include an adherence parameter based on the method of Zhang’s reference. The suggestion/motivation would have been to personalize content in the generated image by finetuning the image diffusion model as suggested by Zhang at Pg. 3815, Section: Controlling Image Diffusion Models, lines 12-14.
Further, one skilled in the art could have combined the elements as described above by known method with no change in their respective functions, and the combination would have yielded nothing more than predictable results.
Therefore, it would have been obvious to combine Zhang with Yu to obtain the invention as specified in claim 18.
Regarding claim 19, in which claim 18 is incorporated, Yu discloses wherein generating the synthetic image (See Yu, ¶ [0086], second latent representation) comprises:
performing, using the second image generation model (See Yu, ¶ [0086], a second neural network based image model (e.g., diffusion model 212)), a second diffusion process (See Yu, ¶ [0088], a second neural network based image model (e.g., diffusion model 212)) after the timestep threshold (See Yu, ¶ [0025], This process is repeated T times (e.g., 50 iterations)); ¶ [0027], reverse the process of noising latents (i.e., perform reverse diffusion) from zT 118t to z0 118a). Examiner considers after the timestep threshold to be once 50 iterations has been reached back down to 0 iterations in the reverse process).
Claims 6, 13, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Yu et al. (US 2024/0386623 A1, hereinafter, “Yu”) in view of Zhang et al. (Adding Conditional Control to Text-to-Image Diffusion Models, 2023, hereinafter, “Zhang”), and further in view of Bie et al. (RenAIssance: A Survey into AI Text-to-Image Generation in the Era of Large Model, 2023, hereinafter, “Bie”).
Regarding claim 6, in which claim 4 is incorporated, Yu in combination with Zhang discloses wherein determining the timestep threshold (See Yu, ¶ [0025], This process is repeated T times (e.g., 50 iterations). Examiner considers the 50 iterations to be a timestep threshold) comprises: [computing a product of a total number of timesteps and the adherence parameter.]
However, Yu does not teach computing a product of a total number of timesteps and the adherence parameter.
Zhang teaches [computing a product of a total number of timesteps and] the adherence parameter (See Zhang, Pg. 3817 Section: Classifier-free guidance resolution weighting, line 6, a user-specified weight).
Thus, it would have been obvious to one of ordinary skills in the art before the effective filing date of the claimed invention to modify Yu’s reference to obtain an adherence parameter. The suggestion/motivation would have been to personalize content in the generated image by finetuning the image diffusion model as suggested by Zhang at Pg. 3815, Section: Controlling Image Diffusion Models, lines 12-14.
However, neither Yu nor Zhang disclose computing a product of a total number of timesteps [and the adherence parameter].
Bie teaches computing a product (See Bie, Pg. 5, left col., Eqn. (11), indented par. 3, lines 5-7, The loss function’s purpose is to calculate the MSE error between the predicted noise at a certain timestep and the noise that is sampled from the forward process. Examiner considers the loss function, which uses multiplication and a timestep value, to be used to determine how close the predicted and actual output are, to computing a product which can be guided by another parameter, such as an adherence parameter) of a total number of timesteps (See Bie, Pg. 5, left col., Eqn. (11), indented par. 3, lines 5-7, The loss function’s purpose is to calculate the MSE error between the predicted noise at a certain timestep and the noise that is sampled from the forward process. Examiner considers the total number of timestep to be a certain timestep, which can vary and be set to be 50 iterations for example) [and the adherence parameter.]
Thus, it would have been obvious to one of ordinary skills in the art before the effective filing date of the claimed invention to modify Yu’s or Zhang’s reference to have computing a product of a total number of timesteps based on the method of Bie’s reference. The suggestion/motivation would have been to not only ensure a high fidelity of the output image but also provide varieties of the image styles as suggested by Bie at Pg. 4, right col., lines 7-10.
Further, one skilled in the art could have combined the elements as described above by known method with no change in their respective functions, and the combination would have yielded nothing more than predictable results.
Therefore, it would have been obvious to combine Bie with Yu and Zhang to obtain the invention as specified in claim 6.
Regarding claim 13, in which claim 11 is incorporated, Yu in combination with Zhang discloses wherein determining the timestep threshold (See Yu, ¶ [0025], This process is repeated T times (e.g., 50 iterations). Examiner considers the 50 iterations to be a timestep threshold) comprises: [computing a product of a total number of timesteps and the adherence parameter.]
However, Yu does not teach computing a product of a total number of timesteps and the adherence parameter.
Zhang teaches [computing a product of a total number of timesteps and] the adherence parameter (See Zhang, Pg. 3817 Section: Classifier-free guidance resolution weighting, line 6, a user-specified weight).
Thus, it would have been obvious to one of ordinary skills in the art before the effective filing date of the claimed invention to modify Yu’s reference to obtain an adherence parameter. The suggestion/motivation would have been to personalize content in the generated image by finetuning the image diffusion model as suggested by Zhang at Pg. 3815, Section: Controlling Image Diffusion Models, lines 12-14.
However, neither Yu nor Zhang disclose computing a product of a total number of timesteps [and the adherence parameter].
Bie teaches computing a product (See Bie, Pg. 5, left col., Eqn. (11), indented par. 3, lines 5-7, The loss function’s purpose is to calculate the MSE error between the predicted noise at a certain timestep and the noise that is sampled from the forward process. Examiner considers the loss function, which uses multiplication and a timestep value, to be used to determine how close the predicted and actual output are, to computing a product which can be guided by another parameter, such as an adherence parameter) of a total number of timesteps (See Bie, Pg. 5, left col., Eqn. (11), indented par. 3, lines 5-7, The loss function’s purpose is to calculate the MSE error between the predicted noise at a certain timestep and the noise that is sampled from the forward process. Examiner considers the total number of timestep to be a certain timestep, which can vary and be set to be 50 iterations for example) [and the adherence parameter.]
Thus, it would have been obvious to one of ordinary skills in the art before the effective filing date of the claimed invention to modify Yu’s or Zhang’s reference to have computing a product of a total number of timesteps based on the method of Bie’s reference. The suggestion/motivation would have been to not only ensure a high fidelity of the output image but also provide varieties of the image styles as suggested by Bie at Pg. 4, right col., lines 7-10.
Further, one skilled in the art could have combined the elements as described above by known method with no change in their respective functions, and the combination would have yielded nothing more than predictable results.
Therefore, it would have been obvious to combine Bie with Yu and Zhang to obtain the invention as specified in claim 13.
Regarding claim 20, in which claim 18 is incorporated, Yu in combination with Zhang discloses wherein determining the timestep threshold (See Yu, ¶ [0025], This process is repeated T times (e.g., 50 iterations). Examiner considers the 50 iterations to be a timestep threshold) comprises: [computing a product of a total number of timesteps and the adherence parameter.]
However, Yu does not teach computing a product of a total number of timesteps and the adherence parameter.
Zhang teaches [computing a product of a total number of timesteps and] the adherence parameter (See Zhang, Pg. 3817 Section: Classifier-free guidance resolution weighting, line 6, a user-specified weight).
Thus, it would have been obvious to one of ordinary skills in the art before the effective filing date of the claimed invention to modify Yu’s reference to obtain an adherence parameter. The suggestion/motivation would have been to personalize content in the generated image by finetuning the image diffusion model as suggested by Zhang at Pg. 3815, Section: Controlling Image Diffusion Models, lines 12-14.
However, neither Yu nor Zhang disclose computing a product of a total number of timesteps [and the adherence parameter].
Bie teaches computing a product (See Bie, Pg. 5, left col., Eqn. (11), indented par. 3, lines 5-7, The loss function’s purpose is to calculate the MSE error between the predicted noise at a certain timestep and the noise that is sampled from the forward process. Examiner considers the loss function, which uses multiplication and a timestep value, to be used to determine how close the predicted and actual output are, to computing a product which can be guided by another parameter, such as an adherence parameter) of a total number of timesteps (See Bie, Pg. 5, left col., Eqn. (11), indented par. 3, lines 5-7, The loss function’s purpose is to calculate the MSE error between the predicted noise at a certain timestep and the noise that is sampled from the forward process. Examiner considers the total number of timestep to be a certain timestep, which can vary and be set to be 50 iterations for example) [and the adherence parameter.]
Thus, it would have been obvious to one of ordinary skills in the art before the effective filing date of the claimed invention to modify Yu’s or Zhang’s reference to have computing a product of a total number of timesteps based on the method of Bie’s reference. The suggestion/motivation would have been to not only ensure a high fidelity of the output image but also provide varieties of the image styles as suggested by Bie at Pg. 4, right col., lines 7-10.
Further, one skilled in the art could have combined the elements as described above by known method with no change in their respective functions, and the combination would have yielded nothing more than predictable results.
Therefore, it would have been obvious to combine Bie with Yu and Zhang to obtain the invention as specified in claim 20.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure:
Balaji et al. (eDiff-I: Text-to-Image Diffusion Models with an Ensemble of Expert Denoisers, 2023) discloses a diffu8sion model for text to image generation. The proposed solution involved training a diffusion model specialized for different synthesis stages. A single model is trained, and then split into specialized models that are then further trained. The goal is to keep cost the same while generating high visual quality images.
Yamada (US 2025/0104291 A1) discloses a method of generating an image using text and layout information as input. Trained neural networks predict noise from the first image representation. The image is denoised to create another, or next, image representation. The aim is to improve accuracy of the neural networks, or diffusion models, used to create reliable images even when two distinct objects are given as inputs while also reducing trial costs.
Wang et al. (US 2025/0068298 A1) discloses an image generation method capable of multiple modes, one being text-to image generation, using text information given by a user. Sample images or style references can also be used as input to help generate at least one target image aligned to the text prompt. The invention aims to accurately express the image a user wants to generate given multiple inputs to the artificial intelligence model.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Jasmin Marcelino Hernandez whose telephone number is (571) 270-0211. The examiner can normally be reached 7am-3pm EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Henok Shiferaw can be reached at (571) 272-4637. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/JASMIN MARCELINO HERNAND/Examiner, Art Unit 2676
/Henok Shiferaw/Supervisory Patent Examiner, Art Unit 2676