DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to non-statutory subject matter. The claimed invention is directed to non-statutory subject matter because the claim(s) as a whole, considering all claim elements both individually and in combination, do not amount to significantly more than an abstract idea. As summarized in the 2019 Revised Patent Subject Matter Eligibility Guidance, examiners must perform a Two-Part Analysis for Judicial Exceptions.
Step 1
In Step 1, it must be determined whether the claimed invention is directed to a process, machine, manufacture or composition of matter. The instant invention encompasses a method in claims 1-11 (i.e., a process); an electronic device in claims 12-19 (i.e. a machine); a non-transitory computer program product in claim 20 (i.e., a manufacture). All claims are directed to one of the four statutory categories and meet the requirements of step 1.
Step 2A
Prong One
The claimed invention is directed to an abstract idea without significantly more. The instant invention is broadly directed to image generation method. Claim 1 recites the following (with emphasis added):
Claim 1: A method for image generation, comprises:
obtaining a plurality of images by an image generating model based on a prompt, the plurality of images comprising a plurality of instances of an object, respectively, the object being specified by the prompt;
determining a plurality of attributes of the plurality of instances of the object, respectively; and
updating the image generating model based on the plurality of attributes and a predetermined distribution of a plurality of predetermined attributes related to the object.
Claim 1 recites the steps for data obtaining and updating a model based on mathematical calculations, which is directed to mathematical relationships and calculations.
Claim 1 encompasses the abstract idea, which is also encompassed by the dependent claims 2-11, which are about more mathematical data and calculations.
Claims 12-20 recite similar limitations of claim 1-11, thus are abstract ideas.
Prong Two
This judicial exception is not integrated into a practical application because mere instruction to implement on a computer or a computer model, or merely using a computer or computer model as a tool to perform the abstract idea, adding insignificant extra solution activity, and/or generally linking the use of the abstract idea to a technological environment or field of use is not considered integration into a practical application. Claim 1 recites updating an image generation model, and further in claim 3 defines the model as a diffusion model, i.e. a machine learning model. Claim 1 is using data to train a machine learning model. Using training data to train a machine learning model is a generic feature of machine learning model, which does not represent technological improvement. Claim 12 and 20 recite using a generic computer to implement the abstract idea. The using of the computer and the machine learning model does not add improvement to the functioning of a computer or to any other technology field, which failed to enable the abstract idea to integrate into a practical application. Claims 2-10, 12-19 are about mathematical relationships and calculations, which are abstract idea. Claim 3, 14 recites a specific network model. Claim 11 recites using the machine learning model to acquire data and specific data, which is data obtaining. The claims do not include additional elements that are sufficient to enable the abstract idea to integrate into a practical application.
Step 2B
Step 2B in the analysis requires us to determine whether the claims do significantly more than simply describe that abstract method. Mayo, 132 S. Ct. at 1297. We must examine the limitations of the claims to determine whether the claims contain an "inventive concept" to "transform" the claimed abstract idea into patent-eligible subject matter. Alice, 134 S. Ct. at 2357 (quoting Mayo, 132 S. Ct. at 1294, 1298). The transformation of an abstract idea into patent-eligible subject matter "requires 'more than simply stat[ing] the [abstract idea] while adding the words 'apply it."' Id. (quoting Mayo, 132 S. Ct. at 1294) (alterations in original). "A claim that recites an abstract idea must include 'additional features' to ensure 'that the [claim] is more than a drafting effort designed to monopolize the [abstract idea].'" Id. (quoting Mayo, 132 S. Ct. at 1297) (alterations in original). Those "additional features" must be more than "well-understood, routine, conventional activity." Mayo, 132 S. Ct. at 1298.
The present claims include the additional elements other than the abstract idea which include a computer. These additional elements are merely conventional computer. Any potentially technical aspects of the claims are well-known generic computer components performing conventional functions (e.g., a processor performing generic data handling using mathematical concepts). The present claims have been analyzed both individually and in combination and, the instant claims do not provide any improvement of the functioning of the computer or improvement to computer technology or any other technical field. There do not appear to be any meaningful limitations other than those that are well-understood, routine and conventional in the field. Thus, the present claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception.
The claims are generally linked to implementing an abstract idea on a computer. When looked at individually and as a whole, the claim limitations are determined to be an abstract idea without "significantly more," and thus not patent eligible.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-7, 9-18, 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Shen et al. (US 2025/0111553 A1).
Regarding claim 1, Shen teaches:
A method for image generation, comprises:
obtaining a plurality of images by an image generating model based on a prompt, ( [0080], “The diffusion model generates an initial batch of images based on the user input prompts.”) the plurality of images comprising a plurality of instances of an object, respectively,([0127], “FIG. 8 shows the images generated from the original stable diffusion model (8a) and the jointly debiased stable diffusion model (8b) for gender and race, according to an embodiment of the invention. It demonstrates the effectiveness of the debiasing process. The models are evaluated using an unseen occupation prompt: “a photo of the face of an electrical and electronics repairer, a person.””) the object being specified by the prompt; ([0077], “At step 401, it receives the text input. The method receives a user-provided text input through a user terminal device. This text serves as the basis for generating an image. For example, the text input can be “a photo of the face of an electrical and electronics repairer”.”)
determining a plurality of attributes of the plurality of instances of the object, respectively; ([0081], “pre-trained classifiers are used to detect and identify specific attributes in the generated images, such as age, gender, race, or other relevant features.”) and
updating the image generating model based on the plurality of attributes and a predetermined distribution of a plurality of predetermined attributes related to the object.([0080], “ In some embodiments, the user defines a target distribution for one or more attributes, such as gender or race. The diffusion model generates an initial batch of images based on the user input prompts. The generated images are compared to the target distribution, and a loss is calculated based on the deviation from the target. The model is then adjusted iteratively to minimize the distributional alignment loss, steering the characteristics of the images toward the user-defined target. This loss function dynamically steers image generation towards a user-defined distribution of attributes like gender, race, or age while maintaining the integrity and semantics of the generated images.”)
The above citations are from different embodiments/examples of Shen. However, it have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to have combined the different embodiments/examples of Shen to generate an image generating model that allows for debiasing prompts, improving the inclusiveness of diffusion model outputs across varied demographics. (Shen, Abstract)
Regarding claim 2, Shen teaches:
The method of claim 1, wherein updating the image generating model comprises: obtaining a loss based on a distribution of the plurality of attributes and the predetermined distribution of the plurality of predetermined attributes; ([0081], “In some embodiments, the step of aligning the generated images involves the following process: pre-trained classifiers are used to detect and identify specific attributes in the generated images, such as age, gender, race, or other relevant features. The identified attributes are then aligned toward a predefined target attribute distribution. This alignment is achieved by applying the distributional alignment loss, which adjusts the generated images to ensure that their attributes match the target distribution.”) and updating the image generating model based on the loss. ([0080], “In some embodiments, the user defines a target distribution for one or more attributes, such as gender or race. The diffusion model generates an initial batch of images based on the user input prompts. The generated images are compared to the target distribution, and a loss is calculated based on the deviation from the target. The model is then adjusted iteratively to minimize the distributional alignment loss, steering the characteristics of the images toward the user-defined target. This loss function dynamically steers image generation towards a user-defined distribution of attributes like gender, race, or age while maintaining the integrity and semantics of the generated images.”)
Regarding claim 3, Shen teaches:
The method of claim 2, wherein the image generating model comprises a diffusion model (FIG. 3, 310) and a distribution guidance parameter for adjusting the diffusion model, the distribution guidance parameter comprises a weight vector, and a plurality of weights in the weight vector corresponding to the plurality of predetermined attributes respectively, ([0092], “At step 501, it prepares the gradient coefficient. In this step, a gradient coefficient is prepared. This coefficient serves as a key factor in adjusting the model's gradient, ensuring that the model is guided to reduce specific types of biases during the optimization process.” For a neural network, the gradient are vectors. ) and obtaining the loss comprises: determining the loss based on the weight vector, the distribution of the plurality of attributes, and the predetermined distribution of the plurality of predetermined attributes. (the distribution output of the a diffusion model is based on the gradient values.
“
PNG
media_image1.png
634
480
media_image1.png
Greyscale
”)
Regarding claim 4, Shen teaches:
The method of claim 3, wherein determining the loss comprises: determining a distribution vector based on the weight vector and the plurality of attributes of the plurality of instances; ((the distribution output of the diffusion model is based on the gradient values and its K classes [0106]) and determining the loss based on a distance between the distribution vector and the predetermined distribution of the plurality of predetermined attributes.
(
PNG
media_image1.png
634
480
media_image1.png
Greyscale
)
Regarding claim 5, Shen teaches:
The method of claim 4, wherein updating the image generating model based on the loss comprises: determining the weight vector for updating the diffusion model by minimizing the loss.(
PNG
media_image1.png
634
480
media_image1.png
Greyscale
)
Regarding claim 6, Shen teaches:
The method of claim 5, wherein determining weight vector comprises: setting the weight vector comprised in the image generating model to an initial weight vector; ([0092], “At step 501, it prepares the gradient coefficient. In this step, a gradient coefficient is prepared. This coefficient serves as a key factor in adjusting the model's gradient, ensuring that the model is guided to reduce specific types of biases during the optimization process.”) determining the distribution vector based on the image generating model that comprises the weight vector; ([0106], “The system first generates a batch of images custom-character={x.sup.(i)}i∈|N|] using the diffusion model being finetuned and some prompt P. For every generated image x.sup.(i), the system uses a pre-trained classifier h to produce a class probability vector p.sup.(i)=[p.sub.1.sup.(i), . . . , p.sub.K.sup.(i)=h(x.sup.(i)), with P.sub.k.sup.(i) denoting the estimated probability that x.sup.(i) is from class k.”) and updating the weight vector based on the loss that is determined based on the distribution vector and the predetermined distribution of the plurality of predetermined attributes. ([0096], “At step 503, in this final step, the adjusted gradient is backpropagated through the diffusion model. After computing the gradient from the generated image, the system backpropagates through the model, adjusting not just the final layers responsible for image creation, but also the earlier layers linked to interpreting the text prompt. This way, the entire model, from text input to final image output, is refined to produce fairer images.” And [0106])
Regarding claim 7, Shen teaches:
The method of claim 6, wherein updating the weight vector based on the distribution vector comprises: in at least one round, in response to determining that the loss determined based on the updated weight vector does not meet a stopping criterion, updating the weight vector based on a difference between the distribution vector and the predetermined distribution. ([0080], “In some embodiments, the user defines a target distribution for one or more attributes, such as gender or race. The diffusion model generates an initial batch of images based on the user input prompts. The generated images are compared to the target distribution, and a loss is calculated based on the deviation from the target. The model is then adjusted iteratively to minimize the distributional alignment loss, steering the characteristics of the images toward the user-defined target. This loss function dynamically steers image generation towards a user-defined distribution of attributes like gender, race, or age while maintaining the integrity and semantics of the generated images.” [0084], “In some embodiments, the process iteratively adjusts the model until the generated images align closely with the target demographic distribution. The distributional alignment loss is applied iteratively during the image generation process. With each iteration, the generated image features are progressively adjusted to reduce the discrepancy between the current distribution and the target distribution. This iterative application continues until the features of the generated images align with the predefined target distribution, ensuring that the final outputs reflect the desired balance of attributes.”)
Regarding claim9, Shen teaches:
The method of claim 1, wherein determining the plurality of attributes of the plurality of instances comprises: with respect to an instance in an image in the plurality of images: detecting a region of interest from the image based on image recognition;([0109], “For the current invention, the system focuses on face-centric attributes such as gender, race, and age. It is found that the following adaptation from the general case yields the best results. First, the system uses a face detector d.sub.face to retrieve the face region d.sub.face(x.sup.(i)) from every generated image x.sup.(i).”) determining a plurality of similarities between an image content in the region of interest and the plurality of predetermined attributes related to the object; ([0109], “The system applies the classifier h and the DAL custom-character.sub.align only on the face regions. Second, the system introduces another face realism preserving loss custom-character.sub.face, which penalize the dissimilarity between the generated face d.sub.face(x.sup.(i)) and the closest face from a set of external real faces custom-character.sup.F,”) and determining the attribute of the instance based on the plurality of similarities.([0109]: “… [00009]ℒface=1N(1-minF∈𝒟Fcos(emb(dface(x(i))),emb(F)).(8)
where emb(.) is a face embedding model. [AltContent: rect].sub.face helps retain realism of the faces, which can be substantially edited by the DAL. In the implementation, the system uses the CelebA and the FairFace dataset as external faces. The system uses the SFNet-20 (as the face embedding model). The CelebA dataset is a large-scale face dataset commonly used for facial recognition, attribute prediction, and other computer vision tasks. ”)
Regarding claim 10, Shen teaches:
The method of claim 6, wherein the predetermined distribution of the plurality of predetermined attributes is determined by any of: a uniform distribution; or a distribution determined by respective frequencies of respective predetermined attributes among the plurality of predetermined attributes.([0083], “A target distribution for each attribute is specified. For example, the user may define a uniform distribution that 50% of the generated images should represent women and 50% men,”)
Regarding claim 11, Shen teaches:
The method of claim 1, further comprises: inputting a target prompt into the image generating model, the target prompt instructing the image generating model to generate a target image that comprises a target object, and the target object being specified in the prompt; (FIG. 2, prompt input from user; [0052], “he method 200 begins with receiving an input text from a user. The input is first processed by a text encoder 210, which parses and understands the textual input, extracting key semantic and contextual features. This understanding helps in identifying the specific characteristics the user wants in the generated image, such as objects, settings, or specific traits like gender or race.”) and receiving the target image from the image generating model. (FIG. 2, output image 230)
Regarding claim 12, Shen teaches:
An electronic device, comprising a computer processor coupled to a computer-readable memory unit, the memory unit comprising instructions that when executed by the computer processor implements a method for image generation, ([0156], “A system for generating images from a text to reduce bias using a diffusion model, comprising: a processor; a memory in electronic communication with the processor; and instructions stored in the memory and executable by the processor to cause the system to perform operations for:”) the rest of claim 12 recites similar limitations of claim 1, thus is rejected accordingly.
claim 20 recites similar limitations of claim 12, thus is rejected accordingly.
claim 13-18 recite similar limitations of claim 2-7 respectively, thus is rejected accordingly.
Claim(s) 8, 19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Shen in view of El-Khamy et al. (US 2024/0176986 A1).
Regarding claim 8, Shen teaches:
The method of claim 1, wherein determining the weight vector comprises: determining a weight vector space with a center at an initial weight vector, the weight vector space comprises a plurality of weight vectors that follow a predetermined distribution; (The diffusion model gradients change center around those at the time step 0 following the noise distribution Z) . selecting a group of weight vectors from the plurality of weight vectors; (select one group of adjusted gradients)
However, Shen does not explicitly, but El-Khamy teaches:
determining a group of rewards for the image generating model based on the loss and the group of weight vectors; determining the weight vector by updating the initial weight vector with the group of rewards.([0056], “A loss function (e.g., a loss function based on cross entropy, size of the network, complexity of the parameter) of the super-network may be used as a reward function to update the controller.”)
Shen teaches using calculated loss to update the model gradients. EL-Khamy teaches rewards determined based on the loss are used to update the model gradients.
It would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to have replaced using loss taught by Shen by using reward taught by EL-Khamy to update the model gradients with reasonable expectation of success.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to YANNA WU whose telephone number is (571)270-0725. The examiner can normally be reached Monday-Thursday 8:00-5:30 ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Alicia Harrington can be reached at 5712722330. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/YANNA WU/Primary Examiner, Art Unit 2615 Any inquiry concerning this communication or earlier communications from the examiner should be directed to YANNA WU whose telephone number is (571)270-0725. The examiner can normally be reached Monday-Thursday 8:00-5:30 ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Alicia Harrington can be reached at 5712722330. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/YANNA WU/Primary Examiner, Art Unit 2615