Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is
directed to an abstract idea without reciting elements that amount to significant more than
the abstract idea. The rationale for this rejection, under MPEP § 2106, for this finding is
explained below.
Step 1: Under step 1, the claims are analyzed to determine if the claim is directed to a
process, machine, article of manufacture, or composition of matter. For the claims in question,
claims 1-10 and 20 are directed towards a machine, and claims 11-19 are directed towards a method.
Step 2A, Prong 1: Under step 2A, prong 1, the claims are evaluated to determine if the
claim recites a judicial exception, which includes the laws of nature, physical phenomena, or an
abstract idea. For independent claim 1 (and corresponding dependent claims 11 and 20), the limitations regarding minimizing a loss function (i.e., “a first loss function”, “a second loss function”, and “a third loss function”) are directed towards a mathematical concept. The claimed loss functions relate to mathematical relationships used to train machine learning models.
Step 2A, Prong 2: Under step 2A, prong 2, the claims are evaluated to determine
whether the claim as a whole integrates the recited judicial exception into a practical application
of the exception (see MPEP 2106.04(d)). The examiner notes that MPEP 2106.05(a) -(c) and (e)
generally concern limitations that are indicative of integration, whereas 2106.05(f)-(h) generally
concern limitations that are not indicative of integration.
Regarding independent claims 1, 11, and 20, the limitations as described above are directed towards a mathematical concept. The additional limitations regarding utilizing a processor, memory component, and/or non-transitory computer-readable storage medium are mere instructions to implement an idea on a computer and/or use a computer to perform an abstract idea, and are not indicative of integration into a practical application (see MPEP 2106.05(f)). The limitation relating to generating an output image using a machine learning model is considered to be extra-solution activity and is not indicative of integration into a practical application (see MPEP 2106.05(g)). The limitations relating to training a first machine learning model and second machine learning model are broadly recited and generally linked to the field of performing image processing using machine learning models, and is not indicative of integration into a practical application (see MPEP 2106.05(h)).
Regarding dependent claims 2-10 and 12-19, the additional limitations are broadly recited and further disclose steps used to perform the judicial exception with any clear indication or detail which would indicate integration into a practical application as noted in MPEP 2106.05(a) or MPEP 2106.05(e). Additionally, multiple dependent claims further claim additional loss functions, which are directed towards a mathematical concept. Therefore, the additional limitations of claims 2-10 and 12-19 do not constitute integration into a practical application.
The examiner emphasizes MPEP 2106.05(a), which states that a limitation is indicative of integration into a practical application if the limitation identifies a manner in which an improvement is explicitly and specifically achieved and recited in the claims. The current claim language all are recited at a high level of generality which do not serve to integrate the limitations in view of MPEP 2106.05(f), and furthermore nothing precludes the current limitations from being interpreted under the mental processes grouping.
The Examiner notes the Ex Parte Desjardins decision issued on 09/26/2025 (https://www.uspto.gov/sites/default/files/documents/202400567-arp-rehearing-decision-20250926.pdf), and how the claims discussed in Desjardins are different from the claims currently presented. The invention in Desjardins pertains to a training method used to train a machine learning model, and claim 1 specifically discloses the computation of several parameters associated with the machine learning model and consequently adjusting the parameters to optimize an objective function. While the limitations are directed towards a mathematical concept at prong one, the limitations are specific and furthermore the Appeals Review Panel directs the limitations towards an improvement to the training of the machine learning model (noting paragraph 21 of the specification of the invention). In contrast to the present application, the limitations currently presented in the independent claims are broad, and the mathematical concepts involving the three loss functions are simply presented as being based on various images produced by the two machine learning models, and at present is unclear how these various loss functions are involved and consequently direct towards an improvement to the technical field.
The Examiner notes that specific portions of the Applicant’s specification as well as the dependent claims do suggest potential improvements/advantages, which are important to note when consider the 35 U.S.C. 101 analysis at Step 2A prong one. Notably, [0096] discloses the one of the claimed loss functions as an annealing identity loss function, which “advantageously maintains strong identity preservations early in the image generation process…while gradually transitioning to allow more style modifications to be made as the image generation process progresses.”. The gradual transition process “provides improvements to the image generation process by reducing visible artifacts…while facilitating uniform style modifications”. Similarly, [0102] and [0110] provide further disclosure regarding how the other claimed loss functions advantageously improve the claimed image generation process.
Step 2B: Under step 2B, the claims are evaluated as a whole to determine if it amounts to
significantly more than the recited exception (i.e., whether any additional element, or
combination of additional elements, adds an inventive concept to the claim). The considerations
of step 2A, prong 2 and step 2B overlap, but differ in that 2B also requires considering the claim
as a whole/combination of limitations, and with reference to MPEP 2106.05(d) whether the
claims feature any “specific limitation(s) other than what is well - understood, routine,
conventional activity in the field” (WURC). The examiner asserts that, even when considered in
combination, the additional elements of claims 1-20 represent mere instructions to train machine learning models utilizing mathematical concepts in the form of loss functions, at a high level of generality that is generally linked to the training of any machine learning model, and therefore does not provide a specifically recited inventive concept.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-3, 6, 9-13, 16, and 19-20 are rejected as being unpatentable over Sauer et al. (“Adversarial Diffusion Distillation”, DOI: 10.48550/arXiv.2311.17042, Publication Year: 2023; hereinafter “Sauer”) in view of Brooks et al. (“InstructPix2Pix: Learning to Follow Image Editing Instructions”, DOI: 10.48550/arXiv.2211.09800, Publication Year: 2022; hereinafter “Brooks”) in view of Feng et al. (“Triplet Distillation for Deep Face Recognition”, DOI: 10.48550/arXiv.1905.04457, Publication Year: 2019; hereinafter “Feng”).
Regarding Claim 1, Sauer discloses a system comprising:
at least one processor; and at least one memory component storing instructions that, when executed by the at least one processor, cause the at least one processor to perform operations comprising (The Examiner notes that the work disclosed by Sauer consists of utilizing various deep learning models with 108-109 model parameters, which is being performed on a computing device consisting of the claimed processor and memory component.):
training a first machine learning model to generate a first output image based on (Fig. 2, Sauer discloses training a DM-teacher model to generate an output image
x
^
φ
based on a target image
x
^
θ
.);
training a second machine learning model to generate a second output image based on model comprising (Fig. 2, Sauer discloses training an ADD-student model to generate an output image
x
^
θ
based on target image
x
0
.):
minimizing a first loss function based on the first output image and the second output image (Fig. 2, Sauer discloses minimizing a distillation loss between
x
^
θ
(i.e., the output from the student/second machine learning model) and
x
^
φ
(i.e., the output from the teacher/first machine learning model).);
minimizing a second loss function based on the second output image and the target image (Fig. 2, 3.1. Training Procedure, Sauer discloses minimizing an adversarial loss between
x
0
(i.e., a real/target image) and
x
^
θ
(i.e., a fake/output image).); and
generating a third output image using the second machine learning model (Figs. 3-4, Appendix D. Additional Samples, Sauer discloses training a model based on the adversarial diffusion distillation technique to generate an image.).
Sauer does not explicitly disclose training a first machine learning model to generate a first output image based on an input image and a target image (italicized for context), training a second machine learning model to generate a second output image based on the input image and the target image (italicized for context), and minimizing a third loss function based on the second output image, the input image, and the target image.
Brooks discloses training a first machine learning model to generate a first output image based on an input image and a target image (italicized for context), training a second machine learning model to generate a second output image based on the input image and the target image (italicized for context) (3.2. InstructPix2Pix, Brooks discloses training a conditional diffusion model using generated training data. The Examiner notes the pair of training images shown in Fig. 3b, wherein the left image (i.e., the photograph of the girl on a horse) is analogous to the claimed “input image” and the right image (i.e., the image of the girl on a dragon) is analogous to the claimed “target image”.),
Sauer and Brooks are considered to be analogous to the claimed invention as they are in the same field of applying deep learning methods for generate modified images. Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Sauer such that the teacher and student networks disclosed by Sauer take in the pair of input and target image inputs disclosed by Brooks. The motivation for this combination being the ability to utilize the pair of images to guide the image editing performed by the teacher and student models.
Sauer in view of Brooks does not explicitly teach minimizing a third loss function based on the second output image, the input image, and the target image.
Feng discloses minimizing a third loss function based on the second output image, the input image, and the target image (3.2. Triplet Distillation, Feng discloses minimizing a triplet loss between an anchor image, a positive image, and a negative image.).
Sauer, Brooks and Feng are considered to be analogous to the claimed invention as they are in the same field of training deep learning methods based on an image input and minimizing a training loss. Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Sauer in view of Brooks by incorporating Feng’s disclosure of minimizing a triplet loss. The motivation for this combination being the ability to incorporate the similarity information distilled from the teacher model, which improves the overall model performance (see Abstract, Feng).
Claims 11 and 20 are the method and non-transitory computer-readable storage medium claims, respectively, corresponding to claim 1, and are similarly rejected (see 3. Method, Sauer regarding the method claim, and additionally the assertion made above regarding the disclosure of Sauer being performed on a computing device including a processor, memory, and non-transitory computer-readable storage medium.).
Regarding Claim 2, Sauer in view of Brooks in view of Feng teaches the system of claim 1, the operations further comprising:
generating a set of identity prompts describing attributes of a subject (3.1.1. Generating Instructions and Paired Captions, Fig. 2, Brooks discloses an input caption describe attributes of an image.);
generating a set of instruction prompts describing modifications (3.1.1. Generating Instructions and Paired Captions, Fig. 2, Brooks discloses feeding the input caption into a GPT-3 model to generate an modification instruction);
generating a set of target prompts based on the set of identity prompts and the set of instruction prompts (3.1.1. Generating Instructions and Paired Captions, Fig. 2, Brooks discloses combining the input caption and the modification instruction to generate an edited caption.); and
generating a training data set comprising the input image and the target image based on the set of target prompts and the set of identity prompts (3.1.1. Generating Instructions and Paired Captions, Fig. 2, Brooks discloses generating pairs of training examples based on the input captions and edited captions.).
Claim 12 is the method claim corresponding to claim 2, and is similarly rejected.
Regarding Claim 3, Sauer in view of Brooks in view of Feng teaches the system of claim 2, wherein generating the training data set comprises: generating a set of target images comprising the target image based on the set of target prompts; and generating a set of input images comprising the input image based on the set of target images and the set of identity prompts (3.1.1. Generating Instructions and Paired Captions, Fig. 2, Brooks discloses generating pairs of training examples including an image based on the input caption and an image based on the edited caption.).
Claim 13 is the method claim corresponding to claim 3, and is similarly rejected.
Regarding Claim 6, Sauer in view of Brooks in view of Feng teaches the system of claim 1, wherein minimizing the first loss function comprises: training a discriminator model to distinguish between images generated by the second machine learning model and target images (Fig. 2, 3.1. Training Procedure, 3.2. Adversarial Loss, Sauer discloses training a discriminator to distinguish between generated images
x
^
θ
and real target images
x
0
.); and minimizing a difference between outputs of the discriminator model for the images generated by the second machine learning model (Fig. 2, 3.1. Training Procedure, 3.2. Adversarial Loss, Equation 3, Sauer discloses minimizing a discriminator loss function.).
Claim 16 is the method claim corresponding to claim 6, and is similarly rejected.
Regarding Claim 9, Sauer in view of Brooks in view of Feng teaches the system of claim 1, wherein generating the third output image is based on an input instruction (Figs. 3-4, Appendix D. Additional Samples, Sauer discloses training a model based on the adversarial diffusion distillation technique to generate an image. The Examiner notes that the images generated are based on an input instruction (i.e., “A cinematic shot of a professor sloth wearing a tuxedo at a BBQ party.” as seen in Fig. 3.)).
Regarding Claim 10, Sauer in view of Brooks in view of Feng teaches the system of claim 9.
The current combination of Sauer in view of Brooks in view of Feng does not explicitly teach the operations further comprising: generating a fourth output image using the second machine learning model based on the third output image and the input instruction.
Brooks further discloses the operations further comprising: generating a fourth output image using the second machine learning model based on the third output image and the input instruction (Fig. 11, Brooks discloses repeatedly modifying a base image based on an input instruction. The Examiner notes the current specification for the application does not provide any detail or further clarification on what the claimed “fourth output image” is, how the “fourth output image” differs from any of the other output images generated by the second machine learning model.).
It would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to further modify the current combination of Sauer in view of Brooks in view of Feng to include Brook’s disclosure of producing multiple edits to the same image. The motivation for this combination being the ability to modify an initial image several times and not be limited to only one modification.
Claim 19 is the method claim corresponding to claim 10, and is similarly rejected.
Claims 4-5 and 14-15 are rejected as being unpatentable over Sauer in view of Brooks in view of Huang et al. (“DreamTime: An Improved Optimization Strategy for Diffusion-Guided 3D Generation”, DOI: 10.48550/arXiv.2306.12422, Publication Year: 2024; hereinafter “Huang”).
Regarding Claim 4, Sauer in view of Brooks in view of Feng teaches the system of claim 1.
Sauer in view of Brooks in view of Feng does not explicitly disclose wherein training the first machine learning model comprises: minimizing a fourth loss function based on the input image and the first output image, the fourth loss function comprising a decreasing weight.
Huang discloses wherein training the first machine learning model comprises: minimizing a fourth loss function based on the input image and the first output image, the fourth loss function comprising a decreasing weight (Fig. 6, 3.3. Time Prioritized Score Distillation, Huang discloses a prior weight function W(t) applied to a loss function. The Examiner notes in Fig. 6 wherein Huang discloses the weight function W(t) includes a decreasing portion.).
Sauer, Brooks, Feng, and Huang are considered to be analogous to the claimed invention as they are in the same field of training deep learning models for generating 2D images. Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Sauer in view of Brooks in view of Feng such that it further incorporated the weighted loss function disclosed by Huang. The motivation for this combination being the ability to further specific loss function which can apply weighted emphasis during particular portions of the model training process.
Claim 14 is the method claim corresponding to claim 4, and is similarly rejected.
Regarding Claim 5, the current combination of Sauer in view of Brooks in view of Feng in view of Huang teaches the system of claim 4.
The current combination of Sauer in view of Brooks in view of Feng in view of Huang does not explicitly teach wherein training the first machine learning model further comprises: minimizing a fifth loss function based on the first output image and the target image, the fifth loss function comprising a comparison of predicted noise associated with the first output image with noise associated with the target image.
Brooks further discloses wherein training the first machine learning model further comprises: minimizing a fifth loss function based on the first output image and the target image, the fifth loss function comprising a comparison of predicted noise associated with the first output image with noise associated with the target image (3.2. InstructPix2Pix, Equation 1, Brooks discloses minimizing a latent diffusion objective.).
It would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to further modify the current invention of Sauer in view of Brooks in view of Feng such that it further incorporated latent diffusion loss function disclosed by Brooks. The motivation for this combination being the ability to train the model by also taking into factor the differences in predicted noise and true noise in the target image.
Claim 15 is the method claim corresponding to claim 5, and is similarly rejected.
Claims 7 and 17 are rejected as being unpatentable over Sauer in view of Brooks in view of Feng in view of Pan et al. (“Effective Real Image Editing with Accelerated Iterative Diffusion Inversion”, DOI: 10.48550/arXiv.2309.04907, Publication Year: 2023; hereinafter “Pan”).
Regarding Claim 7, Sauer in view of Brooks in view of Feng teaches the system of claim 1, wherein minimizing the second loss function comprises: applying a sampling process to the second output image (Fig. 2, Sauer discloses adding noise to an output image by sampling from a normal distribution
ε
,
ε
'
~
N
(
0
.
I
)
.);
Sauer in view of Brooks in view of Feng does not explicitly teach and applying an inversion process to the second output image.
Pan discloses and applying an inversion process to the second output image (3.1. Diffusion Inversion Preliminaries, Pan discloses a DDIM inversion process that can be applied to an image.).
Sauer, Brooks, Feng, and Pan are considered to be analogous to the claimed invention as they are in the same field of generating images using deep learning model. Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Sauer in view of Brooks in view of Feng such that it further included Pan’s DDIM inversion process which can be applied to the second output image taught by Sauer in view of Brooks in view of Feng. The motivation for this combination being the ability to reliably reconstruct an image.
Claim 17 is the method claim corresponding to claim 7, and is similarly rejected.
Claims 8 and 18 are rejected as being unpatentable over Sauer in view of Brooks in view of Feng in view of Philbin et al. (US 2016/0180151; hereinafter “Philbin”).
Regarding Claim 8, Sauer in view of Brooks in view of Feng teaches the system of claim 1, wherein minimizing the third loss function comprises:
minimizing a difference between a first distance and a second distance to within a margin parameter, the first distance being between the input image embedding and the output image embedding, the second distance being between the output image embedding and the target image embedding (3.2. Triplet Distillation, Equation 1, 2, and 4, Feng discloses a triplet loss, wherein minimizing the triplet loss involves margin term. The Examiner notes that Feng discloses both a traditional margin term m as well as a dynamic margin term
F
(
d
)
.).
Sauer in view of Brooks in view of Feng does not explicitly disclose generating an input image embedding based on the input image; generating an output image embedding based on the second output image; generating a target image embedding based on the target image.
Philbin discloses generating an input image embedding based on the input image; generating an output image embedding based on the second output image; generating a target image embedding based on the target image ([0041-0042], Philbin discloses obtaining embedding vectors for an anchor image, a positive image, and a negative image to be used to minimize a triplet loss (i.e., the third loss function).).
Sauer, Brooks, Feng, and Philbin are considered to be analogous to the claimed invention as they are in the same field of generating and comparing images using deep learning models. Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Sauer in view of Brooks in view of Feng such that it further included Philbin’s disclosure regarding obtaining embedding vectors associated with an anchor image, positive image, and negative image, such that the embedding vectors are then used to minimize the triplet loss taught by Sauer in view of Brooks in view of Feng. The motivation for this combination being the ability to convert the images into a lower dimensional embedded space, improving computational efficiency when minimizing the loss function.
Claim 18 is the method claim corresponding to claim 8, and is similarly rejected.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Hsiao et al. (“ReF-LDM: A Latent Diffusion Model for Reference-based Face Image Restoration”, DOI: 10.48550/arXiv.2412.05043, Publication Year: 2024)
Ganguly et al. (“AdaKD: Dynamic Knowledge Distillation of ASR Models using Adaptive Loss Weighting”, DOI: 10.48550/arXiv.2405.08019, Publication Year: 2024)
Lai et al. (“InstantPortrait: One-Step Portrait Editing via Diffusion Multi-Objective Distillation”, Publication Year: 2025)
Any inquiry concerning this communication or earlier communications from the examiner should be directed to PROMOTTO TAJRIAN ISLAM whose telephone number is (703)756-5584. The examiner can normally be reached Monday - Friday 8:30 am - 5:00 pm EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Chan Park can be reached at (571) 272-7409. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/PROMOTTO TAJRIAN ISLAM/Examiner, Art Unit 2669 /CHAN S PARK/Supervisory Patent Examiner, Art Unit 2669