Prosecution Insights
Last updated: October 01, 2026
Application No. 18/424,689

NOISE SCHEDULING FOR DIFFUSION NEURAL NETWORKS

Non-Final OA §101§103
Filed
Jan 26, 2024
Priority
Jan 26, 2023 — provisional 63/441,417
Examiner
ZENG, WENWEI
Art Unit
Tech Center
Assignee
Google LLC
OA Round
1 (Non-Final)
Grant Probability
Favorable
1-2
OA Rounds

Examiner Intelligence

Grants only 0% of cases
0%
Career Allowance Rate
0 granted / 0 resolved
-60.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
Avg Prosecution
23 currently pending
Career history
18
Total Applications
across all art units
This examiner has no resolved cases yet (career too new); statute-level performance unavailable. The Grant Probability card shows Tech Center averages instead.

Office Action

§101 §103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Information Disclosure Statement The information disclosure statement (IDS) submitted on January 16, 2025, was considered by the examiner. The submission is in compliance with the provisions of 37 CFR 1.97. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefore, subject to the conditions and requirements of this title. Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea (math concept) without significantly more. Claim 1: Regarding claim 1, in step 1 of the 101-analysis set forth in MPEP 2106, the claim recites “1. A method of training a diffusion neural network, the method comprising: obtaining a set of one or more training network outputs; for each training network output: sampling a time step by sampling from a time step distribution over time steps between a lower bound and an upper bound of the time step distribution; generating a new noise component; generating a new noisy network output by combining the training network output and the new noise component in accordance with a noise schedule that depends on the sampled time step and a scaling factor that is not equal to one; processing a new diffusion input comprising (i) the new noisy network output and (ii) data specifying the sampled time step using the diffusion neural network to generate a new diffusion output that defines an estimate of the new noise component for the sampled time step; and training the diffusion neural network on an objective that measures, for each training network output, an error between the estimate of the new noise component for the sampled time step generated by processing the new diffusion input comprising the new noisy network output generated from the training network output and the new noise component for the sampled time step,” and a method is one of the four statutory categories of invention. In step 2A prong 1 of the 101-analysis set forth in the MPEP 2106, the examiner has determined that the following limitations recite a process that, under the broadest reasonable interpretation, covers a math concept but for recitation of generic computer components: sampling a time step by sampling from a time step distribution over time steps between a lower bound and an upper bound of the time step distribution; This recites a mathematical relationship, mathematical formula or equation, or mathematical calculation, see in paragraphs [0069, 0070, 0073] from specification, state “The system samples a time step by sampling from a time step distribution over time steps between a lower bound and an upper bound of the time step distribution (step 204). For example, the time step distribution can be a continuous uniform distribution over the interval between zero and one, inclusive. That is, the time step has a value between zero and one, inclusive,” and see [0070] state “The system generates a new noise component (step 206). The noise component generally has the same dimensionality as the training network output but has noisy values. For example, the system can generate the new noise component by sampling each value in the new noise component from a specified noise distribution, e.g., a Normal distribution,” and see PNG media_image1.png 226 858 media_image1.png Greyscale , see MPEP 2106.04(a)(2), subsection I), generating a new noise component; (This recites a mathematical relationship, mathematical formula or equation, or mathematical calculation, see in paragraph [0073] from the specification similar to above limitation, see MPEP 2106.04(a)(2), subsection I), generating a new noisy network output by combining the training network output and the new noise component in accordance with a noise schedule that depends on the sampled time step and a scaling factor that is not equal to one; (This recites a mathematical relationship, mathematical formula or equation, or mathematical calculation, see in paragraph [0073] from the specification similar to above limitation, see MPEP 2106.04(a)(2), subsection I), processing a new diffusion input comprising (i) the new noisy network output and (ii) data specifying the sampled time step using the diffusion neural network to generate a new diffusion output that defines an estimate of the new noise component for the sampled time step, (This recites a mathematical relationship, mathematical formula or equation, or mathematical calculation, see in paragraph [0073] from the specification similar to above limitation, see MPEP 2106.04(a)(2), subsection I), If claim limitations, under their broadest reasonable interpretation, cover performance of the limitations as a math concept but for the recitation of generic computer components, then it falls within the math concept grouping of abstract ideas. Accordingly, the claim “recites” an abstract idea. In step 2A prong 2 of the 101-analysis set forth in MPEP 2106, the examiner has determined that the following additional elements do not integrate this judicial exception into a practical application: A method of training a diffusion neural network, the method comprising: obtaining a set of one or more training network outputs; … (In step 2A, prong 2, obtaining a set of outputs recites mere data gathering, which is considered insignificant extra-solution activity – see MPEP 2106.05(g)), for each training network output: … (Mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)), and training the diffusion neural network on an objective that measures, for each training network output, an error between the estimate of the new noise component for the sampled time step generated by processing the new diffusion input comprising the new noisy network output generated from the training network output and the new noise component for the sampled time step. (Mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)), Since the claim as a whole, looking at the additional elements individually and in combination, does not contain any other additional elements that are indicative of integration into a practical application, the claim is “directed” to an abstract idea. In step 2B of the 101-analysis set forth in the 2019 PEG, the examiner has determined that the claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above, additional elements vi and vii recite mere instructions to apply the judicial exception using generic computer components, which are not indicative of significantly more. The additional element v recites mere data gathering, and is considered an insignificant extra-solution activity. In step 2B, this insignificant extra-solution activity is a well understood routine and conventional activity, which includes receiving or transmitting data over a network from court case Symantec, 838 F.3d at 1321, 120 USPQ2d at 1362 (utilizing an intermediary computer to forward information); TLI Communications LLC v. AV Auto. LLC, 823 F.3d 607, 610, 118 USPQ2d 1744, 1745 (Fed. Cir. 2016), – see MPEP 2106.05(d) (II)(i)), Considering the additional elements individually and in combination, and the claim as a whole, the additional elements do not provide significantly more than the abstract idea. Therefore, the claim is not patent eligible. Claim 2: Regarding claim 2, it is dependent upon claim 1, and thereby incorporates the limitations of, and corresponding analysis applied to claim 1. Further, claim 2 recites the following additional element: 2. The method of claim 1, wherein processing a new diffusion input comprising (i) the new noisy network output and (ii) data specifying the sampled time step using the diffusion neural network to generate a new diffusion output that defines an estimate of the new noise component for the sampled time step comprises: (In step 2A, prong 2, this is considered mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)), (In step 2B, this is also considered mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)), normalizing the new noisy network output by a variance of the new noisy network output, (In step 2A, prong 2, this is considered mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)), (In step 2B, this is also considered mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)), Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible. Claim 3: Regarding claim 3, it is dependent upon claim 1, and thereby incorporates the limitations of, and corresponding analysis applied to claim 1. Further, claim 3 recites the following abstract idea: The method of claim 1, wherein generating a new noisy network output by combining the training network output and the new noise component in accordance with a noise schedule that depends on the sampled time step and a scaling factor that is not equal to one comprises generating a new noisy network output xt that satisfies: PNG media_image2.png 79 353 media_image2.png Greyscale where ϵ is the noise component, b is the scaling factor, γ(t) is the output of the noise schedule for the sampled time step t, and x0 is the training network input, (This recites a mathematical relationship, mathematical formula or equation, or mathematical calculation, see in paragraphs [0073, 0077] mention “[0073] For example, the system can combine the training network output and the new noise component as follows: PNG media_image2.png 79 353 media_image2.png Greyscale where ϵ is the noise component, b is the scaling factor, γ(t) is the noise level, i.e., the output of the noise schedule for the sampled time step t, and x0 is the training network input, and in [0077] stated “In particular, by reducing the scaling factor (to a number less than 1), the noise level within the noisy network output is increased. Thus, given that tasks at higher resolutions require higher noise levels, the scaling factor can be set to smaller values when output resolutions are higher. For example, the scaling factor can be set to one of 0.1, 2, 3, 4, 0.5, 0.6, 0.7, 0.8, or 0.9,” see MPEP 2106.04(a)(2), subsection I), If claim limitations, under their broadest reasonable interpretation, cover performance of the limitations as a mental process and math concept but for the recitation of generic computer components, then it falls within the mental process or math concept grouping of abstract ideas. Accordingly, the claim “recites” an abstract idea. Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible. Claim 4: Regarding claim 4, it is dependent upon claim 3, and thereby incorporates the limitations of, and corresponding analysis applied to claim 3. Further, claim 4 recites the following abstract idea: The method of claim 3, wherein γ(t)=1−t. (This recites a mathematical relationship, mathematical formula or equation, or mathematical calculation, see in paragraph [0080] note “ For example, the noise schedule can be a linear function of the sampled time step and, more specifically, γ(t)=1−t. Using this noise schedule can ensure that the training covers all noise levels during training, improving the performance of the diffusion neural network after training.” Also, see specification in paragraph [0120] for details, see MPEP 2106.04(a)(2), subsection I), If claim limitations, under their broadest reasonable interpretation, cover performance of the limitations as a mental process and math concept but for the recitation of generic computer components, then it falls within the mental process or math concept grouping of abstract ideas. Accordingly, the claim “recites” an abstract idea. Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible. Claim 5: Regarding claim 5, it is dependent upon claim 1, and thereby incorporates the limitations of, and corresponding analysis applied to claim 1. Further, claim 5 recites the following abstract idea: The method of claim 1, wherein the noise schedule is a cosine schedule or a sigmoid schedule. (This recites a mathematical relationship, mathematical formula or equation, or mathematical calculation, see in paragraphs [0081 and 0120] note [0081] stating “As another example, the noise schedule can be a cosine schedule or a sigmoid schedule,” and in [0120] mentioned an equation, see PNG media_image3.png 363 847 media_image3.png Greyscale , see MPEP 2106.04(a)(2), subsection I), If claim limitations, under their broadest reasonable interpretation, cover performance of the limitations as a mental process and math concept but for the recitation of generic computer components, then it falls within the mental process or math concept grouping of abstract ideas. Accordingly, the claim “recites” an abstract idea. Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible. Claim 6: Regarding claim 6, it is dependent upon claim 1, and thereby incorporates the limitations of, and corresponding analysis applied to claim 1. Further, claim 6 recites the following additional element: The method of claim 1, wherein the training network outputs are images, (In step 2A, prong 2, this recites an indication to a field of use or technological environment – see MPEP 2106.05(h)), (In step 2B, this also recites a field of use or technological environment – see MPEP 2106.05(h)), Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible. Claim 7: Regarding claim 7, it is dependent upon claim 1, and thereby incorporates the limitations of, and corresponding analysis applied to claim 1. Further, claim 7 recites the following additional element: The method of claim 1, wherein each training network output is associated with a conditioning input and wherein the new diffusion input comprises a representation of the conditioning input that is associated with the training network output, (In step 2A, prong 2, this is considered mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)), (In step 2B, this is also considered mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)), Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible. Claim 8: Regarding claim 8, it is dependent upon claim 7, and thereby incorporates the limitations of, and corresponding analysis applied to claim 7. Further, claim 8 recites the following additional elements: The method of claim 7, wherein the conditioning input is a text prompt, (In step 2A, prong 2, this recites an indication to a field of use or technological environment – see MPEP 2106.05(h)), (In step 2B, this also recites a field of use or technological environment – see MPEP 2106.05(h)), Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible. Claim 9: Regarding claim 9, it is dependent upon claim 1, and thereby incorporates the limitations of, and corresponding analysis applied to claim 1. Claim 9 recites the following abstract ideas: generating a final diffusion output for the iteration …, (This recites a mathematical relationship, mathematical formula or equation, or mathematical calculation, see in paragraphs in specification [0116, 0120], where in [0116] stated “For example, the system can set the final diffusion output equal to (1+w)*the first diffusion output−w*the additional diffusion output or, when there are multiple additional diffusion outputs, the sum of the additional diffusion outputs,” and in [0120] mentioned an equation, see the final diffusion output used in an equation PNG media_image3.png 363 847 media_image3.png Greyscale , see MPEP 2106.04(a)(2), subsection I), comprising processing a first diffusion input comprising a current network output as of the iteration using the diffusion neural network to generate a first diffusion output, the processing comprising normalizing the current network output; (This recites a mathematical relationship, mathematical formula or equation, or mathematical calculation, see in paragraphs in specification [0116, 0120] similar to above limitation, see MPEP 2106.04(a)(2), subsection I), Further, claim 9 recites the following additional elements: The method of claim 1, further comprising: after the training, using the trained diffusion neural network to generate a new network output, comprising, at each of a plurality of iterations: … (In step 2A, prong 2, this is considered mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)), (In step 2B, this is also considered mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)), and updating the current network output using the final diffusion output for the iteration, (In step 2A, prong 2, this is considered mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)), (In step 2B, this is also considered mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)), If claim limitations, under their broadest reasonable interpretation, cover performance of the limitations as a mental process and math concept but for the recitation of generic computer components, then it falls within the mental process or math concept grouping of abstract ideas. Accordingly, the claim “recites” an abstract idea. Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible. Claim 10: Regarding claim 10, it is dependent upon claim 9, and thereby incorporates the limitations of, and corresponding analysis applied to claim 9. Further, claim 10 recites the following additional element: The method of claim 9, wherein normalizing the current network output comprises normalizing the current network output based on a variance of the current network output, (In step 2A, prong 2, this is considered mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)), (In step 2B, this is also considered mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)), Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible. Claim 11: Regarding claim 11, in step 1 of the 101-analysis set forth in MPEP 2106, the claim recites “A system comprising: one or more computers; and one or more storage devices storing instructions that, when executed by the one or more computers, cause the one or more computers to perform operations for training a diffusion neural network, the operations comprising: obtaining a set of one or more training network outputs; for each training network output: sampling a time step by sampling from a time step distribution over time steps between a lower bound and an upper bound of the time step distribution; …” and a system is one of the four statutory categories of invention. In step 2A prong 1 of the 101-analysis set forth in the MPEP 2106, the examiner has determined that the following limitations recite a process that, under the broadest reasonable interpretation, covers a math concept but for recitation of generic computer components: sampling a time step by sampling from a time step distribution over time steps between a lower bound and an upper bound of the time step distribution; This recites a mathematical relationship, mathematical formula or equation, or mathematical calculation, see in paragraphs [0069, 0070, 0073] from specification, state “The system samples a time step by sampling from a time step distribution over time steps between a lower bound and an upper bound of the time step distribution (step 204). For example, the time step distribution can be a continuous uniform distribution over the interval between zero and one, inclusive. That is, the time step has a value between zero and one, inclusive,” and see [0070] state “The system generates a new noise component (step 206). The noise component generally has the same dimensionality as the training network output but has noisy values. For example, the system can generate the new noise component by sampling each value in the new noise component from a specified noise distribution, e.g., a Normal distribution,” and see PNG media_image1.png 226 858 media_image1.png Greyscale , see MPEP 2106.04(a)(2), subsection I), generating a new noise component; (This recites a mathematical relationship, mathematical formula or equation, or mathematical calculation, see in paragraph [0073] from the specification similar to above limitation, see MPEP 2106.04(a)(2), subsection I), generating a new noisy network output by combining the training network output and the new noise component in accordance with a noise schedule that depends on the sampled time step and a scaling factor that is not equal to one; (This recites a mathematical relationship, mathematical formula or equation, or mathematical calculation, see in paragraph [0073] from the specification similar to above limitation, see MPEP 2106.04(a)(2), subsection I), processing a new diffusion input comprising (i) the new noisy network output and (ii) data specifying the sampled time step using the diffusion neural network to generate a new diffusion output that defines an estimate of the new noise component for the sampled time step, (This recites a mathematical relationship, mathematical formula or equation, or mathematical calculation, see in paragraph [0073] from the specification similar to above limitation, see MPEP 2106.04(a)(2), subsection I), If claim limitations, under their broadest reasonable interpretation, cover performance of the limitations as a math concept but for the recitation of generic computer components, then it falls within the math concept grouping of abstract ideas. Accordingly, the claim “recites” an abstract idea. In step 2A prong 2 of the 101-analysis set forth in MPEP 2106, the examiner has determined that the following additional elements do not integrate this judicial exception into a practical application: A system comprising: one or more computers; and one or more storage devices storing instructions that, when executed by the one or more computers, cause the one or more computers to perform operations for training a diffusion neural network, the operations comprising... (In step 2A, prong 2, this is considered a generic computer component being used as a tool. – see MPEP 2106.05(f)), obtaining a set of one or more training network outputs; … (In step 2A, prong 2, obtaining a set of outputs recites mere data gathering, which is considered insignificant extra-solution activity – see MPEP 2106.05(g)), for each training network output: … (Mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)), and training the diffusion neural network on an objective that measures, for each training network output, an error between the estimate of the new noise component for the sampled time step generated by processing the new diffusion input comprising the new noisy network output generated from the training network output and the new noise component for the sampled time step. (Mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)), Since the claim as a whole, looking at the additional elements individually and in combination, does not contain any other additional elements that are indicative of integration into a practical application, the claim is “directed” to an abstract idea. In step 2B of the 101-analysis set forth in the 2019 PEG, the examiner has determined that the claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above, additional element v recites a generic computer component being used as a tool, and additional elements vii and viii recite mere instructions to apply the judicial exception using generic computer components, which are not indicative of significantly more. The additional element vi recites mere data gathering, and is considered an insignificant extra-solution activity. In step 2B, this insignificant extra-solution activity is a well understood routine and conventional activity, which includes receiving or transmitting data over a network from court case Symantec, 838 F.3d at 1321, 120 USPQ2d at 1362 (utilizing an intermediary computer to forward information); TLI Communications LLC v. AV Auto. LLC, 823 F.3d 607, 610, 118 USPQ2d 1744, 1745 (Fed. Cir. 2016), – see MPEP 2106.05(d) (II)(i)), Considering the additional elements individually and in combination, and the claim as a whole, the additional elements do not provide significantly more than the abstract idea. Therefore, the claim is not patent eligible. Claims 12-19: Regarding claims 12-19, claims 12-19 recite similar limitations as corresponding claims 2-9 listed above, and are rejected for similar reasons under 35 U.S.C. 101. Claim 20: Regarding claim 20, in step 1 of the 101-analysis set forth in MPEP 2106, the claim recites “One or more non-transitory computer storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations for training a diffusion neural network, the operations comprising: obtaining a set of one or more training network outputs; for each training network output: sampling a time step by sampling from a time step distribution over time steps between a lower bound and an upper bound of the time step distribution; generating a new noise component; generating a new noisy network output by combining the training network output and the new noise component in accordance with a noise schedule that depends on the sampled time step and a scaling factor that is not equal to one; processing a new diffusion input comprising (i) the new noisy network output and (ii) data specifying the sampled time step using the diffusion neural network to generate a new diffusion output that defines an estimate of the new noise component for the sampled time step …” and a non-transitory computer storage media recites a system and is one of the four statutory categories of invention. In step 2A prong 1 of the 101-analysis set forth in the MPEP 2106, the examiner has determined that the following limitations recite a process that, under the broadest reasonable interpretation, covers a math concept but for recitation of generic computer components: sampling a time step by sampling from a time step distribution over time steps between a lower bound and an upper bound of the time step distribution; This recites a mathematical relationship, mathematical formula or equation, or mathematical calculation, see in paragraphs [0069, 0070, 0073] from specification, state “The system samples a time step by sampling from a time step distribution over time steps between a lower bound and an upper bound of the time step distribution (step 204). For example, the time step distribution can be a continuous uniform distribution over the interval between zero and one, inclusive. That is, the time step has a value between zero and one, inclusive,” and see [0070] state “The system generates a new noise component (step 206). The noise component generally has the same dimensionality as the training network output but has noisy values. For example, the system can generate the new noise component by sampling each value in the new noise component from a specified noise distribution, e.g., a Normal distribution,” and see [0073] state PNG media_image1.png 226 858 media_image1.png Greyscale , see MPEP 2106.04(a)(2), subsection I), generating a new noise component; (This recites a mathematical relationship, mathematical formula or equation, or mathematical calculation, see in paragraph [0073] from the specification similar to above limitation, see MPEP 2106.04(a)(2), subsection I), generating a new noisy network output by combining the training network output and the new noise component in accordance with a noise schedule that depends on the sampled time step and a scaling factor that is not equal to one; (This recites a mathematical relationship, mathematical formula or equation, or mathematical calculation, see in paragraph [0073] from the specification similar to above limitation, see MPEP 2106.04(a)(2), subsection I), processing a new diffusion input comprising (i) the new noisy network output and (ii) data specifying the sampled time step using the diffusion neural network to generate a new diffusion output that defines an estimate of the new noise component for the sampled time step, (This recites a mathematical relationship, mathematical formula or equation, or mathematical calculation, see in paragraph [0073] from the specification similar to above limitation, see MPEP 2106.04(a)(2), subsection I), If claim limitations, under their broadest reasonable interpretation, cover performance of the limitations as a math concept but for the recitation of generic computer components, then it falls within the math concept grouping of abstract ideas. Accordingly, the claim “recites” an abstract idea. In step 2A prong 2 of the 101-analysis set forth in MPEP 2106, the examiner has determined that the following additional elements do not integrate this judicial exception into a practical application: One or more non-transitory computer storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations for training a diffusion neural network, the operations comprising … (In step 2A, prong 2, this is considered a generic computer component being used as a tool. – see MPEP 2106.05(f)), obtaining a set of one or more training network outputs; … (In step 2A, prong 2, obtaining a set of outputs recites mere data gathering, which is considered insignificant extra-solution activity – see MPEP 2106.05(g)), for each training network output: … (Mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)), and training the diffusion neural network on an objective that measures, for each training network output, an error between the estimate of the new noise component for the sampled time step generated by processing the new diffusion input comprising the new noisy network output generated from the training network output and the new noise component for the sampled time step. (Mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)), Since the claim as a whole, looking at the additional elements individually and in combination, does not contain any other additional elements that are indicative of integration into a practical application, the claim is “directed” to an abstract idea. In step 2B of the 101-analysis set forth in the 2019 PEG, the examiner has determined that the claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above, additional element v recites a generic computer component being used as a tool, and additional elements vii and viii recite mere instructions to apply the judicial exception using generic computer components, which are not indicative of significantly more. The additional element vi recites mere data gathering, and is considered an insignificant extra-solution activity. In step 2B, this insignificant extra-solution activity is a well understood routine and conventional activity, which includes receiving or transmitting data over a network from court case Symantec, 838 F.3d at 1321, 120 USPQ2d at 1362 (utilizing an intermediary computer to forward information); TLI Communications LLC v. AV Auto. LLC, 823 F.3d 607, 610, 118 USPQ2d 1744, 1745 (Fed. Cir. 2016), – see MPEP 2106.05(d) (II)(i)), Considering the additional elements individually and in combination, and the claim as a whole, the additional elements do not provide significantly more than the abstract idea. Therefore, the claim is not patent eligible. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claims 1, 2, 11, 12, and 20 are rejected under 35 U.S.C. 103 over Lin, Y. et al., in Pub. No. CN113822321A, published on December 21, 2021, (hereafter, Lin), in view of Ho, J. et al., in “Denoising diffusion probabilistic models,” published on December 6, 2020, cited in the January 16, 2025, IDS listed as item 27, (hereafter, Ho), further in view of Karras, T., et al. (2022). Elucidating the design space of diffusion-based generative models,” published on October 11, 2022, available in the January 16, 2025, IDS item 33, (hereafter, Karras). Claim 1: Regarding claim 1, Lin teaches “1. A method of training a diffusion neural network, the method comprising:” See Lin in [n0009] describe a “ training module for training a noise removal network and a noise scheduling network using each randomly selected training sample and the corresponding noise level, the noise removal network and the noise scheduling network being included in the generative model; wherein the noise removal network corresponds to the reverse process from a noisy input to a desired output, and the noise scheduling network corresponds to the forward process from the training samples from the training sample set to output a noisy output.” Here, Lin describes a training method of neural networks involved in noise scheduling and noise removal on image data. Also, see Lin in [n0003] describe "Generative models have been widely used in high-fidelity image generation, high-quality speech synthesis, natural language generation (1), and unsupervised representation learning, and have made great progress." Here, Lin applies using the model in generating images. Further, see Lin in [n0040] mention "For example, the Denoising Diffusion Implicit Model (DDIM) formulates a non Markov generative process that uses only a portion of the model in sample generation. By defining a prediction function to directly predict the observed variables of a given latent variable as sample outputs, samples can be generated from a subsequence of the entire inference trajectory of DDIM." Here, Lin shows that the model is a diffusion model that is involved in processing images. Further, Lin teaches “obtaining a set of one or more training network outputs; See Lin in [n0175] describe " the sample generation model uses a trained noise removal network to generate multiple inference samples based on random noise input, and outputs the final inference sample as a new sample." Here, Lin describes using a trained diffusion neural network to generate multiple inference samples (i.e. one or more training network outputs). Samples here is construed to mean a set of outputs of the neural network model from training. Also, see Lin in [n0003] describe "Generative models have been widely used in high-fidelity image generation, high-quality speech synthesis, natural language generation (1), and unsupervised representation learning, and have made great progress." Here, Lin applies using the model in generating image outputs. Further, Lin teaches “for each training network output: …” See Lin in paragraph [n0099] describe “Furthermore, since for each training sample, it is determined whether the parameters of the noise removal network and the noise scheduling network need to be updated based on that training sample, the training of the noise removal network and the noise scheduling network can also be considered as joint training.” Lin here mentions the noise scheduling and noise removing steps are applied for each training data of image data. See Lin in [n0003] describe from above limitation that this model can be applied to image data. See Lin in figures 1-3 for details. However, Lin did not teach “iii. sampling a time step by sampling from a time step distribution over time steps between a lower bound and an upper bound of the time step distribution;” or “generating a new noise component;” or “…generating a new noisy network output by combining the training network output and the new noise component in accordance with a noise schedule that depends on the sampled time step and a scaling factor that is not equal to one” or "processing a new diffusion input comprising (i) the new noisy network output and (ii) data specifying the sampled time step using the diffusion neural network to generate a new diffusion output that defines an estimate of the new noise component for the sampled time step; " or "and training the diffusion neural network on an objective that measures, for each training network output, an error between the estimate of the new noise component for the sampled time step generated by processing the new diffusion input comprising the new noisy network output generated from the training network output and the new noise component for the sampled time step." In an analogous art, Ho teaches “sampling a time step by sampling from a time step distribution over time steps between a lower bound and an upper bound of the time step distribution;” See Ho in page 4, algorithm 1, PNG media_image4.png 3 85 media_image4.png Greyscale where Ho mentions t ~ Uniform({1,…, T}), where a time step variable called t is sampled from a uniform time step distribution. Here, examiner construes lower bound to be a value in the lower side of a range of values, and upper bound to be a value in the maximum or highest value within a range of values. Here, Ho describes taking time steps from a lower value of 1, through an upper value of T as a range of the distribution. Examiner construes bound to mean a lower value and a high value within a range for any distribution. Further, Ho teaches “ generating a new noise component;” See Ho in page 4, algorithm 1 describe “ PNG media_image4.png 3 85 media_image4.png Greyscale PNG media_image4.png 3 85 media_image4.png Greyscale ”, where Ho in step 4 shows creating a new noise component, with every iteration or round of training, there is a new noise with a Gaussian distribution created. Since the algorithm has a step of repeat (see step 1) until converged, this algorithm runs generating new noise every round until the training process of the model has converged. Further, see Ho in page 4, equation 11, describe “where εΘ is a function approximator intended to predict ε from xt.” PNG media_image5.png 359 1071 media_image5.png Greyscale Here, Ho shows predicting ε or noise component from the function εΘ. Further, Ho teaches “generating a new noisy network output by combining the training network output and the new noise component in accordance with a noise schedule that depends on the sampled time step…” See Ho in page 3, equation 10 . PNG media_image6.png 826 1160 media_image6.png Greyscale Here, Ho mentions the variables x0, which relates to the output from training (i.e. training network output) and ε relates to the new noise component, and both are variables involved in equation 10 as part of (x0, ε), and this is involved in generating a new output for the model (i.e. new noisy network output). Also, see Ho in page 2, equation 4, in section 2. Background, describe “The forward process variance βt can be learned by reparameterization [33] or held constant as hyperparameters, and expressiveness of the reverse process is ensured in part by the choice of Gaussian conditionals in pθ(xt−1|xt), because both processes have the same functional form when βt are small [53]. A notable property of the forward process is that it admits sampling xt at an arbitrary timestep t in closed form: using the notation” PNG media_image7.png 207 1095 media_image7.png Greyscale Here, Ho describes a forward process distribution (also called the diffusion process) in a diffusion model, which directly adds noise to the original data of the training data output x0 to produce a noisy version xt at any given time step t. See details in algorithm 2, where Ho describes generating output of network x0. PNG media_image8.png 268 542 media_image8.png Greyscale See Ho in page 5, section 4, Experiments, describe “we set the forward process variances to constants increasing linearly from β1=10−4 to βT =0.02.” Here, Ho shows by using this linear beta schedule, this shows that amount of noise added changes depending on the timestep t that is sampled. Ho mentions that the noise variances (β) increase linearly from the first timestep (β₁ = 10⁻⁴) to the final timestep (βT = 0.02). This means the amount of noise added changes depending on which specific timestep t is being sampled. Further, Ho teaches “processing a new diffusion input comprising (i) the new noisy network output and (ii) data specifying the sampled time step using the diffusion neural network…” See Ho in page 4, algorithm 1, discuss an input of a new noise output and data of the sampled time step, in PNG media_image4.png 3 85 media_image4.png Greyscale , where εΘ shows the new noise output, and the t variable shows the sampled time step. Further, Ho teaches “… to generate a new diffusion output that defines an estimate of the new noise component for the sampled time step;” See Ho in page 4, section 3.2, paragraph after equation 12, mention “to summarize, we can train the reverse process mean function approximator µθ to predict˜µt, or by modifying its parameterization, we can train it to predict ε. Here, is where Ho mentions to predict noise ε, where ε is construed as a new diffusion output that is an estimate of the new noise component for the sampled time step. Further, Ho teaches “ and training the diffusion neural network on an objective that measures, for each training network output, an error between the estimate of the new noise component for the sampled time step generated by processing the new diffusion input comprising the new noisy network output generated from the training network output and the new noise component for the sampled time step” See Ho in page 5, first paragraph, describe “we found it beneficial to sample quality (and simpler to implement) to train on the following variant of the variational bound.” PNG media_image9.png 61 868 media_image9.png Greyscale Examiner construes limitation to mean train the model using error between predicted noise and actual noise. Here, Ho mentions in equation 14 training the model on an objective that measures the variable εΘ, an error of the new noise (i.e. predicted noise), and the variable ε, or noise for the sampled time step (i.e. actual noise). See Ho in equation 12 for details. The variable x0 relates to new model output (i.e. new noisy network output generated from the training network output) from the training process in algorithms 1 and 2 from page 4. Further, Ho describes the model takes in the unprocessed image with noise and the time step as inputs. Then, the model tries to guess the exact noise that was added. Ho mentions in equation 14 that ε - εΘ is the squared error between actual noise and model’s predicted noise, and the equation provides the loss function which calculates the squared error between the actual noise and the model's predicted noise. It would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to combine the base reference of Lin and incorporate into the teachings of Ho because both references teach training a model that generates a noisy network output that includes an error between the new noise estimate and the sampled time step. One of ordinary skill in the art would be motivated to do so because “efficient training is therefore possible by optimizing random terms of L with stochastic gradient descent,” (see Ho in page 3, section 2. Background, and equation 5 for details). However, Ho did not explicitly teach “ generating … a scaling factor that is not equal to one” In an analogous art, Karras teaches “ generating … a scaling factor that is not equal to one” See Karras in page 3, in table 1, PNG media_image10.png 3 63 media_image10.png Greyscale Here, Karras describes that s(t) takes the value of PNG media_image10.png 3 63 media_image10.png Greyscale shows this is a value that is larger than 1, and can be used as a scaling factor for a time step dependent function such as s(t). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the references of Lin and Ho, and incorporate with the teachings of Karras, by using the teachings from Lin and Ho, of training a model that generates a noisy network output that includes an error between the new noise estimate and the sampled time step, with the teaching of Karras of a scaling factor not equal to one. One of ordinary skill in the art would be motivated to do so because by integrating the framework of Karras into the methods of Lin and Ho, one with ordinary skill in the art would achieve the goal of providing a “design changes dramatically improve both the efficiency and quality obtainable with pre-trained score networks from previous work, including improving the FID of a previously trained ImageNet-64 model from 2.07 to near-SOTA 1.55, and after re-training with our proposed improvements to a new SOTA of 1.36,” (see Karras in page 1, abstract ). Claim 2: Regarding claim 2, Lin in view of Ho, further in view of Karras, teach the limitations of claim 1. Further, Karras teaches “2. The method of claim 1, wherein processing a new diffusion input comprising (i) the new noisy network output and (ii) data specifying the sampled time step using the diffusion neural network to generate a new diffusion output that defines an estimate of the new noise component for the sampled time step comprises: normalizing the new noisy network output by a variance of the new noisy network output” See Karras in page 8, section 5 Preconditioning and training, describe “Previous methods [36, 46, 48] address the input scaling via a σ-dependent normalization factor and attempt to precondition the output by training Fθ to predict n scaled to unit variance, from which the signal is then reconstructed via Dθ(x;σ) = x − σFθ(·). “ Here, Karras shows that methods use a sigma dependent factor for normalization, scaled to unit variance. Karras further explains from page 8, section 5, “needs to fine-tune its output carefully to cancel out the existing noise n exactly and give the output at the correct scale; note that any errors made by the network are amplified by a factor of σ. In this situation, it would seem much easier to predict the expected output D(x;σ) directly. In the same spirit as previous parameterizations that adaptively mix signal and noise (e.g., [10, 43, 50]), we propose to precondition the neural network with a σ-dependent skip connection”. Here, Karras shows that the normalization is performed on an output, which includes image data, to give the output at the correct scale. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the references of Lin and Ho, and incorporate with the teachings of Karras, by using the teachings from Lin and Ho, of training a model that generates a noisy network output that includes an error between the new noise estimate and the sampled time step, with the teaching of Karras of a scaling factor not equal to one. One of ordinary skill in the art would be motivated to do so because by integrating the framework of Karras into the methods of Lin and Ho, one with ordinary skill in the art would achieve the goal of providing a “design changes dramatically improve both the efficiency and quality obtainable with pre-trained score networks from previous work, including improving the FID of a previously trained ImageNet-64 model from 2.07 to near-SOTA 1.55, and after re-training with our proposed improvements to a new SOTA of 1.36,” (see Karras in page 1, abstract ). Claim 11: Regarding claim 11, Lin teaches “A system comprising: one or more computers; and one or more storage devices storing instructions that, when executed by the one or more computers, cause the one or more computers to perform operations for training a diffusion neural network, the operations comprising:…” See Lin in [n0234] describe “The computing device may be a server (including a cloud server) and/or a user terminal, or a computing device at each node in a blockchain system.” Also, see Lin [n0239-n0240] describe “according to another aspect of this application, a computer program product is also provided, comprising a computer program that, when executed by a processor, implements the steps of the training method as described with reference to Figures 3-4 and the steps of the generation method as described with reference to Figure 5. Computer programs can be stored in computer-readable storage media. The storage media mentioned above can be non-volatile storage media such as read-only memory, disks, or optical discs.” Here, Lin shows using a computer device which includes a computer-readable storage media to run programs such as running the steps of a training method for a model. Regarding claim 11, the claim recites similar limitations as corresponding independent claim 1 and is rejected for similar reasons as claim 1 using similar teachings and rationale. Claim 12: Regarding claim 12, the claim recites similar limitations as corresponding claim 2 and is rejected for similar reasons as claim 2 using similar teachings and rationale. Claim 20: Regarding claim 20, Lin teaches “One or more non-transitory computer storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations for training a diffusion neural network, the operations comprising:…” See Lin [n0239-n0240] describe “according to another aspect of this application, a computer program product is also provided, comprising a computer program that, when executed by a processor, implements the steps of the training method as described with reference to Figures 3-4 and the steps of the generation method as described with reference to Figure 5. Computer programs can be stored in computer-readable storage media. The storage media mentioned above can be non-volatile storage media such as read-only memory, disks, or optical discs.” Here, Lin shows using a computer device which includes a computer-readable storage media to run programs such as running the steps of a training method for a model. Also, see Lin in [n0236] describe “The memory can be non-volatile memory, such as read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory.” Here, Lin shows the memory is non-volatile which relates to a non-transitory storage that is part of an overall computer program. Regarding claim 20, the claim recites similar limitations as corresponding independent claim 1 and is rejected for similar reasons as claim 1 using similar teachings and rationale. Claims 3 and 13 are rejected under 35 U.S.C. 103 over Lin in view of Ho, further in view of Karras, and further in view of Nachmani, E., et al., in “Denoising diffusion gamma models,” published on October 10, 2021, available at: https://arxiv.org/pdf/2110.05948, (hereafter, Nachmani). Claim 3: Regarding claim 3, Lin in view of Ho, further in view of Karras, teach the limitations of claim 1. Further, Karras teaches “… and a scaling factor that is not equal to one” See Karras in page 3, in table 1, PNG media_image10.png 3 63 media_image10.png Greyscale Here, Karras describes that s(t) takes the value of PNG media_image10.png 3 63 media_image10.png Greyscale shows this is a value that is larger than 1, and can be used as a scaling factor for a time step dependent function such as s(t). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the references of Lin and Ho, and incorporate with the teachings of Karras, by using the teachings from Lin and Ho, of training a model that generates a noisy network output that includes an error between the new noise estimate and the sampled time step, with the teaching of Karras of a scaling factor not equal to one. One of ordinary skill in the art would be motivated to do so because by integrating the framework of Karras into the methods of Lin and Ho, one with ordinary skill in the art would achieve the goal of providing a “design changes dramatically improve both the efficiency and quality obtainable with pre-trained score networks from previous work, including improving the FID of a previously trained ImageNet-64 model from 2.07 to near-SOTA 1.55, and after re-training with our proposed improvements to a new SOTA of 1.36,” (see Karras in page 1, abstract ). However, Lin in view of Ho, further in view of Karras, did not teach “3. The method of claim 1, wherein generating a new noisy network output by combining the training network output and the new noise component in accordance with a noise schedule that depends on the sampled time step … comprises generating a new noisy network output xt that satisfies: PNG media_image11.png 45 308 media_image11.png Greyscale In an analogous art, Nachmani teaches ““3. The method of claim 1, wherein generating a new noisy network output by combining the training network output and the new noise component in accordance with a noise schedule that depends on the sampled time step … comprises generating a new noisy network output xt that satisfies: PNG media_image11.png 45 308 media_image11.png Greyscale ” See Nachmani in Page 3, in algorithm 1, mention “ PNG media_image12.png 29 386 media_image12.png Greyscale ”, and in page 4 where Nachmani mentions “The training procedure of εθ is defined in Alg.1. Given the input dataset d, the algorithm samples , x0 and t. The noisy latent state xt is calculated and fed to the DDPM neural network εθ. A gradient descent step is taken in order to estimate the ε noise with the DDPM network εθ.” Nachmani here describes a denoising diffusion model used as part of the training process. The equation features in step 6 PNG media_image12.png 29 386 media_image12.png Greyscale , which relates to the equation in the limitation PNG media_image11.png 45 308 media_image11.png Greyscale where αt or alpha bar t, which is similar to the ƴ(t) or gamma(t) variable, and includes similar variables such as ε and the (square root of (1-gamma(t)). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the references of Lin, Ho, and Karras, and incorporate with the teachings of Nachmani, by using the teachings from Lin, Ho, and Karras, of training a model that generates a noisy network output that includes an error between the new noise estimate and the sampled time step, with Nachmani’s teaching of the generating a new noisy output equation. One of ordinary skill in the art would be motivated to do so because by integrating Nachmani’s framework into the methods of Lin, Ho, and Karras, one with ordinary skill in the art would achieve using “the Denoising Diffusion Gamma Model (DDGM)… show that noise from Gamma distribution provides improved results for image and speech generation,” (see Nachmani in page 1, abstract). Claim 13: Regarding claim 13, the claim recites similar limitations as corresponding claim 3 and is rejected for similar reasons as claim 3 using similar teachings and rationale. Claims 4 and 14 are rejected under 35 U.S.C. 103 over Lin in view of Ho, further in view of Karras, and further in view of Nachmani, and further in view of Nichol, A. et al., in “Improved Denoising Diffusion Probabilistic Models,” published on February 18, 2021, cited in the January 16, 2025, IDS listed as item 41, (hereafter, Nichol). Claim 4: Regarding claim 4, Lin in view of Ho, further in view of Karras, and further in view of Nachmani, teach the limitations of claim 3. However, Lin in view of Ho, further in view of Karras, and further in view of Nachmani, did not teach “4. The method of claim 3, wherein γ(t)=1−t.” In an analogous field, Nichol teaches “4. The method of claim 3, wherein γ(t)=1−t” See Nichol in page 4, describe "Figure 3. Latent samples from linear (top) and cosine (bottom) schedules respectively at linearly spaced values of t from 0 to T. The latents in the last quarter of the linear schedule are almost purely noise, whereas the cosine schedule adds noise more slowly." Here, Nichol shows that the output of noise schedule for sampled time step t can use a linear noise schedule. Examiner construes γ(t)=1−t to mean a linear noise schedule as stated in the specification from paragraph [0080] “For example, the noise schedule can be a linear function of the sampled time step and, more specifically, γ(t)=1−t. Using this noise schedule can ensure that the training covers all noise levels during training, improving the performance of the diffusion neural network after training.” PNG media_image13.png 804 697 media_image13.png Greyscale It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the references of Lin, Ho, Karras, and Nachmani, and incorporate with the teachings of Nichol, by using the teachings from Lin, Ho, Karras, and Nachmani, of training a model that generates a noisy network output that includes an error between the new noise estimate and the sampled time step, with Nichol’s teaching of a linear noise schedule. One of ordinary skill in the art would be motivated to do so because by integrating Nichol’s framework into the methods of Lin, Ho, Karras, and Nachmani, one with ordinary skill in the art would achieve a method " found in early experiments that we could get a boost in log-likelihood by increasing T from 1000 to 4000; with this change, the log-likelihood improves to 3.77," (see Nichol in page 3, section 3, Improving the Log-likelihood). Claim 14: Regarding claim 14, the claim recites similar limitations as corresponding claim 4 and is rejected for similar reasons as claim 4 using similar teachings and rationale. Claims 5, 6, 15 and 16 are rejected under 35 U.S.C. 103 over Lin in view of Ho, further in view of Karras, and further in view of Nichol, A. et al., in “Improved Denoising Diffusion Probabilistic Models,” published on February 18, 2021, cited in the January 16, 2025, IDS listed as item 41, (hereafter, Nichol). Claim 5: Regarding claim 5, Lin in view of Ho, further in view of Karras, teach the limitations of claim 1. However, Lin in view of Ho, further in view of Karras, did not teach “5. The method of claim 1, wherein the noise schedule is a cosine schedule or a sigmoid schedule.” In an analogous art, Nichol teaches “5. The method of claim 1, wherein the noise schedule is a cosine schedule or a sigmoid schedule.” See Nichol in page 4, describe "Figure 3. Latent samples from linear (top) and cosine (bottom) schedules respectively at linearly spaced values of t from 0 to T. The latents in the last quarter of the linear schedule are almost purely noise, whereas the cosine schedule adds noise more slowly." Here, Nichol shows that the output of noise schedule for sampled time step t can use a cosine noise schedule. Since the claim limitation recites ‘or’, the examiner construes finding either a cosine or sigmoid schedule, and not both. PNG media_image13.png 804 697 media_image13.png Greyscale It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the references of Lin, Ho, Karras, and Nachmani, and incorporate with the teachings of Nichol, by using the teachings from Lin, Ho, Karras, and Nachmani, of training a model that generates a noisy network output that includes an error between the new noise estimate and the sampled time step, with Nichol’s teaching of a cosine noise schedule. One of ordinary skill in the art would be motivated to do so because by integrating Nichol’s framework into the methods of Lin, Ho, Karras, and Nachmani, one with ordinary skill in the art would achieve a method " found in early experiments that we could get a boost in log-likelihood by increasing T from 1000 to 4000; with this change, the log-likelihood improves to 3.77," (see Nichol in page 3, section 3, Improving the Log-likelihood). Claim 6: Regarding claim 6, Lin in view of Ho, further in view of Karras, teach the limitations of claim 1. However, Lin in view of Ho, further in view of Karras, did not teach “6. The method of claim 1, wherein the training network outputs are images.” In an analogous art, Nichol teaches “6. The method of claim 1, wherein the training network outputs are images” See Nichol in page 11, section A. Hyperparameters describe "For computing the reference distribution statistics we follow prior work (Ho et al., 2020; Brock et al., 2018) and use the full training set for CIFAR-10 and ImageNet, and 50K training samples for LSUN. Note that unconditional ImageNet 64×64 models are trained and evaluated using the official ImageNet-64 dataset ". Here, Nichol shows that the ImageNet dataset, which consist of image data, are trained and evaluated to create outputs. See Nichols in page 15 for details. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the references of Lin, Ho, Karras, and Nachmani, and incorporate with the teachings of Nichol, by using the teachings from Lin, Ho, Karras, and Nachmani, of training a model that generates a noisy network output that includes an error between the new noise estimate and the sampled time step, with Nichol’s teaching of images being used as training outputs. One of ordinary skill in the art would be motivated to do so because by integrating Nichol’s framework into the methods of Lin, Ho, Karras, and Nachmani, one with ordinary skill in the art would achieve a method " found in early experiments that we could get a boost in log-likelihood by increasing T from 1000 to 4000; with this change, the log-likelihood improves to 3.77," (see Nichol in page 3, section 3, Improving the Log-likelihood). Claim 15: Regarding claim 15, the claim recites similar limitations as corresponding claim 5 and is rejected for similar reasons as claim 5 using similar teachings and rationale. Claim 16: Regarding claim 16, the claim recites similar limitations as corresponding claim 6 and is rejected for similar reasons as claim 6 using similar teachings and rationale. Claims 7 and 17 are rejected under 35 U.S.C. 103 over Lin in view of Ho, further in view of Karras, and further in view of Rombach, R., et al., in “High-resolution image synthesis with latent diffusion models,” published between 18-24 June 2022 for a conference, available at: https://deepsense.ai/wp-content/uploads/2023/04/Rombach_High-Resolution_Image_Synthesis_With_Latent_Diffusion_Models_CVPR_2022_paper.pdf , (hereafter, Rombach). Claim 7: Regarding claim 7, Lin in view of Ho, further in view of Karras, teach the limitations of claim 1. However, Lin in view of Ho, further in view of Karras did not teach “7. The method of claim 1, wherein each training network output is associated with a conditioning input and wherein the new diffusion input comprises a representation of the conditioning input that is associated with the training network output.” In an analogous art, Rombach teaches “7. The method of claim 1, wherein each training network output is associated with a conditioning input and wherein the new diffusion input comprises a representation of the conditioning input that is associated with the training network output” See Rombach in page 10685, Introduction, in second to last paragraph on page, describe “(v) Moreover, we design a general-purpose conditioning mechanism based on cross-attention, enabling multi-modal training. We use it to train class-conditional, text-to-image and layout-to-image models.” Here, Rombach shows training models that generate image results using a text input as a condition (i.e. wherein each training network output is associated with a conditioning input). Further, see Rombach in page 10687, section 3.3. Conditioning Mechanisms describe “Similar to other types of generative models [55, 80], diffusion models are in principle capable of modeling conditional distributions of the form p(z|y). This can be implemented with a conditional denoising autoencoder ǫθ(zt, t, y) and paves the way to controlling the synthesis process through inputs y such as text [66], semantic maps [32,59] or other image-to-image translation tasks… we turn DMs into more flexible conditional image generators by augmenting their underlying UNet backbone with the cross-attention mechanism [94], which is effective for learning attention-based models of various input modalities. We introduce a domain specific encoder τθ that projects y to an intermediate representation τθ(y) ∈ RM× dτ, which is then mapped to the intermediate layers of the UNet via a cross-attention layer". Here, Rombach shows that using the ‘conditional denoising autoencoder’ of inputs y such as text and is then mapped to the intermediate layers of the UNet using a cross-attention layer, is a form of a representation of the conditioning input that connects to the training image data. Further, see Rombach in figure 4 describe training of the models were performed on images. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the references of Lin, Ho, and Karras, and incorporate with the teachings of Rombach, by using the teachings from Lin, Ho, and Karras, of training a model that generates a noisy network output that includes an error between the new noise estimate and the sampled time step, with Rombach’s teaching of each training network output is associated with a conditioning input. One of ordinary skill in the art would be motivated to do so because by integrating Rombach’s framework into the methods of Lin, Ho, and Karras, one with ordinary skill in the art would achieve “we turn DMs into more flexible conditional image generators by augmenting their underlying UNet backbone with the cross-attention mechanism [94], which is effective for learning attention-based models of various input modalities,” (see Rombach in page 10687, section 3.3. Conditioning Mechanisms). Claim 17: Regarding claim 17, the claim recites similar limitations as corresponding claim 7 and is rejected for similar reasons as claim 7 using similar teachings and rationale. Claims 8 and 18 are rejected under 35 U.S.C. 103 over Lin in view of Ho, further in view of Karras, and further in view of Rombach, and further in view of Liew, J. et al., in “Magicmix: Semantic mixing with diffusion models,” published on October 28, 2022, available at: https://arxiv.org/pdf/2210.16056 , (hereafter, Liew). Claim 8: Regarding claim 8, Lin in view of Ho, further in view of Karras, and further in view of Rombach, teach the limitations of claim 7. However, Lin in view of Ho, further in view of Karras, and further in view of Rombach, did not teach “8. The method of claim 7, wherein the conditioning input is a text prompt.” In an analogous art, Liew teaches “8. The method of claim 7, wherein the conditioning input is a text prompt” See Liew in page 1, abstract describe "Unlike style transfer where an image is stylized according to the reference style without changing the image content, semantic blending mixes two different concepts in a semantic manner to synthesize a novel concept ... our method first obtains a coarse layout (either by corrupting an image or denoising from a pure Gaussian noise given a text prompt), followed by injection of conditional prompt for semantic mixing." Here, Liew explicitly shows using a text prompt as an input to condition or train the model for processing images. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the references of Lin, Ho, Karras, and Rombach, and incorporate with the teachings of Liew, by using the teachings from Lin, Ho, Karras, and Rombach, of training a model that generates a noisy network output that includes an error between the new noise estimate and the sampled time step, with Liew’s teaching of a conditioning input is a text prompt. One of ordinary skill in the art would be motivated to do so because by integrating Liew’s framework into the methods of Lin, Ho, Karras, and Rombach, one with ordinary skill in the art would achieve “present MagicMix, a simple yet effective solution based on pre-trained text-conditioned diffusion models… our method first obtains a coarse layout (either by corrupting an image or denoising from a pure Gaussian noise given a text prompt), followed by injection of conditional prompt for semantic mixing. Our method does not require any spatial mask or re-training, yet is able to synthesize novel objects with high fidelity. To improve the mixing quality, we further devise two simple strategies to provide better control and flexibility over the synthesized content,” (see Liew in page 1, abstract). Claim 18: Regarding claim 18, the claim recites similar limitations as corresponding claim 8 and is rejected for similar reasons as claim 8 using similar teachings and rationale. Claims 9 and 19 are rejected under 35 U.S.C. 103 over Lin in view of Ho, further in view of Karras, and further in view of Chen, N. et al., in Pub. No. WO2022051548A1, published on March 10, 2022, (hereafter, Chen), and further in view of Wu, Y., et al., in “Group normalization,” published on June 11, 2018, available at https://openaccess.thecvf.com/content_ECCV_2018/papers/Yuxin_Wu_Group_Normalization_ECCV_2018_paper.pdf. Claim 9: Regarding claim 9, Lin in view of Ho, further in view of Karras, teach the limitations of claim 1. Further, Lin teaches “9. The method of claim 1, further comprising: after the training, using the trained diffusion neural network to generate a new network output, comprising, at each of a plurality of iterations:” See Lin in [n0012] describe "According to an embodiment of this application, the generated noise scheduling sequence can be used by a sample generation device to: generate multiple inference samples based on random noise input using a trained noise removal network, and output the final inference sample as a new sample." Here, Lin shows using a trained diffusion model (i.e. trained diffusion neural network) to produce an output as a new sample (i.e. new network output). Also, Lin elaborates at [n0063] that "The trained ∈θ(x<sub>t</sub>, α<sub>t</sub>) network can use Langevin dynamics to iteratively generate new samples based on the conditional distribution probability pθ(x<sub>t-1</sub>|x<sub>t</sub>) from random inputs." Here, Lin shows how the trained model can output new samples or results after training, and this is done iteratively or repeatedly (i.e. at each of a plurality of iterations). See also Lin in [n0040] mention "For example, the Denoising Diffusion Implicit Model (DDIM) formulates a non Markov generative process that uses only a portion of the model in sample generation. By defining a prediction function to directly predict the observed variables of a given latent variable as sample outputs, samples can be generated from a subsequence of the entire inference trajectory of DDIM." Lin here explicitly mentions using a diffusion model. To predict outputs. Further, see Lin in [n0104] note "For example, as mentioned above, since for the above α<sub>s</sub>, the corresponding α<sub>s</sub> can be obtained below as an example. After obtaining the noise level corresponding to the training sample, the training sample x<sub>0</sub> can generate an intermediate sample x<sub>s</sub>, which is based on α<sub>s</sub> and has been subjected to noise data.” Here, Lin describes that after training, the model generates an intermediate sample or outputs. See Lin in [n0065] and [n0153] for details. Further, Lin teaches “generating a final diffusion output … comprising processing a first diffusion input …” See Lin in [n0040] mention "For example, the Denoising Diffusion Implicit Model (DDIM) formulates a non Markov generative process that uses only a portion of the model in sample generation. By defining a prediction function to directly predict the observed variables of a given latent variable as sample outputs, samples can be generated from a subsequence of the entire inference trajectory of DDIM." Lin here explicitly mentions using a diffusion model, to predict outputs. Also, see Lin in [n0082] show “Figure 3 shows a schematic flowchart of a training method for generating a generative model to generate a desired output according to an embodiment of this application. The generative model can include a noise removal network and a noise scheduling network.” Lin here mentions the model contains a network model. Further, see Lin in [n0012] describe “According to an embodiment of this application, the generated noise scheduling sequence can be used by a sample generation device to: generate multiple inference samples based on random noise input using a trained noise removal network, and output the final inference sample as a new sample.” Here, Lin mentions generating final inference sample, where sample relates to final output data in the diffusion model (i.e. final diffusion output). Further, see Lin in [n0083] describe “As shown in Figure 3, in step S310, a training sample set is obtained, which includes multiple training samples, and the multiple training samples are independent and identically distributed samples.” Here, Lin shows that the model trains multiple samples or multiple image data, which includes a first, second or third, or subsequent image input (i.e. includes processing a first diffusion input). See Lin in figure 3 for details. PNG media_image14.png 570 707 media_image14.png Greyscale However, Lin in view of Ho, further in view of Karras, did not teach “generating a final diffusion output for the iteration, comprising processing a first diffusion input comprising a current network output as of the iteration using the diffusion neural network to generate a first diffusion output, the processing comprising normalizing the current network output;” or “and updating the current network output using the final diffusion output for the iteration.” In an analogous art, Chen teaches “generating a final diffusion output for the iteration, comprising processing a first diffusion input comprising a current network output as of the iteration using the diffusion neural network to generate a first diffusion output, …” See Chen describe in [0066] "At each of the multiple iterations, the system generates a noise output for the iteration by processing a model input including (1) the cu[rr]ent network output, (2) the network input, and optionally (3) iteration-specific data for the iteration (206) using a noise estimation neural network. The iteration-specific data is generally derived from noise levels for the iterations, where each noise level corresponds to a particular iteration. The noise output can include a noise estimate for each value in the current network output. For example, the respective noise estimate for a particular value in the current network output can represent an estimate of the noise that has been added to the corresponding actual value in an actual network output for the network input to generate the particular value." Here, Chen mentions taking a current network output, and network input as a model input (i.e. relates to first diffusion input), for the current iteration or round to generate a final noise output for that iteration (i.e. final output for the iteration). Also, see Chen in [0029] note "The described techniques, on the other hand, start from an initial network output, e.g., a noisy output that includes values sampled from a noise distribution, and iteratively refine the network output via a gradient-based sampler conditioned on the network input, i.e. an iterative denoising process may be used." Here, Chen describes using a network input to iteratively refine the network output, which uses the network model to process an input to refine the output. By iteratively, Chen shows this process is performed multiple times in repeat, and can process a first, second, third, or subsequent input. See Chen describe in [0074] "The noise estimation network 300 processes a model input including (1) a current network output 114, (2) a network input 102, and (3) iteration-specific data including aggregate noise level 306 corresponding to the cu[rr]ent iteration to generate a noise output 110. " Here, Chen shows the workflow of taking input and generating an output. Further, see Chen in [0043] note "As another example, the system can be configured to perform an image processing task on the network input to generate the network output.” Here, Chen mentions performing the method on image processing tasks. Further, Chen teaches “and updating the current network output using the final diffusion output for the iteration” See Chen in [0050] The system 100 then generates the final network output 104 by updating the current network output 114 at each of multiple iterations. In other words, the final network output 104 is the current network output 114 after the last iteration of the multiple iterations." Here, Chen explicitly shows updating the current network output after multiple iterations of running the model, and later creates the final output for the model. This shows a system that continuously updates the current output using the latest or recent final output for a specific iteration. Also, see Chen N. in [0064] describe “ The system initializes a current network output (204). For a final network output including multiple values, the system can sample each value in an initial current network output having the same number of values as the final network output from a noise distribution. For example, the system can initialize a current network output using a noise distribution (e.g., a Gaussian noise distribution), represented by yN ~ N(0, 1), where /is an identity matrix and the N in yN represents the intended number of iterations. The system can update the initial current network output over the N iterations, from iteration N to iteration 1, in descending order.” Here, Chen shows that the system repeatedly updates the initial current network output to indicate updating the current network output using the latest or recent current network output (i.e. viewed here as a final network output for the iteration). Further, see Chen in [0059-0060] describe “In particular, the update engine 112 updates the current network output 114 using the noise estimate and the corresponding noise level for the iteration. That is, the update engine 112 updates each value of the current network output 114 using the corresponding noise estimate of the noise output 110 and the corresponding noise level at the iteration, as is discussed in further detail with respect to FIG. 2. [0060] After the final iteration, the conditional output generation system 100 outputs the updated network output 114 as the final network output 104.” Here, Chen mentions that with each iteration, each output is considered to be a recent or final output of the last iteration. Final output is construed to mean the result from the most recent round of the model. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the references of Lin, Ho, and Karras, and incorporate with the teachings of Chen, by using the teachings from Lin, Ho, and Karras, of training a model that generates a noisy network output that includes an error between the new noise estimate and the sampled time step, with Chen’s teaching of updating the current network output. One of ordinary skill in the art would be motivated to do so because by integrating Chen’s framework into the methods of Lin, Ho, and Karras, one with ordinary skill in the art would achieve “The described techniques, on the other hand, start from an initial network output, e.g., a noisy output that includes values sampled from a noise distribution, and iteratively refine the network output via a gradient-based sampler conditioned on the network input, i.e. an iterative denoising process may be used. As a result, the approach is non-autoregressive and requires only a constant number of generation steps during inference. For example, for audio synthesis conditioned on a spectrogram, the described techniques can generate high fidelity audio samples in very few iterations, e.g., six or fewer, that compare to or even exceed those generated by state of the art autoregressive models with greatly reduced latency and while using many fewer computational resources. In addition, the described techniques can generate higher quality (e.g. higher fidelity) samples than those produced by existing non-autoregressive models,” (see Chen in [0029]). However, Lin in view of Ho, further in view of Karras, and further in view of Chen did not teach “… the processing comprising normalizing the current network output;” In an analogous art, Wu teaches “… the processing comprising normalizing the current network output;” See Wu in page 1, abstract describe “In this paper, we present Group Normalization (GN) as a simple alternative to BN. GN divides the channels into groups and computes within each group the mean and variance for normalization. GN’s computation is independent of batch sizes, and its accuracy is stable in a wide range of batch sizes.” Wu here shows a group normalization method that is used to run the normalization step within neural network layers or various groups of the model. This relates to the processing includes normalizing the current network output. Further, see Wu in page 10, first paragraph, Group division section, mention “we also evaluate fixing the number of channels per group (Table 3 ,left panel). Note that because the layers can have different channel numbers, the group number G can change across layers in this setting. In the extreme case of 1 channel per group ,GN is equivalent to IN. Even if using as few as 2channels per group, GN has substantially lower error than IN(25.6%vs. 28.4%). This result shows the effect of grouping channels when performing normalization.” Wu here shows using group normalization on various layers or output layers of a currently used network model. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the references of Lin, Ho, Karras, and Chen, and incorporate with the teachings of Wu, by using the teachings from Lin, Ho, Karras, and Chen, of training a model that generates a noisy network output that includes an error between the new noise estimate and the sampled time step, with Wu’s teaching of normalizing the current network output. One of ordinary skill in the art would be motivated to do so because by integrating Wu’s framework into the methods of Lin, Ho, Karras, and Chen, one with ordinary skill in the art would achieve "the improvement of GN on detection, segmentation, and video classification demonstrates that GN is a strong alternative to the powerful and currently dominant BN technique in these tasks," (see Wu in page 14, third paragraph of page, from section Results of 64-frame inputs). Claim 19: Regarding claim 19, the claim recites similar limitations as corresponding claim 9 and is rejected for similar reasons as claim 9 using similar teachings and rationale. Claim 10 is rejected under 35 U.S.C. 103 over Lin in view of Ho, further in view of Karras, further in view of Chen, further in view of Wu, and further in view of Lefkimmiatis, S., et al., in “Universal denoising networks: a novel CNN architecture for image denoising,” published for a conference from June 18-23, 2018, available at: https://ieeexplore.ieee.org/abstract/document/8578436 , (hereafter, Lefkimmiatis). Claim 10: Regarding claim 10, Lin in view of Ho, further in view of Karras, further in view of Chen, and further in view of Wu, teach the limitations of claim 9. However, Lin in view of Ho, further in view of Karras, further in view of Chen, and further in view of Wu, did not teach "10. The method of claim 9, wherein normalizing the current network output comprises normalizing the current network output based on a variance of the current network output." In an analogous art, Lefkimmiatis teaches "10. The method of claim 9, wherein normalizing the current network output comprises normalizing the current network output based on a variance of the current network output." See Lefkimmiatis in page 3207, section 2.3. Minimization strategy, describe “Here h1(y) can be interpreted as a noise estimator, which infers from the input the noise realization that distorts it. The noise realization estimate is further normalized to ensure that it has the correct variance and then it is subtracted from the noisy input. This leads to an output which consists of the latent image plus some residual noise, n−n1.” Here, Lefkimmiatis mentions that each noise estimate (i.e. viewed as current network output) is normalized to ensure it has a correct variance. Also, see Lefkimmiatis in page 3205, section 2, Image Restoration, describe “While the additive white Gaussian noise (AWGN) assumption is not frequently met in practice, an efficient solution of this problem is extremely valuable for two main reasons. The first one is that even in cases where the noise is signal dependent, there are several techniques available in the literature, such as variance stabilization transforms (VST) [1], [12], [30], which are able to transform the input data in a different domain so that the noise follows a Gaussian distribution with a fixed variance. Therefore, the solution can be obtained by first performing Gaussian denoising in the transform domain and then mapping the solution back to the original domain using the inverse VST.” Here, Lefkimmiatis shows ensuring data is transformed so every data point has a fixed variance, shows based on a variance. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the references of Lin, Ho, Karras, Chen, and Wu, and incorporate with the teachings of Lefkimmiatis, by using the teachings from Lin, Ho, Karras, Chen, and Wu, of training a model that generates a noisy network output that includes an error between the new noise estimate and the sampled time step, with the teaching of Lefkimmiatis of normalizing the current network output based on a variance of the current network output. One of ordinary skill in the art would be motivated to do so because by integrating the framework of Lefkimmiatis into the methods of Lin, Ho, Karras, Chen, and Wu, one with ordinary skill in the art would achieve "in the variational framework the choice of the regularizer has an important effect on the quality of the restored image. Equally important is our ability to efficiently compute the minimizer of the overall objective function," (see Lefkimmiatis, in page 3206, section 2.2. Constrained Optimization), and "a novel network architecture for learning discriminative image models that are employed to efficiently tackle the problem of grayscale and color image denoising," (see Lefkimmiatis, in page 3204, abstract). Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to WENWEI ZENG whose telephone number is (571)272-7111. The examiner can normally be reached Monday-Friday, 8am-5pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Usmaan Saeed can be reached at (571) 272-4046. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /WenWei Zeng/Examiner, Art Unit 2146 /USMAAN SAEED/Supervisory Patent Examiner, Art Unit 2146
Read full office action

Prosecution Timeline

Jan 26, 2024
Application Filed
Aug 24, 2026
Non-Final Rejection mailed — §101, §103 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
Grant Probability
Low
PTA Risk
Based on 0 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month