Prosecution Insights
Last updated: August 17, 2026
Application No. 17/862,779

APPARATUS AND METHOD WITH NEURAL NETWORK TRAINING BASED ON KNOWLEDGE DISTILLATION

Non-Final OA §103§112
Filed
Jul 12, 2022
Priority
Nov 01, 2021 — RE 10-2021-0148167 +1 more
Examiner
JABLON, ASHER H.
Art Unit
2127
Tech Center
2100 — Computer Architecture & Software
Assignee
Seoul National University R&DB Foundation
OA Round
3 (Non-Final)
43%
Grant Probability
Moderate
3-4
OA Rounds
3m
Est. Remaining
87%
With Interview

Examiner Intelligence

Grants 43% of resolved cases
43%
Career Allowance Rate
40 granted / 94 resolved
-12.4% vs TC avg
Strong +44% interview lift
Without
With
+44.5%
Interview Lift
resolved cases with interview
Typical timeline
4y 4m
Avg Prosecution
24 currently pending
Career history
121
Total Applications
across all art units

Statute-Specific Performance

§101
25.0%
-15.0% vs TC avg
§103
37.2%
-2.8% vs TC avg
§102
9.6%
-30.4% vs TC avg
§112
26.1%
-13.9% vs TC avg
Black line = Tech Center average estimate • Based on career data from 94 resolved cases

Office Action

§103 §112
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Continued Examination Under 37 CFR 1.114 A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 06/16/2026 has been entered. Status of the Claims Claims 1, 4-7, 11-15, 17, and 19-23 have been amended. Claims 2-3, 9, and 18 have been cancelled. Claims 1, 4-8, 10-17, and 19-23 are currently pending and have been considered by the Examiner. Claim Objections Claim 5 is objected to because of the following informalities: In claim 5, line 2, “second parameters” should recite “the second parameters”. Appropriate correction is required. Claim Rejections - 35 USC § 112 The following is a quotation of the first paragraph of 35 U.S.C. 112(a): (a) IN GENERAL.—The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor or joint inventor of carrying out the invention. The following is a quotation of the first paragraph of pre-AIA 35 U.S.C. 112: The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor of carrying out his invention. Claims 14-15 are rejected under 35 U.S.C. 112(a) or 35 U.S.C. 112 (pre-AIA ), first paragraph, as failing to comply with the written description requirement. The claim(s) contains subject matter which was not described in the specification in such a way as to reasonably convey to one skilled in the relevant art that the inventor or a joint inventor, or for applications subject to pre-AIA 35 U.S.C. 112, the inventor(s), at the time the application was filed, had possession of the claimed invention. Claim 14 recites the limitation "the another teacher network" in line 7. The written disclosure recites no more than one teacher network. Claim 15 is rejected for failing to cure the deficiencies of claim 14. The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 14-15 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. In claim 14, the limitations in lines 4-7 render the claim indefinite. There is insufficient antecedent basis for the limitation “the teacher network result of the another teacher network provided with the network input” in the claim because another teacher network provided with the network input lacks sufficient antecedent basis. The semicolon at the end of line 7 has been deleted in the filed claims, and it is unclear if the end of line 7 and the start of line 8 comprise the same sentence (i.e., “… the network input training the first …”). It is unclear if the limitation “the teacher network result of the another teacher network provided with the network input” should instead recite “another teacher network result of the teacher network provided with the network input;”. Examiner treats claim 14 as if it had recited this limitation. Claim 15 is rejected for failing to cure the deficiencies of claim 14. Claim 15 recites the limitation "the trained model parameters" in line 3. There is insufficient antecedent basis for this limitation in the claim. Parent claim 11 recites training first parameters of the energy-based model on page 5, line 5 and training second parameters of the implemented student network in line 10. It is unclear if “the trained model parameters” refers to either of these parameters. In claim 15, Examiner treats “the trained model parameters” as “the trained first parameters”. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention. Claims 1, 4, 7, 10, and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Ahn et al. (“Variational Information Distillation for Knowledge Transfer”, cited in IDS filed 07/12/2022) in view of Strauss et al. (“Arbitrary Conditional Distributions with Energy”, cited in PTO-892 issued 04/16/2026) and Li et al. (US 20220004803 A1, cited in PTO-892 issued 12/03/2025). Regarding claim 1, Ahn teaches: A processor-implemented method, the method comprising: applying an input to at least one image t and s. The sentence on P. 9164, end of col. 2, second line below equation 1 to P. 9165, col. 1, line 16, and lines 5-11 below equation 3 discloses the limitations. A student network result is s, a teacher network result is t, and a conditional distribution is q(t|s).) training first parameters of the [objective function] (k)|s(k)) in equation 4.) wherein the first value is associated with a difference between mutual information of the implemented teacher network and the implemented student network, and a variational lower bound of the mutual information; (Page 9164, col. 2, second line above equation 1 to equation 1 discloses “I” represents mutual information, and page 9615, col. 1, line 4 to the fourth line below equation 3 discloses minimizing a loss function which includes equation 3. In the instant specification, equation 1 in paragraph [0062] is identical to Ahn’s equation 3. Instant specification paragraph [0071] states that DKL is a difference between actual mutual information and a variational lower bound.) training second parameters of the implemented student network to increase a second value based on the conditional distribution associated with the variational lower bound; and (Page 9163, caption for Fig. 1; Page 9164, col. 1, in the paragraph that starts with “Evidently”, lines 4-14 in this paragraph; and Page 9165, col. 1, lines 4-10 discloses training the student network. Page 9165, below equation 4, lines 5-9 teaches training the student network to maximize conditional likelihood, which is the “second value based on the conditional distribution” as claimed.) outputting [a label] However, Ahn does not explicitly teach: applying an input to at least one image generation network, wherein the at least one image generation network comprises an implemented teacher network and an implemented student network; generating, based on a student network result of the implemented student network provided with the input and a teacher network result of the implemented teacher network provided with the input, a sample sampled from a conditional distribution of the teacher network result conditioned on the student network result, the conditional distribution being represented by an energy-based model; training first parameters of the energy-based model to decrease a first value based on the conditional distribution of the energy-based model by using the sample, training second parameters of the implemented student network to increase a second value based on the conditional distribution of the energy-based model by using the sample, outputting an image But Strauss teaches: generating, PNG media_image1.png 51 112 media_image1.png Greyscale by an energy-based model. On page 5, the sentence above equation 4 states “a proposal distribution PNG media_image2.png 48 161 media_image2.png Greyscale which is similar to the target distribution.” On page 5, line 1 below equation 6 states PNG media_image3.png 52 576 media_image3.png Greyscale .) training first parameters of the energy-based model to decrease a first value based on the conditional distribution of the energy-based model by using the sample, (Page 3, § 3.2, lines 1-7 discloses learning consists of finding an energy function that outputs low energies for correct prediction values. A “first value” is the energy output from the energy-based model.) It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to have used Strauss’ energy-based model to represent Ahn’s conditional distribution q(t|s). It is noted that equation 2 in specification paragraph [0068] is identical to equation 3 in Strauss, page 4 when Strauss’ unobserved variable xui and observed variables xo are replaced with t and s, respectively. In the combination, sampling from q(t|s) represented by an energy-based model would result in “generating… a sample sampled from a conditional distribution of the teacher network result conditioned on the student network result” as claimed, training first parameters of the energy-based model to decrease a first value (i.e., its energy) would be based on the teacher network result and the student network result, and training second parameters of the implemented student network to increase a second value based on the conditional distribution of the energy-based model would be based on the sample. A motivation for the combination is that energy functions are naturally capable of representing non-smooth distributions with low-density regions or discontinuities. (Strauss, Page 3, § 3.2, final sentence) However, Ahn and Strauss do not explicitly teach: applying an input to at least one image generation network, wherein the at least one image generation network comprises an implemented teacher network and an implemented student network; outputting an image But Li teaches: applying an input to at least one image generation network, wherein the at least one image generation network comprises an implemented teacher network and an implemented student network; ([0037], lines 1-8 teaches inputting a zebra image to a student model and a teacher model.) outputting an image based on the implemented student network ([0037], lines 1-8 teaches outputting a horse image from the student model.) It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to have replaced Ahn’s student and teacher models with Li’s image generation student and teacher model in the combination of Ahn and Strauss. Ahn’s models classify an input image, and Li’s models transform an input image into a new image. A motivation for the combination is to apply Ahn and Strauss’ system to image generation models. Li discloses that practical consumer applications incorporating image-to-image translation tasks are desirable and popular. It is thus desirable to provide GANs-based models for use on typical user devices such as smartphones, tablets, etc. to meet user demands and enhance the user experience. (Li, [0004]-[0005]) Regarding claim 4, the combination of Ahn, Strauss, and Li teaches: The method of claim 1, Ahn teaches: wherein the training of the second parameters of the implemented student network comprises training the second parameters of the implemented student network to increase the second value of the [conditional likelihood] However, Ahn does not explicitly teach: increase the second value of the energy-based model based on the trained first parameters, based on the sample and the student network result. But Strauss teaches: PNG media_image3.png 52 576 media_image3.png Greyscale .) It would have been obvious to a person having ordinary skill in the art to have represented Ahn’s conditional likelihood with Strauss’ energy-based function. A motivation for the combination is that energy functions are naturally capable of representing non-smooth distributions with low-density regions or discontinuities. (Strauss, Page 3, § 3.2, final sentence) Regarding claim 7, the combination of Ahn, Strauss, and Li teaches: The method of claim 1, Ahn teaches: qθ(t|s) denotes an approximate distribution of a conditional distribution p(t|s), (Page 9165, below equation 2, lines 7-9) (t, s) denotes an input of the implemented teacher network and the implemented student network respectively, (Page 9164, col. 2, § 2, lines 17-24 discloses feedforwarding an input x through the teacher and student networks. Ahn’s input x corresponds to “(t, s)” as claimed.) the teacher network result and the student network result are obtained based on (t, s), and (Page 9164, col. 2, § 2, lines 17-24 discloses feedforwarding an input x through the teacher and student networks to obtain activations of the teacher and student layers.) However, Ahn and Li do not explicitly teach: wherein the energy-based model is represented by PNG media_image4.png 71 391 media_image4.png Greyscale , Eθ denotes an energy function parameterized by the student network, θ denotes a first parameter of the energy-based model, … Zθ denotes a partition function representing a sum of probabilities that each of inputs of the implemented student network is present. But Strauss teaches: wherein the energy-based model is represented by PNG media_image4.png 71 391 media_image4.png Greyscale , (Page 3, § 3.2, lines 1-7 and Page 4, § 4.2, lines 1-3 disclose equations for energy-based models having the same terms as the claimed equation.) Eθ denotes an energy function parameterized by [a neural] θ denotes a first parameter of the energy-based model, (Page 4, § 4.2, lines 1-3 discloses representing the energy function as a neural network. A neural network has parameters.) Zθ denotes a partition function representing a sum of probabilities that each of inputs It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to have applied Strauss’ energy-based model, with the input being the same as the input to Ahn’s student and teacher networks. A motivation for the combination is the same as the motivation given for claim 1. Regarding claim 10, the combination of Ahn, Strauss, and Li teaches: the method of claim 1. However, Ahn and Strauss do not explicitly teach: A non-transitory computer-readable storage medium storing instructions that, when executed by a processor, cause the processor to perform the method of claim 1. But Li teaches: A non-transitory computer-readable storage medium storing instructions that, when executed by a processor, cause the processor to perform (Li, [0081], lines 13-18) It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to have incorporated Li’s computer program product into the combination of Ahn, Strauss, and Li. A motivation for the combination is to perform the method on a client computing device in the real world. (Li, [0081], lines 13-18) Claim 17 recites an apparatus which implements the same features as the method of claim 1 and is therefore rejected for at least the same reasons. However, Ahn and Strauss do not explicitly teach: An apparatus comprising: one or more processors; and a memory configured to store instructions, wherein the one or more processors are configured to execute the instructions, which configures the one or more processors to perform: But Li teaches: An apparatus comprising: one or more processors; and a memory configured to store instructions, wherein the one or more processors are configured to execute the instructions, which configures the one or more processors to perform: (Li, [0081], lines 13-18) It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to have incorporated Li’s computer program product into the combination of Ahn, Strauss, and Li. A motivation for the combination is to perform the method on a client computing device in the real world. (Li, [0081], lines 13-18) Claims 5, 8, and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Ahn et al. (“Variational Information Distillation for Knowledge Transfer”, cited in IDS filed 07/12/2022) in view of Strauss et al. (“Arbitrary Conditional Distributions with Energy”, cited in PTO-892 issued 04/16/2026), Li et al. (US 20220004803 A1, cited in PTO-892 issued 12/03/2025), and Smith et al. (US 20210117842 A1, cited in PTO-892 issued 12/03/2025). Regarding claim 5, the combination of Ahn, Strauss, and Li teaches: The method of claim 1, Ahn teaches the training of second parameters of the implemented student network at page 9165, col. 1, line 4 to the second line below equation 4. Strauss teaches the training of the first parameters at Page 3, § 3.2, lines 1-7. However, Ahn, Strauss, and Li do not explicitly teach: wherein the training of the first parameters and the training of second parameters of the implemented student network are repeatedly performed. But Smith teaches: wherein the training of the first parameters… are repeatedly performed. ([0046], lines 1-7 and [0085], lines 12-17 and discloses iteratively updating the energy function of an energy-based model.) It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to have updated the energy function iteratively until convergence in the combination of Ahn, Strauss, and Li, and similarly it would have been obvious to a person having ordinary skill in the art before the effective filing date to have iteratively trained Li’s student model until convergence. A motivation for the combination is to generate models with high prediction accuracy. Regarding claim 8, the combination of Ahn, Strauss, and Li teaches: The method of claim 1, However, Ahn, Strauss, and Li not explicitly teach: wherein the generating of the sample comprises generating the sample based on a Markov chain Monte Carlo (MCMC) scheme. But Smith teaches: wherein the generating of the sample comprises generating the sample based on a Markov chain Monte Carlo (MCMC) scheme. ([0046], lines 1-7 and [0048] on page 4, right column, lines 1-3 discloses using a MCMC to generate samples from an energy-based model.) It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to have used an MCMC to generate samples from an energy-based model. A motivation for the combination is that the partition function Z cannot be computed exactly, and sampling methods can be used to approximate it. (Smith, [0048]) Claim 19 recites an apparatus which implements the same features as the method of claim 5 and is therefore rejected for at least the same reasons. Claims 6 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Ahn et al. (“Variational Information Distillation for Knowledge Transfer”, cited in IDS filed 07/12/2022) in view of Strauss et al. (“Arbitrary Conditional Distributions with Energy”, cited in PTO-892 issued 04/16/2026), Li et al. (US 20220004803 A1, cited in PTO-892 issued 12/03/2025), Smith et al. (US 20210117842 A1, cited in PTO-892 issued 12/03/2025), and Fan et al. (“Learning to Teach”, cited in PTO-892 issued 12/03/2025). Regarding claim 6, the combination of Ahn, Strauss, Li, and Smith teaches: The method of claim 5, wherein, while the training of the first parameters and the training of the second parameters are repeatedly performed, However, Ahn, Strauss, and Smith do not explicitly teach: the first parameters are trained based on another student network result of the trained student network provided with the input. But Li teaches: It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to have generated multiple image results using the trained student network in the combination of Ahn, Strauss, Li, and Smith. A motivation for the combination is that once a network has been trained, a user can invoke the network as many times as needed to generate predictions. However, Ahn, Strauss, Li, and Smith do not explicitly teach: the first parameters are trained based on another student network result of the trained student network But Fan teaches: the first parameters are trained based on another student network result (P. 5, lines 1-7 and 20-28, where the limitation “first parameters” corresponds to Fan’s teacher model parameters, and the limitation “another student network result” corresponds to the reward rt.) It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to have incorporated Fan’s interactive process between a teacher and a learner into the combination of Ahn, Strauss, Li, and Smith. The result is an interactive process between the energy-based model and the student model. A motivation for the combination is that modifying a teacher model based on feedback from a student model facilitates learning of the student model. (Fan, P. 2, lines 16-25) Claim 20 recites an apparatus which implements the same features as the method of claim 6 and is therefore rejected for at least the same reasons. Claims 11-12, 14-16, and 21-22 are rejected under 35 U.S.C. 103 as being unpatentable over Li et al. (US 20220004803 A1, cited in PTO-892 issued 12/03/2025) in view of Ahn et al. (“Variational Information Distillation for Knowledge Transfer”, cited in IDS filed 07/12/2022) and Strauss et al. (“Arbitrary Conditional Distributions with Energy”, cited in PTO-892 issued 04/16/2026). Regarding claim 11, Li teaches: A processor-implemented method, comprising: ([0078], lines 4-6) applying an input to at least one image generation network, wherein the at least one image generation network comprises an implemented teacher network and an implemented student network; and ([0037], lines 1-8 teaches inputting a zebra image to a student model and a teacher model.) outputting an image based on the implemented student network, ([0037], lines 1-8 teaches outputting a horse image from the student model.) wherein the at least one image generation network is trained by a knowledge distillation scheme using an [objective function] … the teacher network result being from the implemented teacher network, and the student network result being from the implemented student network, ([0037], lines 1-8 discloses each of the teacher model and student model outputs a horse image. The horse image from the teacher model is “the teacher network result” and the horse image from the student model is “the student network result”.) wherein the training of the at least one image generation network comprises: ([0039], [0041] discloses training a student generator using the objective function in Equation (2).) However, Li does not explicitly teach: a knowledge distillation scheme using an energy-based model, wherein the training of the at least one image generation network is based on a sample sampled from a conditional distribution of a teacher network result conditioned on a student network result, the conditional distribution being represented by the energy-based model, … training first parameters of the energy-based model to decrease a first value based on the conditional distribution of the energy-based model by using the sample, the teacher network result and the student network result, wherein the first value is associated with a difference between mutual information of the implemented teacher network and the implemented student network, and a variational lower bound of the mutual information; and training second parameters of the implemented student network to increase a second value based on the conditional distribution of the energy-based model by using the sample, the teacher network result and the student network result, wherein the second value is associated with the variational lower bound. But Ahn teaches: wherein the training of the at least one t and s. The sentence on P. 9164, end of col. 2, second line below equation 1 to P. 9165, col. 1, line 16, and lines 5-11 below equation 3 discloses the limitations. A student network result is s, a teacher network result is t, and a conditional distribution is q(t|s). Training a network means training the student network, a teacher network result is t, a student network result is s, and a conditional distribution is q(t|s). It is noted that equation 1 in the specification paragraph [0062] is the same as equation 3 in Ahn, page 9165, and variables are defined below equation 1 on page 9164.) … training first parameters of the [objective function] (k)|s(k)) in equation 4.) wherein the first value is associated with a difference between mutual information of the implemented teacher network and the implemented student network, and a variational lower bound of the mutual information; and (Page 9164, col. 2, second line above equation 1 to equation 1 discloses “I” represents mutual information, and page 9615, col. 1, line 4 to the fourth line below equation 3 discloses minimizing a loss function which includes equation 3. In the instant specification, equation 1 in paragraph [0062] is identical to Ahn’s equation 3. Instant specification paragraph [0071] states that DKL is a difference between actual mutual information and a variational lower bound.) training second parameters of the implemented student network to increase a second value based on the conditional distribution It would have been obvious to a person having ordinary skill in the art to have trained the student neural network according to Ahn’s technique. A motivation for the combination is that Ahn’s framework proposes an actionable objective for knowledge transfer and allows one to quantify the amount of information that is transferred from a teacher network to a student network. (Ahn, P. 9164, col. 1, in the paragraph that starts with “Evidently”, lines 4-11 in this paragraph) However, Li and Ahn do not explicitly teach: a knowledge distillation scheme using an energy-based model, a sample sampled from a conditional distribution of a teacher network result conditioned on a student network result, the conditional distribution being represented by the energy-based model, training first parameters of the energy-based model to decrease a first value based on the conditional distribution of the energy-based model by using the sample, training second parameters of the implemented student network to increase a second value based on the conditional distribution of the energy-based model by using the sample, But Strauss teaches: “using an energy-based model” and “a sample sampled from a conditional distribution… the conditional distribution being represented by an energy-based model,” (Page 3, § 3.2, lines 1-7; Page 3, § 4, lines 1-6; and Page 4, § 4.2, first sentence in lines 1-3 discloses representing a conditional distribution PNG media_image1.png 51 112 media_image1.png Greyscale by an energy-based model. On page 5, the sentence above equation 4 states “a proposal distribution PNG media_image5.png 50 165 media_image5.png Greyscale which is similar to the target distribution.” On page 5, line 1 below equation 6 states PNG media_image3.png 52 576 media_image3.png Greyscale .) training first parameters of the energy-based model to decrease a first value based on the conditional distribution of the energy-based model by using the sample, (Page 3, § 3.2, lines 1-7 discloses learning consists of finding an energy function that outputs low energies for correct prediction values. A “first value” is the energy output from the energy-based model.) It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to have used Strauss’ energy-based model to represent Ahn’s conditional distribution q(t|s). It is noted that equation 2 in specification paragraph [0068] is identical to equation 3 in Strauss, page 4 when Strauss’ unobserved variable xui and observed variables xo are replaced with t and s, respectively. In the combination, sampling from q(t|s) represented by an energy-based model would result in “a knowledge distillation scheme using an energy-based model” as claimed, training first parameters of the energy-based model to decrease a first value (i.e., its energy) would be based on the teacher network result and the student network result, and training second parameters of the implemented student network to increase a second value based on the conditional distribution of the energy-based model would be based on the sample. A motivation for the combination is that energy functions are naturally capable of representing non-smooth distributions with low-density regions or discontinuities. (Strauss, Page 3, § 3.2, final sentence) Regarding claim 12, the combination of Li, Ahn, and Strauss teaches: The method of claim 11, Li teaches: wherein the implemented student network of the at least one image generation network comprises: a first type student network trained by a first knowledge distillation scheme using the energy-based model; and/or a second type student network trained by a second knowledge distillation scheme However, Li and Strauss do not explicitly teach: using a Gaussian distribution. But Ahn teaches: a second knowledge distillation scheme using a Gaussian distribution. (Page 9165, § 2.1, lines 1-6) A motivation for the combination is the same as the motivation given for claim 11. Regarding claim 14, the combination of Li, Ahn, and Strauss teaches: The method of claim 11, Li teaches: wherein the implemented student network of the at least one image generation network is trained by: ([0037], lines 1-8 and [0039]) generating, based on another student network result of the implemented student network provided with a network input, another sample corresponding to [an objective function] s(x) and the teacher network result Gt(x) when both networks are provided with an input “x”. The feature of “another sample” is another objective function value.) training the first parameters of the [objective function] training the second parameters of the implemented student network to increase the second value of the [objective function] However, Li does not explicitly teach: generating, based on another student network result of the implemented student network provided with a network input, another sample corresponding to a distribution of the energy-based model training the first parameters of the energy-based model to decrease the first value of the energy-based model, training the second parameters of the implemented student network to increase the second value of the energy-based model, But Strauss teaches: generating… another sample corresponding to a distribution of the energy-based model (Page 4, § 4.2, first sentence in lines 1-3 discloses representing a conditional distribution PNG media_image1.png 51 112 media_image1.png Greyscale by an energy-based model. On page 5, the sentence above equation 4 states “a proposal distribution PNG media_image5.png 50 165 media_image5.png Greyscale which is similar to the target distribution.” On page 5, line 1 below equation 6 states PNG media_image3.png 52 576 media_image3.png Greyscale .) training the first parameters of the energy-based model to decrease the first value of the energy-based model, (Page 3, § 3.2, lines 1-7 discloses learning consists of finding an energy function that outputs low energies for correct values.) training [to modify] It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to have generated another sample using Strauss’ energy-based model and to have trained the energy-based model in the combination of Li, Ahn, and Strauss. In the instant application, specification paragraph [0059] discloses the student network may be trained such that a pair of the teacher output and the student output decreases an output of an energy-based function. In other words, the energy-based function acts as an objective function to minimize a difference between the student network result and teacher network result. A motivation for the combination is the same as the motivation given for claim 11. Regarding claim 15, the combination of Li, Ahn, and Strauss teaches: The method of claim 14, Li teaches: wherein the training of the implemented student network comprises training the implemented student network to increase the second value of the [objective function] However, Li and Ahn do not explicitly teach: the energy-based model based on the trained [first] model parameters But Strauss teaches: the energy-based model based on the trained [first] model parameters. (Page 3, § 3.2, lines 1-7 discloses learning consists of finding an energy function that outputs low energies for correct values. Page 5, § 4.3, lines 1-2 and the sentence starting above and including equation 8 teaches training.) In the combination of Li, Ahn, and Strauss, Strauss’ trained energy-based model would replace Li’s objective function, as explained in the rejection of claim 14 above. A motivation for the combination is the same as the motivation for claim 11. Regarding claim 16, the combination of Li, Ahn, and Strauss teaches: the method of claim 11. Li teaches: A non-transitory computer-readable storage medium storing instructions that, when executed by a processor, cause the processor to perform ([0081], lines 13-18 teaches a computer program product.) Claim 21 recites an apparatus which implements the same features as the method of claim 11 and is therefore rejected for at least the same reasons. Li teaches: An apparatus, comprising: one or more processors; and ([0081], line 16 discloses “a processing unit”.) a memory configured to store instructions, wherein the one or more processors are configured to execute the instructions, which configures the one or more processors to perform: (Li, [0081], lines 13-18) Claim 22 recites an apparatus which implements the same features as the method of claim 12 and is therefore rejected for at least the same reasons. Claims 13 and 23 are rejected under 35 U.S.C. 103 as being unpatentable over Li et al. (US 20220004803 A1, cited in PTO-892 issued 12/03/2025) in view of Ahn et al. (“Variational Information Distillation for Knowledge Transfer”, cited in IDS filed 07/12/2022), Strauss et al. (“Arbitrary Conditional Distributions with Energy”, cited in PTO-892 issued 04/16/2026), and Tseng et al. (US 20200302292 A1, cited in PTO-892 issued 12/03/2025). Regarding claim 13, the combination of Li, Ahn, and Strauss teaches: The method of claim 12, Li teaches: wherein each of the at least one image generation network is However, Li, Ahn, and Strauss do not explicitly teach: each of the at least one image generation network is determined to be one of the first type student network and the second type student network But Tseng teaches: each of the at least one lines 1-3, [0095] discloses determining whether each network is a first type network (e.g., high accuracy) or a second type network (e.g., low accuracy).) It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to have incorporated Tseng’s determining a type of network into the combination of Li, Ahn, and Strauss to determine whether a student network translates an image from a first domain to a second domain or vice-versa. A motivation for the combination is to determine whether a model should be used for a task. (Tseng, [0096]) Claim 23 recites an apparatus which implements the same features as the method of claim 13 and is therefore rejected for at least the same reasons. Response to Arguments The following are Examiner’s responses to Applicant’s arguments filed 06/16/2026. Applicant’s Arguments Under 35 U.S.C. 112: The cancellation of claim 9 and the amendments to claims 11-14 and 21-22 obviate the previous rejections. Examiner’s Response: Applicant’s arguments have been fully considered. The previous rejections under 35 U.S.C. 112 have been withdrawn. Applicant’s First Arguments Under 35 U.S.C. 103 (Pages 13-14): The Office's reliance on Ahn for the recited difference therefore conflates Ahn's Gaussian-based optimization, which is carried out by training the student under a fixed-form distribution, with the distinct operation recited in claim 1 of training parameters of an energy-based model to decrease the first value… Strauss thus discloses neither the opposed training directions recited in claim 1 nor the use of an energy-based model as a trainable variational distribution for estimating the mutual information between two networks… Because Ahn supplies the mutual-information objective only in connection with a fixed Gaussian variational distribution and the training of the student, while Strauss trains an energy- based model only to fit the likelihood of a single data vector, the proposed combination does not arrive at the claimed operation of training first parameters of the energy-based model to decrease a value associated with the difference between mutual information and a variational lower bound. Arriving at that operation would require not merely substituting Strauss's energy- based model for Ahn's Gaussian variational distribution, but also repurposing Strauss's data- likelihood training objective into an objective that tightens a mutual-information bound between a teacher network and a student network - an objective that Strauss neither discloses nor suggests. Such a reconstruction is supported only by impermissible hindsight derived from the present application. See MPEP § 2142. Examiner’s Response: Applicant essentially argues the combination of Ahn and Strauss is improper due to impermissible hindsight reasoning, and because the combination allegedly does not make mathematical sense. Applicant's arguments have been fully considered but they are not persuasive. In the combination of at least Ahn and Strauss, Ahn’s conditional distribution of an objective function, namely the term log q(t(k)|s(k)) in equation Ahn’s equation 4, is replaced with Strauss’ energy-based model representing a conditional distribution p(xu|xo). In the instant specification, equation 1 in paragraph [0062] is identical to Ahn’s equation 3. Equation 2 in specification paragraph [0068] is identical to equation 3 in Strauss, page 4 when Strauss’ unobserved variable xui and observed variables xo are replaced with t and s, respectively. In response to applicant's argument that the examiner's conclusion of obviousness is based upon improper hindsight reasoning, it must be recognized that any judgment on obviousness is in a sense necessarily a reconstruction based upon hindsight reasoning. But so long as it takes into account only knowledge which was within the level of ordinary skill at the time the claimed invention was made, and does not include knowledge gleaned only from the applicant's disclosure, such a reconstruction is proper. See In re McLaughlin, 443 F.2d 1392, 170 USPQ 209 (CCPA 1971). In this case, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to have used Strauss’ energy-based model to represent Ahn’s conditional distribution q(t|s). It is noted that equation 2 in specification paragraph [0068] is identical to equation 3 in Strauss, page 4 when Strauss’ unobserved variable xui and observed variables xo are replaced with t and s, respectively. In the combination, sampling from q(t|s) represented by an energy-based model would result in “generating… a sample sampled from a conditional distribution of the teacher network result conditioned on the student network result” as claimed, training first parameters of the energy-based model to decrease a first value (i.e., its energy) would be based on the teacher network result and the student network result, and training second parameters of the implemented student network to increase a second value based on the conditional distribution of the energy-based model would be based on the sample. A motivation for the combination is that energy functions are naturally capable of representing non-smooth distributions with low-density regions or discontinuities. (Strauss, Page 3, § 3.2, final sentence) Claim 1 recites decreasing a first value that is associated with a difference involving mutual information and a variational lower bound of the mutual information, and increasing a second value that is associated with the variational lower bound. These features are recited at a high level of generality, and they lack any details explaining how they might accomplish tightening a mutual-information bound between a teacher network and a student network, as argued in the remarks. Applicant’s Second Arguments Under 35 U.S.C. 103 (Page 14): Applicants further note that the recited “sample sampled from a conditional distribution… represented by an energy-based model” cannot be read, as the Office suggests, as merely “a value… using an energy-based model or some other function.” The claim requires a sample that is sampled from a conditional distribution, and, consistent with the present specification, such a sample is obtained by drawing from the energy-based model, for example by way of a Markov chain Monte Carlo scheme. A deterministic scalar value computed by a loss or objective function is not a sample drawn from a distribution, and the broadest reasonable interpretation consistent with the specification does not encompass such a value. Examiner’s Response: Applicant's arguments have been fully considered but they are not persuasive. Claim 1 recites generating a sample at a high level of generality, and it lacks details explaining how the sample, the conditional distribution, the energy-based model, the student network, and the teacher network are all related to each other. The claim merely provides broad statements about what the generating step is based on and how the conditional distribution is represented. Therefore, it is reasonable to treat the limitation in claim 1, lines 5-9 as a value based on a student network result and a teacher network result using an energy-based model or some other function. With respect to the argument “A deterministic scalar value computed by a loss or objective function is not a sample drawn from a distribution,” it appears this argument comes from the remarks filed 02/24/2026, final paragraph on page 12, in response to a different combination of references and claim mapping. Applicant’s Third Arguments Under 35 U.S.C. 103 (Pages 14-16): On pages 14-15, the arguments explain differences between the present application and Ahn and Strauss. Applicant argues the combination of Ahn, Strauss, and Li do not teach at least the limitations of claim 1, from line 10 to the end of the claim. Examiner’s Response: Applicant's arguments have been fully considered but they are not persuasive for similar reasons given in the responses above. In the training steps in claim 1, relationships among a first value, the conditional distribution, and a difference between mutual information and a variational lower bound of the mutual information are recited at a high level of generality. Relationships among a second value, the conditional distribution, and the variational lower bound are also recited at a high level of generality. Therefore, the claims have been interpreted such that Strauss’ energy-based model can replace Ahn’s objective function, and Strauss’ energy function can represent a conditional distribution from which to sample a value. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to Asher H. Jablon whose telephone number is (571)270-7648. The examiner can normally be reached Monday - Friday, 9:00 am - 6:00 pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Abdullah Al Kawsar can be reached at (571)270-3169. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /A.H.J./Examiner, Art Unit 2127 /ABDULLAH AL KAWSAR/Supervisory Patent Examiner, Art Unit 2127
Read full office action

Prosecution Timeline

Show 2 earlier events
Feb 24, 2026
Response Filed
Mar 18, 2026
Applicant Interview (Telephonic)
Mar 18, 2026
Examiner Interview Summary
Apr 16, 2026
Final Rejection mailed — §103, §112
Jun 16, 2026
Request for Continued Examination
Jun 20, 2026
Response after Non-Final Action
Jun 30, 2026
Examiner Interview (Telephonic)
Jul 17, 2026
Non-Final Rejection mailed — §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12675727
METHOD AND SYSTEM FOR DETERMINING POLICIES, RULES, AND AGENT CHARACTERISTICS, FOR AUTOMATING AGENTS, AND PROTECTION
5y 10m to grant Granted Jul 07, 2026
Patent 12643559
NETWORK FOR DETECTING EDGE CASES FOR USE IN TRAINING AUTONOMOUS VEHICLE CONTROL SYSTEMS
1y 9m to grant Granted Jun 02, 2026
Patent 12626141
AUTOMATED GENERATION OF MACHINE LEARNING MODELS
3y 5m to grant Granted May 12, 2026
Patent 12614076
NEURAL NETWORK OPTIMIZATION DEVICE FOR EDGE DEVICE MEETING ON-DEMAND INSTRUCTION AND METHOD USING THE SAME
1y 9m to grant Granted Apr 28, 2026
Patent 12572794
SYSTEM AND METHOD FOR AUTOMATED OPTIMAZATION OF A NEURAL NETWORK MODEL
5y 4m to grant Granted Mar 10, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
43%
Grant Probability
87%
With Interview (+44.5%)
4y 4m (~3m remaining)
Median Time to Grant
High
PTA Risk
Based on 94 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month