DETAILED ACTION
This final office action is responsive to application 18/235,273 with applicant’s amendments and request for reconsideration as submitted 27 July 2026.
Claim status is currently pending and under examination for claims 1, 3-8, 10-15 and 17-23 in which independent claims are 1, 8 and 15; amended claims are 1, 3-5, 8, 10-11, 15 and 17-18; cancelled claims are 2, 9 and 16; newly presented are claims 21-23.
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Remarks
Examiner thanks applicant for the responsive remarks filed 07/27/26 which are considered together with amendments on the remaining issues as follows.
Applicant’s remarks regarding the prior art have been considered, but they are moot in view of the new grounds of rejection as necessitated by applicant’s amendments. The amendments differ from the proposed amendments discussed during prior interview, though substantively include initialization such as shown Fig 1, and recite an additional neural network (i.e., in addition to first and second, in other words a third neural network). After updated search and consideration, prior art discovery reveals at least newly applied reference Xu as detailed below. Examiner respectfully submits that Xu’s disclosure provides a technique to capture the method as a whole, notwithstanding nominal processor circuit. Accordingly, Xu is found to anticipate the method which does not include processor circuitry, while the system and processor are found obvious over combination of Xu with Liu. The prior art rejections are updated as detailed below.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claims 15 and 17 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by:
Xu et al., “Knowledge distillation guided by multiple homogeneous teachers” hereinafter Xu.
With respect to claim 15, Xu teaches:
A method {Xu [P.232 Sect.3] “Method… we introduce the proposed MHT-KD method” Fig 1 illustrates, and implemented Alg.1 [P.235]}, comprising:
obtain one or more weights of a first neural network {Xu [P.234 Sect. 3.2.2] “wk is the weight corresponding to the kth teacher. W = [w1, w2, …, wk, …, wn] is a vector of n weights… the identified vector Wbest = [w1best, w2best, …, wkbest, …, wnbest] is selected, and its corresponding teacher network is considered the best teacher” such that select and/or identify is obtain, Fig 1 shows student-teacher networks which comprise [P.236 ¶1] “neural networks, namely ResNet20, ResNet32, and WideResNet-28-2” are known CNNs, also [P.233 ¶3,9] “pretraining teacher network… selected pretrained teacher” teacher network is a first neural network};
initialize one or more weights of a second neural network using a set of weights generated by an additional neural network based, at least in part, on the one or more weights of the first neural network {Xu [P.233 Sect. 2.3] “we design a confidence-adaptive initialization strategy to initialize the parameters of the student network based on the confidence of the teacher group” illustrated Fig 1 parameters including w-weights of student (second network, stage II) based on the multi-teacher networks (Teacher-1 is first network, Teacher-2 thru Teacher-n are additional networks of ensemble) in KD - knowledge distillation framework, implemented [P.235] Alg.1 Line 3 “Initialize a student network” from [P.234 Sect. 3.2.2] “weight corresponding to the kth teacher… The student directly inherits all of the network parameters to perform inheritance initialization” Eqs. 13-14. The student-teacher networks being [P.236 ¶1] “three types of neural networks, namely ResNet20, ResNet32, and WideResNet-28-2” are known CNNs}; and
update the one or more weights of the second neural network to match accuracy of the first neural network independently of outputs of the first neural network {Xu see [P.235] Alg.1 Line 15 “Update weights in NetS” subscript S is student network (second network) subject to Alg.1 Line 3 initialization, particularly the initialization’s confidence is an accuracy function per [P.234 Sect. 3.2.1] Eqs. 6-11 describes training and validation accuracies for teachers (first networks) with thresholds, matching may comprise thresholding so as for similarity loss Eq.17 [P.235] and/or minimized difference [P.233 ¶1-2] ”KD aims to minimize the difference between T and S” where “T and S represent the output of the teacher network and the student network… force pS to match pT” probabilities student & teacher. The networks are [P.231 ¶2] “trained independently” again at [P.232 Sect.3 ¶1] and denoted by subscript ind of LKDind – Loss, knowledge distillation (ind)ependent introduced Eq. 5, applied Eqs. 13-14 and combined Eq. 18. The total loss Eq.18 is calculated Alg.1 Line 14 prior to performing Alg.1 Line 15 NetS weight update}.
Claim 16 (Cancelled).
With respect to claim 17, Xu teaches the method of claim 15, further comprising:
cause the first neural network to generate one or more outputs {Xu [P.233 ¶1] “output of the teacher network” is first/teacher network to produce output by Eq.1, used [P.235] Alg.1 Line 10 “output PTi of NetTi”};
compare the one or more outputs with ground truth data {Xu [P.233 ¶1-2] “ground truth label” compared by Eq.3 where yi variable is ground truth and output is pT, “the output pT of the teacher model typically provides a higher probability value on the truth label” similar functions comprise Eqs. 4-5 and 13-14 relating said variables y and pT}; and
train the additional neural network using the comparison {Xu [P.235] Alg.1 Lines 9-16 are training epochs where Line 12 is Eq.14 in which the variables y and pT are applied to k-th teacher which is an additional network in the ensemble of teacher networks [P.234 Sect. 3.2.2], Fig 1-2. Eqs. 14, 13 and 5-4 are generalized from Eq.3 [P.233-34] loss functions for the training}.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 1, 3, 8, 10 and 14 are rejected under 35 U.S.C. 103 as unpatentable over: Xu (above) in view of Liu et al., PCT WO2024/040544A1 hereinafter Liu (Intel).
With respect to claim 1, Xu teaches:
obtain one or more weights of a first neural network {Xu [P.234 Sect. 3.2.2] “wk is the weight corresponding to the kth teacher. W = [w1, w2, …, wk, …, wn] is a vector of n weights… the identified vector Wbest = [w1best, w2best, …, wkbest, …, wnbest] is selected, and its corresponding teacher network is considered the best teacher” such that select and/or identify is obtain, Fig 1 shows student-teacher networks which comprise [P.236 ¶1] “neural networks, namely ResNet20, ResNet32, and WideResNet-28-2” are known CNNs, also [P.233 ¶3,9] “pretraining teacher network… selected pretrained teacher” teacher network is a first neural network};
initialize one or more weights of a second neural network using a set of weights generated by an additional neural network based, at least in part, on the one or more weights of the first neural network {Xu [P.233 Sect. 2.3] “we design a confidence-adaptive initialization strategy to initialize the parameters of the student network based on the confidence of the teacher group” illustrated Fig 1 parameters including w-weights of student (second network, stage II) based on the multi-teacher networks (Teacher-1 is first network, Teacher-2 thru Teacher-n are additional networks of ensemble) in KD - knowledge distillation framework, implemented [P.235] Alg.1 Line 3 “Initialize a student network” from [P.234 Sect. 3.2.2] “weight corresponding to the kth teacher… The student directly inherits all of the network parameters to perform inheritance initialization” Eqs. 13-14. The student-teacher networks being [P.236 ¶1] “three types of neural networks, namely ResNet20, ResNet32, and WideResNet-28-2” are known CNNs}; and
update the one or more weights of the second neural network to match accuracy of the first neural network independently of outputs of the first neural network {Xu see [P.235] Alg.1 Line 15 “Update weights in NetS” subscript S is student network (second network) subject to Alg.1 Line 3 initialization, particularly the initialization’s confidence is an accuracy function per [P.234 Sect. 3.2.1] Eqs. 6-11 describes training and validation accuracies for teachers (first networks) with thresholds, matching may comprise thresholding so as for similarity loss Eq.17 [P.235] and/or minimized difference [P.233 ¶1-2] ”KD aims to minimize the difference between T and S” where “T and S represent the output of the teacher network and the student network… force pS to match pT” probabilities student & teacher. The networks are [P.231 ¶2] “trained independently” again at [P.232 Sect.3 ¶1] and denoted by subscript ind of LKDind – Loss, knowledge distillation (ind)ependent introduced Eq. 5, applied Eqs. 13-14 and combined Eq. 18. The total loss Eq.18 is calculated Alg.1 Line 14 prior to performing Alg.1 Line 15 NetS weight update}.
However, Xu does not expressly disclose processor circuitry which is disclosed by Liu:
A processor, comprising: circuitry {Liu Fig 8:802 [0094] “processing device 802 may include… CPUs, GPUs” comprising [0093-95] “circuitry”. See e.g. [0127] “processor to perform operations including: inserting a first layer into the target neural network”. Additionally, Liu supports further functionality of matching loss minimization [0065] using feature distance [0064] between teacher and student networks Figs 4-5, with accuracy thresholding [0054-55], and weight update with initialization [0029]} to:
Liu is directed to neural network training thus being analogous. A person having ordinary skill in the art would have considered it obvious prior to the effective filing date to employ processor circuitry per Liu for Xu’s computer implementation as applying a known device to a known technique ready for improvement to yield predictable results such that a requisite computer environment reasonably provides “underlying hardware for executing machine learning applications” [0012] for example to “convert computationally intensive DNNs into more lightweight ones with similar accuracy. From a hardware perspective, this can facilitate replacement of a deep, sequential processing with parallel, distributed processing. This structural conversion can enable the acceleration of DNN training” [0017].
Claim 2 (Cancelled).
With respect to claim 3, the combination of Xu and Liu teaches the processor of claim 1, wherein the circuitry is further to:
cause the first neural network to generate one or more outputs {Xu [P.233 ¶1] “output of the teacher network” is first/teacher network to produce output by Eq.1, used [P.235] Alg.1 Line 10 “output PTi of NetTi”};
compare the one or more outputs with ground truth data {Xu [P.233 ¶1-2] “ground truth label” compared by Eq.3 where yi variable is ground truth and output is pT, “the output pT of the teacher model typically provides a higher probability value on the truth label” similar functions comprise Eqs. 4-5 and 13-14 relating said variables y and pT}; and
train the additional neural network using the comparison {Xu [P.235] Alg.1 Lines 9-16 are training epochs where Line 12 is Eq.14 in which the variables y and pT are applied to k-th teacher which is an additional network in the ensemble of teacher networks [P.234 Sect. 3.2.2], Fig 1-2. Eqs. 14, 13 and 5-4 are generalized from Eq.3 [P.233-34] loss functions for the training}.
With respect to claim 8, the rejection of claim 1 is incorporated. The difference in scope being a system comprising processor to perform limitations of claim 1. Liu discloses [0026] “DNN systems, methods and devices” with “processors” e.g. Fig 8:802. Motivation for combination is applied equally as in claim 1 and the remainder of this claim is rejected for the same rationale as claim 1.
Claim 9 (Cancelled).
With respect to claim 10, the combination of Xu and Liu teaches the system of claim 8, and further teaches the limitation of claim 3. Therefore, the rejection of claim 3 is applied to claim 10.
With respect to claim 14, the combination of Xu and Liu teaches the system of claim 8, wherein the one or more processors are to cause
the first neural network to adjust one or more weights after matching accuracy of the second neural network {Liu Figs 4-5 neural networks, [0029] “Weights can be initialized and updated” e.g. subject to performance criteria [0052] “The training module 250 may stop adjusting the parameters in the merged network after a threshold condition is met…. target performance (e.g., an accuracy)” the accuracy is matched upon matching loss being minimized [0065,64]}.
A person having ordinary skill in the art would have considered it obvious prior to the effective filing date to adjust/update weight after threshold accuracy based on matching loss per Liu in combination to arrive at the invention as claimed for a motivation “DNNs with better accuracy” [0017] such that it “provides a teacher-student training framework to train a compact, computationally efficient DNN model having improved prediction accuracy” [0013].
Claims 4 and 11 are rejected under 35 U.S.C. 103 as being unpatentable over: Xu and Liu in view of Xu et al., US PG Pub No 2023/0153615A1 hereinafter XuY (Huawei) and further in view of Yao et al., US PG Pub No 2025/0252318A1 hereinafter Yao (Intel).
With respect to claim 4, the combination of Xu and Liu teaches the processor of claim 1, wherein the circuitry is further to cause the additional neural networks to:
assign the one or more first weights to the first neural network {Xu [P.234 Sect. 3.2.2 ¶2] “Wbest are assigned” from “identified vector Wbest = [w1best, w2best, …, wkbest, …, wnbest] is selected, and its corresponding teacher network is considered the best teacher” where “wk is the weight corresponding to the kth teacher. W = [w1, w2, …, wk, …, wn] is a vector of n weights” illustrated Figs 1-2 ensemble of the teacher networks}.
However, the combination of Xu and Liu does not disclose that second weights of second neural network are received by an additional/third neural network, nor a weight “distribution” which is disclosed by XuY
receive one or more second weights of the second neural network {XuY Fig 9 arrows indicate receiving by second and third/additional neural networks, and discloses [0203-06] “input of the second neural network layer is X2 and the second weight is F2… h(X2,F2) may indicate the second target” emphasis second weights (F2) of the second neural network, cont’d [0237-38] “knowledge distillation… using the updated second neural network as the teacher model and the updated first neural network as the student model to obtain a third neural network”};
identify a weight distribution of the second neural network based, at least in part, on the one or more second weights {XuY [0024] “weight distribution of the second neural network is Gaussian” thus identified or obtained from a Gaussian distribution [0203-06]. See also [0197] “weight distribution of the second neural network. Specifically, there are hundreds or even tens of millions of parameters in the neural network…CNN”};
XuY is directed to neural network training thus being analogous. A person having ordinary skill in the art would have considered it obvious prior to the effective filing date to specify weight distribution of second neural network for a third neural network per XuY in combination for a motivation being it “eliminates network performance degradation caused by different weight distribution of the neural network layers during knowledge distillation” [0019] and such that “progressive distillation method can be used to enable the teacher model and the student model to learn together” [0040] further noting “obvious advantage in classification precision” [0240].
However, the combination Xu, Liu and XuY does not appear to disclose the following limitation that is met by Yao:
generate one or more first weights of the first neural network based, at least in part, on the weight distribution of the second neural network {Yao Figs 4-5 neural networks, weights detailed per [0069] Eq. where “p denotes probability distribution” (of weights/parameter), particularly “θS denotes parameters for the student network, θT denotes parameters for the teacher network), generating may employ bi-directional distillation between student and teacher model that uses backward gradient flow [0066,69] and/or [0031] “Weights can be initialized and updated by backpropagation” further discloses generative models e.g. RNN, GAN and LSTM [0071], Fig 7:710}; and
Yao is directed to neural network training thus being analogous. A person having ordinary skill in the art would have considered it obvious prior to the effective filing date to weight according to the solution of Yao in combination to arrive at the invention as claimed for a motivation “improved techniques for knowledge distillation… better accuracy-efficiency tradeoff” [0015,19] and “the need to design feature distillation losses and to tune weighting factors to balance loss terms can be avoided” [0052].
With respect to claim 11, the combination of Xu and Liu teaches the system of claim 8, and further combination with XuY and Yao teaches the limitations of claim 4. Therefore, the rejection of claim 4 with equal motivation is applied to claim 11.
Claims 5 and 12 are rejected under 35 U.S.C. 103 as being unpatentable over: Xu and Liu in view of Liu et al., “NORM: Knowledge Distillation via N-To-One Representation Matching” hereinafter Liu-NPL (arXiv: 2305.13803v1, ICLR conference paper as NPL version of Liu-PCT/WO cited in claim 1).
With respect to claim 5, the combination of Xu and Liu teaches the processor of claim 1. Liu-NPL teaches wherein the circuitry is further to
generate a learned transformation that projects a layer of the second neural network with a second dimensionality to a layer of the first neural network with a first dimensionality {Liu-NPL see [Abst] “Feature Transform (FT) module consisting of two linear layers… The first linear layer projects the student representation to a feature space having N times feature channels than the teacher” detailed [P.3 ¶5] “(FT) which for the teacher|student network, which projects Ft|Fs to the same feature space” notes channel dimension which is denoted with superscript of feature set memberships and described e.g. [P.4 ¶3] “dimension to project” illustratively Figs 1, 5 and generated per [P.5 ¶2] “independently initialized feature transforms to generate N channel-expanded views of the student feature”}.
Liu-NPL, same author as Liu-PCT/WO, is directed to training neural networks thus being analogous. A person having ordinary skill in the art would have considered it obvious prior to the effective filing date to perform the transformation that projects per Liu-NPL in combination as applying a known technique to a known device ready for improvement to yield predictable results and/or a motivation of a “simple FT module” for “enabling many-to-one feature mimicking” which “allows NORM to introduce many parallel knowledge transfer routes between a single teacher-student pair via simple feature splitting and group-wise feature mimicking” and “the FT module can be directly merged into its subsequent fully connected layer, without introducing any extra parameters or architectural modifications” [P.2 ¶2].
With respect to claim 12, the combination of Xu and Liu teaches the system of claim 8, and further combination with Liu-NPL teaches the limitations of claim 5. Therefore, the rejection of claim 5 with equal motivation is applied to claim 12.
Claims 6 and 13 are rejected under 35 U.S.C. 103 as being unpatentable over: Xu and Liu in view of Jacob et al., US PG Pub No 2024/0037930A1 hereinafter Jacob.
With respect to claim 6, the combination of Xu and Liu teaches the processor of claim 1. Jacob teaches wherein
the second neural network is a trained neural network capable of performing two or more tasks, and the first neural network is to be trained to perform one or more of the two or more tasks {Jacob teaches multi-task distillation, [0033-34] “distillation may include a training strategy for training single-task neural network models (STL) and a multi-task neural network framework (MTL) simultaneously on Nt tasks. The single-task neural network models may guide the optimization of the multi-task network throughout the training process. The multi-task network weights may be tied to the single-task neural network models through distillation loss… task-specific loss” Eq.1 details loss with STL+MTL shown Fig 1 and Fig 3. MTL corresponds to 2nd NN, STL corresponds to 1st NN. See also [0039] online task weighting}.
Jacob is directed to training neural networks thus being analogous. A person having ordinary skill in the art would have considered it obvious prior to the effective filing date to specify tasks for trained neural networks per Jacob in combination for a motivation being “multi-task learning is configured to exploit information in the training data of related tasks to learn a shared representation and improve generalization” [0002] e.g. “The single-task neural network models may guide the optimization of the multi-task network throughout the training process” [0033] and/or “By performing the adaptive feature distillation while training the multi-task neural network framework and/or training the multi-task neural network framework using the online weighting scheme, the reduced storage and increased speed advantages of multi-task learning are preserved along with an improvement in the performance of the multi-task neural network” [0032].
With respect to claim 13, the combination of Xu and Liu teaches the system of claim 8, and further combination with Jacob teaches the limitation of claim 6. Therefore, the rejection of claim 6 with equal motivation is applied to claim 13.
Claim 7 is rejected under 35 U.S.C. 103 as being unpatentable over: Xu and Liu in view of Sundaresan et al., US PG Pub No 2022/0036194A1 hereinafter Sundaresan.
With respect to claim 7, the combination of Xu and Liu teaches the processor of claim 1. Sundaresan teaches wherein
a number of weights in the second neural network is more than a number of weights in the first neural network {Sundaresan see [0026] “The teacher model is much larger than the student model (in terms of parameter space size)” e.g. [0036] “parameter budget 706 may specify that the subnet 201 should not have more than a certain number of parameters” Figs 1-2, 7-9 neural network distillation, note [0018] “supernet contains a smaller subnet that, when training in isolation, can match the accuracy (or other performance metrics) of the original ML model”}.
Sundaresan is directed to neural network training thus being analogous. A person having ordinary skill in the art would have considered it obvious prior to the effective filing date to specify number of weights/parameters (a.k.a. model size, complexity) per Sundaresan in combination to arrive at the invention as claimed for a motivation that large models perform more complex tasks with higher performance (accuracy or speed) at the expense of a tradeoff for additional computing resources, e.g. “reducing the number of parameters may decrease the performance of a model, but may allow the model to run faster and use less memory than it would with a larger number of parameters” [0004]. Additional motivation points to matching accuracy [0018].
Claim 18 is rejected under 35 U.S.C. 103 as being unpatentable over: Xu in view of XuY and Yao.
With respect to claim 18, Xu teaches the method of claim 15, and further combination with XuY and Yao teaches the limitations of claim 4. Therefore, the rejection of claim 4 with equal motivation is applied to claim 18.
Claim 19 is rejected under 35 U.S.C. 103 as being unpatentable over: Xu in view of Liu-NPL.
With respect to claim 19, Xu teaches the method of claim 15, and further combination with Liu-NPL teaches the limitations of claim 5. Claim 19 substitutes term map for project as equivalent terminology. Therefore, the rejection of claim 5 with equal motivation is applied to claim 19.
Claim 20 is rejected under 35 U.S.C. 103 as being unpatentable over: Xu in view of Jacob.
With respect to claim 20, Xu teaches the method of claim 15, and further combination with Jacob teaches the limitations of claim 6. Therefore, the rejection of claim 6 with equal motivation is applied to claim 20.
Claim 21 is rejected under 35 U.S.C. 103 as being unpatentable over: Xu and Liu in view of Xing et al., “Collaborative Consistent Knowledge Distillation Framework for Remote Sensing Image Scene Classification Network” hereinafter Xing.
With respect to claim 21, the combination of Xu and Liu teaches the processor of claim 1, wherein:
the first neural network is a pretrained neural network {Xu discloses [P.233 ¶3,9] “pretrained teacher network… selected pretrained teacher” Fig 1 shows teacher/first neural network}; and
However, the combination of Xu and Liu does not expressly disclose “untrained or uninitialized” which is disclosed by Xing:
the second neural network is untrained or uninitialized prior to being initialized using the set of weights generated by the additional neural network {Xing discloses [P.2 Last2¶] “untrained student sub-networks” preceded “our CKD framework starts with a powerful and pre-trained teacher network” and initializes weights as parameters Alg.1 [P.16] “Initialize parameters: θstu1 for student sub-network 1, θstu2 for student sub-network 2” with Line3 generated teacher sub-networks for student update shown Fig 1 student and teacher sub-networks for distillation framework}.
Xing is directed to neural network training for student-teacher distillation thus being analogous. A person having ordinary skill in the art would have considered it obvious prior to the effective filing date to specify untrained student networks for initialization per Xing for Xu’s random initialization in combination to arrive at the invention as claimed for a motivation that untrained student is initialized from scratch without bias, based on starting from a powerful pretrained teacher, and/or a further motivation “our goal was to solve the parameter redundancy problem in the teacher-student model and obtain small deep neural network with powerful feature extraction capabilities that can be easily deployed on lower-performance hardware devices and meet the accuracy requirements” [P.2 ¶6], contributions.
Claim 22 is rejected under 35 U.S.C. 103 as being unpatentable over: Xu and Liu in view of Wang et al., US PG Pub No 2022/0067274A1 hereinafter Wang.
With respect to claim 22, the combination of Xu and Liu teaches the system of claim 8. Wang teaches wherein
update the one or more weights of the second neural network without further training of the first neural network {Wang [0046] “update the weight parameters of each unit module of the student model” while [0040] “Freezing the weight of the teacher network…weight parameters belonging to the teacher network are frozen” i.e. [0038] “without training” see Figs 1-2 student-teacher distillation}.
Wang is directed to training of deep learning models with distillation thus being analogous. A person having ordinary skill in the art would have considered it obvious prior to the effective filing date to update weights of student with teacher weights frozen per Wang in combination to arrive at the invention as claimed for a motivation “Freezing the weight of the teacher network can also improve the efficiency of the whole model training” [0040].
Claim 23 is rejected under 35 U.S.C. 103 as being unpatentable over: Xu in view of Zhang et al., “Adaptive Multi-Teacher Knowledge Distillation with Meta-Learning” hereinafter Zhang (arXiv: 2306.06634v1, see PTO-892 dated 04/27/26 reference V).
With respect to claim 23, Xu teaches the method of claim 15. Zhang teaches further comprising:
providing, to the additional neural network as input, the one or more weights of the first neural network {Zhang Fig 2 additional neural network is “Meta-Weight network” dashed arrows provide “weight network input” arrows coming from Teacher 1 is a first neural network e.g. ResNet Table 1};
receiving, from the additional neural network, the set of weights {Zhang Fig 2 “Meta-Weight network” solid arrows away from additional/meta-weight network, the weights introduced Eqs.2,4 of meta-weight network are applied to student network training at Eqs.6-7 “obtain the weight vector wc …obtain the weight vector wf”. See also Eq.9 combined loss}; and
updating the one or more weights of the second neural network based, at least in part, on the set of weights {Zhang [P.3 Sect.C ¶3,2] “update the student and meta-weight network” particularly as “we define ϕ as the total parameters of the meta-weight network, which includes Wr1, Wr2, Wf1 and Wf2” cont’d “We then optimize ϕ to minimize the cross-entropy loss… calculate the gradient of ϕ based on the above equation” Eq.10. See also [P.3 Last¶] parameter initialization and ARI accuracy metric}.
Zhang is directed to neural network training with multi-teacher distillation thus being analogous. A person having ordinary skill in the art would have considered it obvious prior to the effective filing date to use the teachings of Zhang in combination to arrive at the invention as claimed for a motivation “benefit of learning from multiple teachers… multiple teachers have more diverse knowledge” thereby “guide the student… improve feature expression ability” [P.4 ¶1,3].
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure:
Shang et al., “Multi-teacher knowledge distillation based on joint Guidance of Probe and Adaptive Corrector” discloses top-k accuracy, intermediate representations and initializing, see Figs 1-2 and Alg.1-3
Sasagawa et al., US PG Pub No 2025/0036951A1 discloses accuracy b/w student-teacher subnetwork layers and with feature matching, [0131] “by matching the size of the feature maps, it is possible to accurately obtain the error between the teacher output and the student output. As a result, it is possible to accurately obtain weight data for each layer of the student neural network”
Chen et al., “D3ETR: Decoder Distillation for Detection Transformer” arXiv: 2211.09768v1 at Fig 2 illustrates matching between teacher and student models as distillation
Dong et al., “DisWOT: Student Architecture Search for Distillation WithOut Training” arXiv: 2303.15678v1 see Fig 3, Alg.1 DisWOT (dis)tillation without training
Chen et al., “MAKD: Multiple Auxiliary Knowledge Distillation” Alibaba Fig 2 matching loss and feature alignment
Hu et al., US PG Pub No 2023/0102489A1 Samsung Fig 3 shows multi-teacher distillation
Rabinowitz et al., US PG Pub No 2017/0337464A1 Google Fig 1 shows Progressive Networks sequence of DNN(s)
Gani et al., US PG Pub No 2024/0212330A1 see Fig 4 ViT vision transformer student-teacher
Huang et al., “Knowledge Diffusion for Distillation” arXiv: 2305.15712v1 Fig 3 diffusion-kd
Liu et al., “Large-scale Knowledge Distillation with Elastic Heterogeneous Computing Resources” arXiv: 2207.06667v1 Alg.1 servers, distributed training PaddlePaddle (Baidu)
Sun et al., “Cross-stage distillation with inverted bottleneck projectors” see Fig 2, Alg.1 discloses ARI accuracy metric and intermediate feature loss
Tancik et al., “Learned Initializations for Optimizing Coordinate-Based Neural Representations” arXiv: 2012.02189v2 Fig 1 meta-initialization
Zaidi et al., “When Does Re-initialization Work?” arXiv: 2206.10011v2 discloses a layer-wise re-initialization for distillation
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Chase P Hinckley whose telephone number is (571)272-7935. The examiner can normally be reached M-F 9:00 - 5:00.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Miranda M. Huang can be reached at 571-270-7092. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/CHASE P. HINCKLEY/Examiner, Art Unit 2124