8DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claim(s) 1, 3-4, 8-9, 11, 15-16, 18, 20 is/are rejected under 35 U.S.C. 102(a)(1) as being anticipated by K L Navaneet et al. (hereinafter Navaneet) (“SimReg: Regression as a Simple Yet Effective Tool for Self-supervised Knowledge Distillation,” 2022-01-13).
Regarding claim 1, Navaneet teaches;
A computer-implemented method for training a first machine learning model ([pg. 9] Multi-teacher Distillation: We train a single student model), the method comprising: processing first data ([Abstract] using the same weakly augmented image as input for both teacher and student networks … [pg. 4] input image x) via a plurality of trained machine learning models ([pg. 9] multiple teacher networks trained with different SSL methods) to generate a plurality of first outputs ([pg. 4] teacher network T … Let ft = T(x) … [pg. 5] Let fkt be the output vector of the kth teacher); processing the first data via the first machine learning model ([Abstract] using the same weakly augmented image as input for both teacher and student networks … [pg. 4] input image x) to generate a second output ([pg. 4] the output of the student network is obtained as fs =S(x)); processing the second output ([pg. 4] fs) via a plurality of projection heads ([pg. 10] using … a separate prediction head for each teacher) to generate a plurality of third outputs ([pg. 5] fks = gk(fs) that of the corresponding student prediction head gk(.)); computing a plurality of losses ([pg. 5] The multi-teacher distillation objective for K teachers is given by: L = 1/K ∑k d(fkt , fks) … loss term from corresponding teachers [Note: i.e., each d(fkt , fks)]) based on the plurality of first outputs ([pg. 5] for k teachers … fkt) and the plurality of third outputs ([pg. 10] using … a separate prediction head for each teacher … [pg. 5] fks); and performing one or more operations to update one or more parameters of the first machine learning model ([pg. 5] the backbone S is trained using the summation in Eq. 5 … … [Equation 5] L = 1/K ∑k d(fkt , fks)) and one or more parameters of the plurality of projection heads ([pg. 5] The prediction heads are trained by the loss term from corresponding teachers) based on the plurality of losses ([pg. 5] L = 1/K ∑k d(fkt , fks)).
Regarding claim 3, Navaneet teaches;
each projection head included in the plurality of projection heads ([pg. 10] using … a separate prediction head for each teacher) comprises a multi-layer perceptron ([pg. 4] prediction head g(.) … is modeled using a multi-layer perceptron (MLP)).
Regarding claim 4, Navaneet teaches;
each of the first machine learning model ([pg. 4] Let S be the student network) and the plurality of trained machine learning models ([pg. 3] Consider a teacher network T … [pg. 9] multiple teacher networks trained with different SSL methods) comprises an encoder model ([pg. 3] ft =T(x), ft ∈ Rc as the output vector (logits) corresponding to input image x … [pg. 4] fs = S(x), fs ∈ Rc the feature vector corresponding to input image x).
Regarding claim 8, Navaneet teaches;
computing a plurality of gradients ([pg. 5] We use SGD optimizer) based on the plurality of losses ([pg. 5] L = 1/K ∑k d(fkt , fks)); and performing one or more backpropagation operations ([pg. 5] We use SGD optimizer) to update the one or more parameters of the first machine learning model ([pg. 5] the backbone S is trained using [L = 1/K ∑k d(fkt , fks)]) and the one or more parameters of the plurality of projection heads ([pg. 5] The prediction heads are trained by the loss term from corresponding teachers) based on the plurality of gradients ([pg. 5] We use SGD optimizer).
NOTE: Navaneet teaches training the student backbone and prediction heads using SGD (Stochastic Gradient Descent) based on the disclosed distillation losses. In the disclosed neural network training, computation of gradients of the disclosed losses with respect to the trainable backbone and prediction head parameters, and propagation of those gradients backward through the networks, is necessarily performed for the disclosed SGD training.
Regarding claim 9, Navaneet teaches;
wherein the plurality of losses are weighted equally when performing the one or more operations to update the one or more parameters of the first machine learning model and the one or more parameters of the plurality of projection heads (Note: Each loss term d(fkt , fks) is weighted equally by 1/K in equation 5 … [equation 5] L = 1/K ∑k d(fkt , fks) … [pg. 5] The prediction heads are trained by the loss term from corresponding teachers while the backbone S is trained using the summation in Eq. 5).
Regarding claim 11,
Claim 11 is a non-transitory computer-readable storage media claim directly corresponding to computer-implemented method claim 1 and is rejected using the same reasoning.
Regarding claim 15,
Claim 15 is a non-transitory computer-readable storage media claim directly corresponding to computer-implemented method claim 8 and is rejected using the same reasoning.
Regarding claim 16,
Claim 16 is a non-transitory computer-readable storage media claim directly corresponding to computer-implemented method claim 9 and is rejected using the same reasoning.
Regarding claim 18, Navaneet teaches;
subsequent to the updating the one or more parameters of the first machine learning model (test stage occurs after the previously taught training stage of distillation, see fig. 1 below); processing second data (the test stage uses different/second image data, see fig. 1 below) via the first machine learning model ([pg. 2] Backbone student) to generate a fourth output (student inference feature, see fig. 1 below);
[pg. 2, fig. 1]
PNG
media_image1.png
531
1305
media_image1.png
Greyscale
and processing the fourth output (student inference feature, see fig. 1 above) via a decoder model ([pg. 5] single linear layer) to generate a fifth output (Note: the classification output generated by the downstream linear classification layer… [pg. 5] evaluate the performance of distilled networks on … classification task … We train a single linear layer atop the frozen backbone network for transfer evaluation).
NOTE: Navaneet teaches that, after distilling / updating the student backbone, subsequent image data is processed by the trained student backbone to generate a student inference feature, which is then processed by a trained downstream linear classification layer to generate a classification output. The trained linear classification layer therefore corresponds to the claimed decoder model, which processes the inference feature to generate downstream output.
Regarding claim 20,
Claim 20 is a system claim directly corresponding to computer-implemented method claim 1 and is rejected using the same reasoning.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 2, 12 is/are rejected under 35 U.S.C. 103 as being unpatentable over Navaneet (“SimReg: Regression as a Simple Yet Effective Tool for Self-supervised Knowledge Distillation,” 2022-01-13) as applied to claims 1 and 11 above, further in view of Haiping Wu et al. (hereinafter Wu) (“CvT: Introducing Convolutions to Vision Transformers,” 2021-03-29).
Regarding claim 2, Navaneet teaches;
the first machine learning model ([Abstract] student network)
Navaneet fails to teach but Wu teaches;
a hybrid convolutional and transformer neural network ([Abstract] We present in this paper a new architecture, named Convolutional vision Transformer (CvT), that improves Vision Transformer (ViT) in performance and efficiency by introducing convolutions into ViT).
OBVIOUSNESS TO COMBINE WU:
Wu is analogous art to the present disclosure as it pertains to hybrid convolutional and transformer neural networks. Navaneet teaches a multi-teacher distillation framework having a first student neural network model, while Wu teaches a hybrid convolutional and transformer neural network architecture which yields the benefits of both designs ([Abstract] These changes introduce desirable properties of convolutional neural networks (CNNs) to the ViT architecture (i.e. shift, scale, and distortion in variance) while maintaining the merits of Transformers (i.e. dynamic attention, global context, and better generalization)). Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date, to implement the student machine learning model of Navaneet using the convolutional vision transformer architecture of Wu to obtain the desirable properties of convolutional neural networks and of vision transformers.
Regarding claim 12,
Claim 12 is a non-transitory computer-readable storage media claim directly corresponding to computer-implemented method claim 2 and is rejected using the same reasoning.
Claim(s) 5 is/are rejected under 35 U.S.C. 103 as being unpatentable over Navaneet (“SimReg: Regression as a Simple Yet Effective Tool for Self-supervised Knowledge Distillation,” 2022-01-13) as applied to claims 1 and 11 above, further in view of Ximeng Sun et al. (hereinafter Sun) (“DIME-FM:DIstilling Multimodal and Efficient Foundation Models,” 2023-10-06).
Regarding claim 5, Navaneet teaches;
the first machine learning model ([Abstract] student network) and the plurality of trained machine learning models ([pg. 9] multiple teacher networks trained with different SSL methods)
Navaneet fails to explicitly teach but Sun teaches;
[teacher and student models] ([pg. 3] distill knowledge from the teacher to the student) comprises a foundation model ([Abstract] Large Vision-Language Foundation Models (VLFM) … [pg. 9] we propose … DIME-FM that distills knowledge in pre-trained VLFMs to small foundation models).
OBVIOUSNESS TO COMBINE SUN:
Sun is analogous art to the present disclosure as it pertains to knowledge distillation using foundational models. It would have been obvious to one of ordinary skill in the art, before the effective filing date, to implement the student and teacher models of Navaneet as foundation models, as taught by Sun, in order to obtain a distilled student model having the transferability and robustness of large pretrained foundation models ([Sun, Abstract] Large Vision-Language Foundation Models (VLFM), … achieve superior transferability and robustness on downstream tasks) while reducing model size and computational requirements ([Sun, Abstract] transfer the knowledge contained in large VLFMs to smaller, customized foundation models using a relatively small amount of inexpensive, unpaired images and sentences). Sun expressly teaches knowledge distillation from large pretrained foundation models to smaller customized foundation models to retain performance across downstream tasks. Such a modification would have constituted the predictable use of known foundation models within Navaneet’s known multi-teacher knowledge distillation framework to obtain the known benefits of generalizable representations and efficient deployment.
Claim(s) 6, 13 is/are rejected under 35 U.S.C. 103 as being unpatentable over Navaneet (“SimReg: Regression as a Simple Yet Effective Tool for Self-supervised Knowledge Distillation,” 2022-01-13) as applied to claims 1 and 11 above, further in view of Ahmet Iscen et al. (hereinafter Iscen) (“Class-Balanced Distillation for Long-Tailed Visual Recognition,” 2022-06-12).
Regarding claim 6, Navaneet teaches;
wherein computing the plurality of losses comprises, for each first output ([pg. 5] fkt ) included in the plurality of first outputs ([pg. 5] L = 1/K ∑k d(fkt , fks)), computing a ([pg. 4] L=Lreg =d(ft, fs) where d(.) is a distance metric) between one or more first summary features included in the first output ([pg. 5] Let fkt be the output vector of the kth teacher) and one or more second summary features included in one of the plurality of third outputs ([pg. 5] fks = gk(fs) that of the corresponding student prediction head gk(.)) that corresponds to the first output ([pg. 10] using … a separate prediction head for each teacher).
Navaneet fails to expressly teach but Iscen teaches;
computing a cosine distance between [corresponding features] ([pg. 6] 1−cos(v, x) tries to minimize the cosine distance between two feature descriptors)
OBVIOUSNESS TO COMBINE ISCEN:
Iscen is analogous art to the present disclosure as it pertains to knowledge distillation utilizing multiple teacher models and computing cosine distance between feature vectors. It would have been obvious to one of ordinary skill in the art, before the effective filing date, to employ the cosine-distance feature loss taught by Iscen as the distance metric d(.) for each corresponding teacher output and student prediction head output pair of Navaneet. Navaneet expressly teaches computing a distance between corresponding teacher and student feature representations and identifies d(.) generally as a distance metric, while Iscen teaches that teacher and student image feature descriptors may be matched during knowledge distillation using a cosine distance loss. A person of ordinary skill therefore would have had reason to employ Iscen’s known cosine distance metric in Navaneet’s feature distillation framework to encourage the student representation to match the corresponding teacher representation, amounting to the predictable use of a known feature matching loss for its established purpose.
Regarding claim 13,
Claim 13 is a non-transitory computer-readable storage media claim directly corresponding to computer-implemented method claim 6 and is rejected using the same reasoning.
Claim(s) 7, 14 is/are rejected under 35 U.S.C. 103 as being unpatentable over Navaneet (“SimReg: Regression as a Simple Yet Effective Tool for Self-supervised Knowledge Distillation,” 2022-01-13) as applied to claims 1 and 11 above, further in view of Hanqiu Deng et al. (hereinafter Deng) (“Anomaly Detection via Reverse Distillation from One-Class Embedding,” 2022), further in view of Yong Yang et al. (hereinafter Yang) (“Low-Light Image Enhancement Network Based on Multi-Scale Feature Complementation,” 2023-06-26).
Regarding claim 7, Navaneet teaches;
for each first output ([pg. 5] fkt ) included in the plurality of first outputs ([pg. 5] L = 1/K ∑k d(fkt , fks)), computing a ([pg. 4] L=Lreg =d(ft, fs) where d(.) is a distance metric) between one or more first ([pg. 5] Let fkt be the output vector of the kth teacher) and one or more second ([pg. 5] fks = gk(fs) that of the corresponding student prediction head gk(.)) that corresponds to the first output ([pg. 10] using … a separate prediction head for each teacher).
Navaneet fails to explicitly teach but Deng teaches;
cosine distance ([pg. 4] Specifically, for feature tensors fkE and fkD , we calculate their vector-wise cosine similarity loss)
[pg. 4]
PNG
media_image2.png
157
1051
media_image2.png
Greyscale
between [corresponding spatial features] ([pg. 4] the paired activation correspondence in our T-S model is {fkE = Ek(I), fkD = Dk(ϕ)}, where Ek and Dk represent the kth encoding and decoding block in the teacher and student model, respectively. fkE, fkD ∈ RCk× Hk× Wk)
OBVIOUSNESS TO COMBINE DENG:
It would have been obvious to one of ordinary skill in the art, before the effective filing date, to modify Navaneet’s knowledge distillation framework to use the spatial feature cosine distance technique taught by Deng because both references are directed to transferring knowledge from trained teacher networks to a student network by minimizing the distance between corresponding teacher and student feature representations. Deng teaches comparing corresponding spatial feature vectors using cosine distance, as it more precisely captures relations in both high and low dimensional information (see Deng, pg. 4). Accordingly, a person of ordinary skill in the art would have been motivated to apply Deng’s known spatial feature cosine loss to the corresponding teacher and projected student outputs of Navaneet to more precisely capture relationships between corresponding spatial feature representations, predictably improving the fidelity of knowledge transferred from the teacher models to the student model.
Navaneet and Deng fail to teach but Yang teaches;
combination of a cosine distance and a smooth L1 distance ([pg. 4] The content loss consists of a smooth L1 loss and a cosine similarity loss)
OBVIOUSNESS TO COMBINE YANG
Yang is analogous art to the present disclosure as it pertains to training neural networks for vision tasks using loss functions that measure differences between corresponding visual outputs. It would have been obvious to one of ordinary skill in the art, before the effective filing date, to modify the cosine-based spatial feature loss of Navaneet as modified by Deng to additionally include the Smooth L1 loss taught by Yang because Yang expressly teaches combining cosine similarity and Smooth L1 in a common loss for more closely aligning corresponding outputs. A person of ordinary skill in the art would therefore have recognized Yang’s combined loss as a known loss formulation that supplements a cosine based loss with Smooth L1 and would have been motivated to apply it to Deng’s comparison of corresponding teacher and student spatial features to provide an additional measure of discrepancy between those features, predictably improving their alignment during knowledge distillation.
Regarding claim 14,
Claim 14 is a non-transitory computer-readable storage media claim directly corresponding to computer-implemented method claim 7 and is rejected using the same reasoning.
Claim(s) 10 is/are rejected under 35 U.S.C. 103 as being unpatentable over Navaneet (“SimReg: Regression as a Simple Yet Effective Tool for Self-supervised Knowledge Distillation,” 2022-01-13) as applied to claim 1 above, further in view of Lingyun Gu et al. (hereinafter Gu) (“Learning Lightweight and Superior Detectors with Feature Distillation for Onboard Remote Sensing Object Detection,” 2023-01-07).
Regarding claim 10, Navaneet teaches;
one or more operations to update the one or more parameters of the first machine learning model and the one or more parameters of the plurality of projection heads (using the same reasoning from claim 1)
Navaneet fails to teach but Gu teaches;
performing one or more automatic loss balancing operations to determine a weight for each loss included in the plurality of losses ([pg. 5] Adaptive Dense Multi-teacherDistillation (ADMD) uses an adaptive weight to balance the loss terms of each student–teacher pair … [pg. 8] the weight values can be dynamically adjusted by network learning)
OBVIOUSNESS TO COMBINE GU:
Gu is analogous art to the present disclosure as it pertains to automatic loss balancing for a multi-teacher distillation framework. It would have been obvious to one of ordinary skill in the art, before the effective filing date, to modify the multi-teacher knowledge distillation framework of Navaneet to perform automatic loss balancing by dynamically determining a respective weight for each teacher specific loss, as taught by Gu. Navaneet teaches distilling knowledge from a plurality of trained teacher models into a student model using a plurality of respective teacher specific distillation losses. Gu similarly teaches multi-teacher feature distillation and teaches an adaptive dense multi-teacher distillation strategy that dynamically learns a respective weight for each student-teacher loss term and uses the weights to balance the loss terms, thereby allowing utilization of the unique knowledge of different teachers ([pg. 7] The straightforward approach to distilling knowledge from multiple teachers is to utilise the supervision signal as the average response from all teachers [32]. However, this approach is obviously oversimplified and cannot effectively utilise the unique knowledge of different teachers. To further efficiently perform knowledge transfer of multiple teachers, an adaptive dense multi-teacher distillation strategy based on adaptive weighted loss is proposed). Accordingly, one of ordinary skill in the art would have been motivated to apply Gu’s adaptive loss weighting technique to Navaneet’s plurality of teacher specific losses to more effectively integrate the differing useful knowledge provided by the respective teacher models, predictably improving the student’s feature representation.
Claim(s) 17 is/are rejected under 35 U.S.C. 103 as being unpatentable over Navaneet (“SimReg: Regression as a Simple Yet Effective Tool for Self-supervised Knowledge Distillation,” 2022-01-13) as applied to claim 11 above, further in view of Abdul Wasay et al. (hereinafter Wasay) (“MOTHERNETS: RAPID DEEP ENSEMBLE LEARNING,” 2020-03-08).
Regarding claim 17, Navaneet teaches;
processing the first data via each trained machine learning model included in the plurality of trained machine learning models (using the same reasoning from claim 1)
Navaneet fails to teach but Wasay teaches;
plurality of … machine learning models using a different set of processors ([pg. 16] an ensemble of multiple networks, … assign them to available GPUs … we assign one network to multiple GPUs dividing idle GPUs equally between networks … assigning them to as distinct set of GPUs as possible)
OBVIOUSNESS TO COMBINE WASAY:
Wasay is analogous art to the present disclosure as it pertains to an ensemble of neural networks each assigned to a different set of processors. It would have been obvious to one of ordinary skill in the art, before the effective filing date, to modify the multi-teacher distillation framework of Navaneet by assigning each of the plurality of teacher models to a different set of processors, using the multi-GPU model assignment method taught by Wasay. Navaneet teaches simultaneously using a plurality of trained teacher neural networks to process input data during knowledge distillation, while Wasay teaches executing a plurality of neural networks on respective distinct sets of GPUs and expressly explains that assigning the networks to as distinct sets of GPUs as possible minimizes communication overhead ([Wasay, pg. 5] minimizes communication overhead by assigning them to as distinct set of GPUs as possible). Accordingly, one of ordinary skill in the art would have been motivated to implement Navaneet’s respective trained teacher models on different sets of GPUs to minimize communication overhead when executing the plurality of teacher networks.
Claim(s) 19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Navaneet (“SimReg: Regression as a Simple Yet Effective Tool for Self-supervised Knowledge Distillation,” 2022-01-13) as applied to claim 11 above, further in view of Jiahui Yu et al. (hereinafter Yu) (“BigNAS: Scaling Up Neural Architecture Search with Big Single-Stage Models,” 2020-07-17).
Regarding claim 19, Navaneet teaches;
processing the first data via the plurality of trained machine learning models (using the same reasoning from claim 1) … trained machine learning model included in the plurality of trained machine learning models (The plurality of trained teacher models, using the same reasoning from claim 1).
Navaneet fails to teach but Yu teaches;
performing one or more interpolation operations on the first data to generate interpolated data ([pg. 5] apply bicubic interpolation to the same patch to transform it into all target resolutions); … and inputting the interpolated data into at least one [teacher model] ([pg. 5] feed the same image patches into both the teacher and the student)
OBVIOUSNESS TO COMBINE YU:
Yu is analogous art to the present disclosure as it pertains to knowledge distillation between teacher and student neural networks. It would have been obvious to one of ordinary skill in the art, before the effective filing date, to modify Navaneet’s knowledge distillation framework to interpolate input image data to a target resolution before processing the interpolated image data with one or more teacher models, as taught by Yu. Both Navaneet and Yu concern transferring knowledge from teacher models to student models using corresponding image inputs, and Yu teaches that when teacher and student models operate at different input resolutions, bicubic interpolation may be applied to the same image patch to generate the respective target resolution inputs. Yu further teaches that providing corresponding interpolated image patches in this manner makes the teacher predictions more compatible with the student inputs and provides a more accurate distillation signal ([Yu, pg. 5] soft labels predicted by the biggest child model (the teacher) are more compatible with the inputs seen by other child models (the students). Therefore this can serve as a more accurate distillation signal). Accordingly, a person of ordinary skill would have been motivated to apply Yu’s known input interpolation technique to one or more of Navaneet’s trained teacher models to accommodate differing input resolutions while preserving correspondence of the image content used for distillation.
CONCLUSION
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Matthew Alan Cady whose telephone number is (571) 272-7229. The examiner can normally be reached Monday - Friday, 7:30 am - 5:00 pm ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Cesar Paula can be reached on (571)272-4128. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC)
at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/MATTHEW ALAN CADY/ Examiner, Art Unit 2145
/CESAR B PAULA/ Supervisory Patent Examiner, Art Unit 2145