DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Objections
Claims 8, 20, 22, 26, 28, and 31 are objected to because of the following informalities:
In claim 8, line 3, “freezing all of a subset…” should read “freezing all or a subset…” in order to correct spelling error.
In claim 20, line 1, “wherein the downstream task…” should read “wherein a downstream task…” in order to use the proper indefinite article for the new element introduced.
In claim 22, line 1, “wherein said probing…” should read “wherein a probing…” in order to use the proper indefinite article for the new element introduced.
In claim 26, line 14, “using at least on the task objective…” should read “using at least in order to remove typo.
In claim 28, line 3, “receiving a input image…” should read “receiving an input image…” in order to use the proper indefinite article.
In claim 31, line 1, “the method of claim 30…” should read “the device of claim 30…” in order to refer to the correct claim type.
Appropriate correction is required.
Claim Interpretation
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked.
As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph:
(A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function;
(B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and
(C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function.
Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function.
Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function.
Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action.
Claim 28 recites limitations that use words like “means” (or “step”) or similar terms with functional language but do not invoke 35 U.S.C. 112(f):
Claim 28; recites the limitation, “processing the input image by a lightweight student network model…,” [Line 4].
Such claim limitation(s) is/are:
(i) “lightweight student network model….” has a structure associated with it a network.
Because this/these claim limitation(s) is/are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are not being interpreted to cover only the corresponding structure, material, or acts described in the specification as performing the claimed function, and equivalents thereof.
If applicant intends to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to remove the structure, materials, or acts that performs the claimed function; or (2) present a sufficient showing that the claim limitation(s) does/do not recite sufficient structure, materials, or acts to perform the claimed function.
Claims 26, 27, 28, and 30 recite limitations that use words like “means” (or “step”) or similar terms with functional language and do invoke 35 U.S.C. 112(f):
Claim 26; recites the limitation, “a model optimization module configured to: optimize…” [Line 9].
Claim 26; recites the limitation, “a training loss determination module configured to: determine…” [Line 17].
Claim 27; recites the limitation, “an extended dataset generation module configured to generate…” [Lines 1-2].
Claim 27; recites the limitation, “a synthetic data generation module for generating…” [Lines 3-4].
Claim 28; recites the limitation, “receiving a input image from an input device…,” [Line 3].
Claim 30; recites the limitation, “an input device for receiving an input image…,” [Line 3].
Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof.
After a careful analysis, as disclosed above, and a careful review of the specification the following limitations in claims 26, 27, 28, and 30:
“model optimization module” (Fig. 1, #114 called model optimization module, Paragraph [0250] – “Each module may be implemented using code. The term code, as used above, may include software, firmware, and/or microcode, and may refer to programs, routines, functions, classes, data structures, and/or objects.” Paragraph [0242-0243] – “The program code or the computer-executable instructions may, for example, be stored on a computer-readable storage medium. In an embodiment, a storage medium (or a data carrier, or a computer- readable medium) comprises, stored thereon, the computer program or the computer-executable instructions for performing one of the methods described herein where it is performed by a processor.” See also Paragraph [0055]. Thus, the model optimization module does have sufficient structure associated with it wherein it is a program whose functions are carried out by a processor.
“training loss determination module” (Fig. 1, #112 called training loss determination module, Paragraph [0250] – “Each module may be implemented using code. The term code, as used above, may include software, firmware, and/or microcode, and may refer to programs, routines, functions, classes, data structures, and/or objects.” Paragraph [0242-0243] – “The program code or the computer-executable instructions may, for example, be stored on a computer-readable storage medium. In an embodiment, a storage medium (or a data carrier, or a computer- readable medium) comprises, stored thereon, the computer program or the computer-executable instructions for performing one of the methods described herein where it is performed by a processor.” See also Paragraph [0055]. Thus, the training loss determination module does have sufficient structure associated with it wherein it is a program whose functions are carried out by a processor.
“extended dataset generation module” (Fig. 1, #120 called extended dataset generation module, Paragraph [0250] – “Each module may be implemented using code. The term code, as used above, may include software, firmware, and/or microcode, and may refer to programs, routines, functions, classes, data structures, and/or objects.” Paragraph [0242-0243] – “The program code or the computer-executable instructions may, for example, be stored on a computer-readable storage medium. In an embodiment, a storage medium (or a data carrier, or a computer- readable medium) comprises, stored thereon, the computer program or the computer-executable instructions for performing one of the methods described herein where it is performed by a processor.” See also Paragraph [0056]. Thus, the extended dataset generation module does have sufficient structure associated with it wherein it is a program whose functions are carried out by a processor.
“synthetic data generation module” (Fig. 1, #126 called synthetic data generation module, Paragraph [0250] – “Each module may be implemented using code. The term code, as used above, may include software, firmware, and/or microcode, and may refer to programs, routines, functions, classes, data structures, and/or objects.” Paragraph [0242-0243] – “The program code or the computer-executable instructions may, for example, be stored on a computer-readable storage medium. In an embodiment, a storage medium (or a data carrier, or a computer- readable medium) comprises, stored thereon, the computer program or the computer-executable instructions for performing one of the methods described herein where it is performed by a processor.” See also Paragraph [0056]. Thus, the synthetic data generation module does have sufficient structure associated with it wherein it is a program whose functions are carried out by a processor.
“input device” (Fig. 15, Paragraph [0241] – “If the trained task is an image processing task, for instance, one or more of the server 1502 or client devices 1504 may be provided with one or more imaging devices (CCDs, energy sensors, cameras, etc.) for directly or indirectly receiving new images (or image signals) of various origins and types for processing by the trained lightweight models.” Thus, the input device does have sufficient structure associated with it wherein it is an imaging device such as a camera, CCD, or energy sensor.
If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed
invention is not identically disclosed as set forth in section 102 of this title, if the
differences between the claimed invention and the prior art are such that the claimed
invention as a whole would have been obvious before the effective filing date of the
claimed invention to a person having ordinary skill in the art to which the claimed
invention pertains. Patentability shall not be negated by the manner in which the
invention was made.
Claims 1-13, 18, 20, 21, 23, and 25 are rejected under 35 U.S.C. 103 as being unpatentable over CHEN (US 20250086462 A1), hereinafter referenced as CHEN in view of GUO (US 20240070454 A1), hereinafter referenced as GUO.
Regarding claim 1, CHEN teaches a computer-implemented method for training a lightweight student network model for performing a task (Fig. 13, Paragraph [0119] - CHEN discloses FIG. 13 depicts an example graphical diagram 1300 for performing distillation training of a student model based on a fine-tuned classification model with one or more projection head neural network layers that have been pretrained using contrastive learning. See also Fig. 14A-B, Paragraph [0127-0128].), the method comprising:
augmenting a pretrained network model (Fig. 13, #1304 called classification model, Paragraph [0122]) with a task-specific prediction head to provide a teacher model (Fig. 13, Paragraph [0122] - CHEN discloses classification model 1304 also includes one or more projection head layer(s) 1308, for example, originally from a projection head neural network pretrained with contrastive learning using unlabeled training data, where the specific projection head layer(s) were preserved after the pretraining and later fine-tuned based on the set of label training data. Further, classification head 1310 may receive and process one or more representations to generate classification output 1312, such as a classification prediction, detection prediction, recognition prediction, segmentation prediction, and/or other types of predictions and prediction tasks.),
the pretrained network model (Fig. 13, #1304 called classification model, Paragraph [0122]) being larger than the lightweight student network model (Fig. 13, Paragraph [0123] - CHEN discloses classification model 1304 is used to train a student network 1314 that is more specialized for a target task. For example, fine-tuned classification model 1304 is used when performing distillation training to distill the model to student network 1314 comprising a relatively smaller number of parameters relative to image classification model. As such, student network 1314 generally is lightweight and better suited to be deployed to client computing devices with limited local computing resources.);
training the teacher model on the task (Fig. 13, Paragraph [0122] - CHEN discloses classification model 1304 also includes one or more projection head layer(s) 1308, for example, originally from a projection head neural network pretrained with contrastive learning using unlabeled training data, where the specific projection head layer(s) were preserved after the pretraining and later fine-tuned based on the set of label training data.),
said training the teacher model using a task objective (Fig. 13, Paragraph [0122] - CHEN discloses classification head 1310 may receive and process one or more representations to generate classification output 1312, such as a classification prediction, detection prediction, recognition prediction, segmentation prediction, and/or other types of predictions and prediction tasks.);
training the lightweight student network model jointly using at least the task objective and a distillation objective (Fig. 13, Paragraph [0123] - CHEN discloses classification model 1304 is used to train a student network 1314 that is more specialized for a target task. For example, fine-tuned classification model 1304 is used when performing distillation training to distill the model to student network 1314 comprising a relatively smaller number of parameters relative to image classification model.),
Although CHEN further teaches and storing the trained lightweight student network model (Fig. 14A, Paragraph [0093] - CHEN discloses one or more models 120 can be stored and implemented at the user computing device 102 and/or one or more models 140 can be stored and implemented at the server computing system 130.).
CHEN fails to explicitly teach the distillation objective comprising a similarity between predictions from the lightweight student model and predictions from the trained teacher model;
However, GUO explicitly teaches the distillation objective comprising a similarity between predictions from the lightweight student model and predictions from the trained teacher model (Fig. 2, Paragraph [0025] - GUO discloses distillation based on relationship, that is, considering the differences between the teacher model and the student model in metrics such as similarity for different samples, thereby guiding the training of the student model.);
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date the claimed invention was made to combine the teachings of CHEN of having a computer-implemented method for training a lightweight student network model for performing a task, the method comprising: augmenting a pretrained network model with a task-specific prediction head to provide a teacher model, the pretrained network model being larger than the lightweight student network model; training the teacher model on the task, said training the teacher model using a task objective; training the lightweight student network model jointly using at least the task objective and a distillation objective, and storing the trained lightweight student network model with the teachings of GUO of having the distillation objective comprising a similarity between predictions from the lightweight student model and predictions from the trained teacher model.
Wherein CHEN’s computer-implemented method wherein the distillation objective comprising a similarity between predictions from the lightweight student model and predictions from the trained teacher model.
The motivation behind this modification would have been to provide an enhanced method for training a lightweight student network model that provides computational efficiency and training precision, since both CHEN and GUO relate to training of a student model, wherein CHEN discloses distillation training is performed based on reusing the unlabeled pretraining data to distill the network to a student network that performs one or more specialized tasks wherein such semi-supervised contrastive learning improves accuracy and computational efficiency over previously known methods, and GUO discloses a lightweight model training solution, which may effectively prevent the underfitting phenomenon in the case where a knowledge distillation strategy is adopted to train the lightweight model, thereby improving training precision of the lightweight model and enhancing precision of the knowledge distillation of the lightweight model. Please see CHEN (US 20250086462 A1), Paragraph [0035], and GUO (US 20240070454 A1), Paragraph [0030].
Regarding claim 2, CHEN in view of GUO teach the method of claim 1,
CHEN further teaches wherein the pretrained network model (Fig. 13, #1304 called classification model, Paragraph [0122] - CHEN discloses classification model 1304 includes a fine-tuned network 1306, such as a network (e.g., a fine-tuned base encoder neural network, large convolutional neural network, etc.) that was first pretrained with contrastive learning using unlabeled training data and later fine-tuned based on a relatively small set of labeled training data.)
includes pretrained network model parameters (Fig. 13, Paragraph [0123] - CHEN discloses fine-tuned classification model 1304 is used when performing distillation training to distill the model to student network 1314 comprising a relatively smaller number of parameters relative to image classification model. See also Fig. 12, Paragraph [0117].),
the task-specific prediction head (Fig. 13, #1310 called classification head, Paragraph [0122]) includes task- specific prediction head parameters (Fig. 13, Paragraph [0122] - CHEN discloses classification head 1310 may receive and process one or more representations to generate classification output 1312, such as a classification prediction, detection prediction, recognition prediction, segmentation prediction, and/or other types of predictions and prediction tasks.),
and the lightweight student network model (Fig. 13, #1314 called student network, Paragraph [0123]) includes lightweight student network model parameters (Fig. 13, Paragraph [0123] - CHEN discloses classification model 1304 is used to train a student network 1314 that is more specialized for a target task. For example, fine-tuned classification model 1304 is used when performing distillation training to distill the model to student network 1314 comprising a relatively smaller number of parameters relative to image classification model.).
Regarding claim 3, CHEN in view of GUO teach the method of claim 2, CHEN further teaches
further comprising: before said training the teacher model, initializing the pretrained network model parameters (Fig. 11, Paragraph [0110] - CHEN discloses the computer system performs unsupervised pretraining of a large model using a modified version SimCLR. For example, where in some examples, SimCLR training generally may involve ResNet-50 (4×) models, the computer system generally performs unsupervised pretraining of larger models with increased depth and width, such as a 152-layer ResNet with 3× wider channels and selective kernels, a channel-wise attention mechanism that improves parameter efficiency, performance, and accuracy.);
and/or before said training the teacher model, pretraining the pretrained network model (Fig. 11, Paragraph [0109] - CHEN discloses the computer system performs unsupervised pretraining of a model using contrastive learning based on a set of unlabeled training data.),
wherein said pretraining is self-supervised (Fig. 11, Paragraph [0109] - CHEN discloses pretraining of a model may be performed using unsupervised or self-supervised contrastive learning based on unlabeled, task agnostic training data without class labels and without being directed or tailored to a specific classification task.).
Regarding claim 4, CHEN in view of GUO teach the method of claim 2,
CHEN further teaches further comprising: initializing the lightweight student network model parameters (Fig. 13, Paragraph [0123] - CHEN discloses fine-tuned classification model 1304 is used when performing distillation training to distill the model to student network 1314 comprising a relatively smaller number of parameters relative to image classification model. In various examples, student network 1314 obtains or otherwise receives and processes unlabeled distillation input data 1302 to generate student classification output 1316. See also Paragraph [0124], Equation [0004].).
Regarding claim 5, CHEN in view of GUO teach the method of claim 2,
CHEN further teaches further comprising; pretraining the lightweight student network model parameters with a task- agnostic distillation (Fig. 13, Paragraph [0118] - CHEN discloses the computing system may reuse the unlabeled pretraining data directly when performing distillation as part of training a lightweight student network specialized for one or more targeted tasks. As such, the unlabeled training data first is used in a task-agnostic fashion for pretraining and then again used in distillation after performing fine-tuning to train a student network for one or more specialized targeted tasks.).
Regarding claim 6, CHEN in view of GUO teach the method of claim 1,
CHEN further teaches wherein the task objective is supervised (Fig. 2B, Paragraph [0051] - CHEN discloses task specific model 250 and/or the base encoder neural network 204 can be additionally trained (e.g., “fine-tuned”) on additional training data (e.g., which may be task specific data). The additional training can be, for example, supervised learning training.).
Regarding claim 7, CHEN in view of GUO teach the method of claim 2,
CHEN further teaches wherein said training the teacher model comprises: freezing all or a subset of the pretrained network model parameters (Fig. 12, Paragraph [0111] - CHEN discloses the computing system generates or configures a classification model for fine-tuning that includes some but not all of multiple projection head neural network layers that have been pretrained using contrastive learning with unlabeled training data, as further described with respect to FIG. 12 and in other examples of the present disclosure. See also Paragraph [0117, 0122].);
and optimizing the task-specific prediction head parameters (Fig. 2B, Paragraph [0051] - CHEN discloses task specific model 250 and/or the base encoder neural network 204 can be additionally trained (e.g., “fine-tuned”) on additional training data (e.g., which may be task specific data).).
Regarding claim 8, CHEN in view of GUO teach the method of claim 7,
CHEN further teaches wherein said training the lightweight student network model (Fig. 13, #1314 called student network, Paragraph [0123]) comprises: freezing all of a subset of the trained teacher model including the pretrained network model parameters and the task-specific prediction head parameters (Fig. 13, Paragraph [0125] - CHEN discloses a teacher network (i.e., classification model 1304) that outputs P.sup.T(y x.sub.i) can be fixed [wherein fixed is frozen] during the distillation, so only student network 1314 is trained. See also Equation [0005].).
Regarding claim 9, CHEN in view of GUO teach the method of claim 1,
CHEN further teaches wherein the task-specific prediction head comprises one or more of: a multilayer perceptron (Fig. 2A, Paragraph [0046] - CHEN discloses the projection head neural network 206 can be a multi-layer perceptron with one hidden layer.);
or a convolutional layer (Fig. 2A, Paragraph [0110] - CHEN discloses a projection head neural network may include three or more layers, a portion of which may be later reused during fine-tuning and distillation, instead discarding the projection head neural network entirely after pretraining.).
Regarding claim 10, CHEN in view of GUO teach the method of claim 1,
CHEN further teaches wherein the task-specific prediction head comprises a linear layer (Fig. 12, Paragraph [0117] - CHEN discloses some but not all pretrained projection head layer(s) 1208 may be added as one or more respective linear transformation layers on top of a pretrained network.).
Regarding claim 11, CHEN in view of GUO teach the method of claim 8,
CHEN further teaches wherein training the lightweight student network model (Fig. 13, Paragraph [0119] - CHEN discloses FIG. 13 depicts an example graphical diagram 1300 for performing distillation training of a student model based on a fine-tuned classification model with one or more projection head neural network layers that have been pretrained using contrastive learning.)
further uses a weighting parameter for the distillation objective and/or the task objective (Fig. 13, Paragraph [0126] - CHEN discloses when distillation training involves labeled training data, distillation loss 1318 may be combined with ground-truth labeled examples using a weighted combination as follows.).
Regarding claim 12, CHEN in view of GUO teach the method of claim 1,
CHEN fails to explicitly teach wherein the distillation objective comprises an average of a dissimilarity measure between predictions from the lightweight student model and predictions from the teacher model.
However, GUO explicitly teaches wherein the distillation objective comprises an average of a dissimilarity measure between predictions from the lightweight student model and predictions from the teacher model (Fig. 2, Paragraph [0025] - GUO discloses distillation based on relationship, that is, considering the differences between the teacher model and the student model in metrics such as similarity for different samples, thereby guiding the training of the student model.).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date the claimed invention was made to combine the teachings of CHEN in view of GUO of having a computer-implemented method for training a lightweight student network model for performing a task, the method comprising: augmenting a pretrained network model with a task-specific prediction head to provide a teacher model, the pretrained network model being larger than the lightweight student network model; training the teacher model on the task, said training the teacher model using a task objective; training the lightweight student network model jointly using at least the task objective and a distillation objective, the distillation objective comprising a similarity between predictions from the lightweight student model and predictions from the trained teacher model; with the teachings of GUO of having wherein the distillation objective comprises an average of a dissimilarity measure between predictions from the lightweight student model and predictions from the teacher model.
Wherein CHEN’s computer-implemented method wherein the distillation objective comprises an average of a dissimilarity measure between predictions from the lightweight student model and predictions from the teacher model.
The motivation behind this modification would have been to provide an enhanced method for training a lightweight student network model that provides computational efficiency and training precision, since both CHEN and GUO relate to training of a student model, wherein CHEN discloses distillation training is performed based on reusing the unlabeled pretraining data to distill the network to a student network that performs one or more specialized tasks wherein such semi-supervised contrastive learning improves accuracy and computational efficiency over previously known methods, and GUO discloses a lightweight model training solution, which may effectively prevent the underfitting phenomenon in the case where a knowledge distillation strategy is adopted to train the lightweight model, thereby improving training precision of the lightweight model and enhancing precision of the knowledge distillation of the lightweight model. Please see CHEN (US 20250086462 A1), Paragraph [0035], and GUO (US 20240070454 A1), Paragraph [0030].
Regarding claim 13, CHEN in view of GUO teach the method of claim 1,
CHEN further teaches wherein the task objective is over a labeled original dataset (Fig. 12, Paragraph [0117] - CHEN discloses fine-tuning of classification model 1204 may be performed by adjusting various parameters based on labeled fine-tuning input data 1202, for example, using a supervised cross-entropy loss or other type of loss function (not shown), allowing classification model 1204 to slightly adjust internal representations for one or more specific tasks.);
and wherein the distillation objective is over an extended dataset generated using data from the labeled original dataset (Fig. 13, Paragraph [0124] - CHEN discloses a fine-tuned classification model 1304 provides labels for training student network 1314 and distillation loss 1318 may be minimized. See also Equation [0004]. Paragraph [0126] - CHEN discloses when distillation training involves labeled training data, distillation loss 1318 may be combined with ground-truth labeled examples using a weighted combination. See also Equation [0005].).
Regarding claim 18, CHEN in view of GUO teach the method of claim 1,
CHEN further teaches wherein: the pretrained network model includes at least two times a number of model parameters compared to the lightweight student network model (Fig. 13, Paragraph [0123] - CHEN discloses student network 1314 comprising a relatively smaller number of parameters relative to image classification model [wherein image classification model is the pretrained network model, see Paragraph [0121].]. As such, student network 1314 generally is lightweight and better suited to be deployed to client computing devices with limited local computing resources.);
and/or the pretrained network model requires at least two times an amount of computing resources compared to the lightweight student network model (Fig. 13, Paragraph [0123] - CHEN discloses student network 1314 comprising a relatively smaller number of parameters relative to image classification model [wherein image classification model is the pretrained network model, see Paragraph [0121].]. As such, student network 1314 generally is lightweight and better suited to be deployed to client computing devices with limited local computing resources.).
Regarding claim 20, CHEN in view of GUO teach the method of claim 1,
CHEN further teaches wherein the downstream task is an image processing task (Fig. 13, Paragraph [0121] - CHEN discloses classification model 1304 may be an image classification model or any other type of classification model. etc. Paragraph [0122] - CHEN further discloses classification head 1310 may receive and process one or more representations to generate classification output 1312, such as a classification prediction, detection prediction, recognition prediction, segmentation prediction, and/or other types of predictions and prediction tasks.);
and wherein the data comprises images (Fig. 13, Paragraph [0109] - CHEN discloses training data generally may include any type of visual and non-visual data including, but not limited to, images, video content, image frames of video content, audio data, textual data, geospatial data, sensor data, etc.).
Regarding claim 21, CHEN in view of GUO teach the method of claim 20,
CHEN further teaches wherein the image processing task comprises image classification, object detection, and/or semantic segmentation (Fig. 13, Paragraph [0121] - CHEN discloses classification model 1304 may be an image classification model or any other type of classification model. etc. Paragraph [0122] - CHEN further discloses classification head 1310 may receive and process one or more representations to generate classification output 1312, such as a classification prediction [wherein classification prediction is image classification], detection prediction [wherein detection prediction is object detection], recognition prediction, segmentation prediction [wherein segmentation prediction is semantic segmentation], and/or other types of predictions and prediction tasks.).
Regarding claim 23, CHEN in view of GUO teach the method of claim 1,
CHEN further teaches wherein the pretrained network model comprises an encoder (Fig. 2, Paragraph [0110] - CHEN discloses pretraining may be performed using a projection head neural network having three or more layers on top of a ResNet encoder or other encoder, such as base encoder neural network 204. See also Fig. 13, Paragraph [0122].).
Regarding claim 25, CHEN teaches a computer-implemented system for training a lightweight student network model for performing a task (Fig. 14A, Paragraph [0127] - CHEN discloses FIG. 14A illustrates one example computing system that can be used to implement the present disclosure. Paragraph [0090] - CHEN further discloses system 100 includes a user computing device 102, a server computing system 130, and a training computing system 150 that are communicatively coupled over a network 180.), the system comprising:
a processor (Fig. 14A, #112 called processors, Paragraph [0090] - CHEN discloses user computing device 102 includes one or more processors 112 and a memory 114.);
a memory (Fig. 14A, #114 called memory, Paragraph [0090] - CHEN discloses user computing device 102 includes one or more processors 112 and a memory 114.);
and executable instructions stored in the memory for causing the processor to perform a method (Fig. 14A, Paragraph [0090] - CHEN discloses memory 114 can store data 116 and instructions 118 which are executed by the processor 112 to cause the user computing device 102 to perform operations.) comprising:
augmenting a pretrained network model (Fig. 13, #1304 called classification model, Paragraph [0122]) with a task-specific prediction head to provide a teacher model (Fig. 13, Paragraph [0123] - CHEN discloses classification model 1304 is used to train a student network 1314 that is more specialized for a target task. For example, fine-tuned classification model 1304 is used when performing distillation training to distill the model to student network 1314 comprising a relatively smaller number of parameters relative to image classification model. As such, student network 1314 generally is lightweight and better suited to be deployed to client computing devices with limited local computing resources.),
the pretrained network model (Fig. 13, #1304 called classification model, Paragraph [0122]) being larger than the lightweight student network model (Fig. 13, Paragraph [0123] - CHEN discloses classification model 1304 is used to train a student network 1314 that is more specialized for a target task. For example, fine-tuned classification model 1304 is used when performing distillation training to distill the model to student network 1314 comprising a relatively smaller number of parameters relative to image classification model. As such, student network 1314 generally is lightweight and better suited to be deployed to client computing devices with limited local computing resources.);
training the teacher model on the task (Fig. 13, Paragraph [0122] - CHEN discloses classification model 1304 also includes one or more projection head layer(s) 1308, for example, originally from a projection head neural network pretrained with contrastive learning using unlabeled training data, where the specific projection head layer(s) were preserved after the pretraining and later fine-tuned based on the set of label training data.)
using a task objective (Fig. 13, Paragraph [0122] - CHEN discloses classification head 1310 may receive and process one or more representations to generate classification output 1312, such as a classification prediction, detection prediction, recognition prediction, segmentation prediction, and/or other types of predictions and prediction tasks.);
Although CHEN further teaches training the lightweight student network model jointly using the task objective and a distillation objective (Fig. 13, Paragraph [0123] - CHEN discloses classification model 1304 is used to train a student network 1314 that is more specialized for a target task. For example, fine-tuned classification model 1304 is used when performing distillation training to distill the model to student network 1314 comprising a relatively smaller number of parameters relative to image classification model.),
CHEN fails to explicitly teach the distillation objective comprising a similarity between predictions from the lightweight student model and predictions from the trained teacher model.
However, GUO explicitly teaches the distillation objective comprising a similarity between predictions from the lightweight student model and predictions from the trained teacher model (Fig. 2, Paragraph [0025] - GUO discloses distillation based on relationship, that is, considering the differences between the teacher model and the student model in metrics such as similarity for different samples, thereby guiding the training of the student model.).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date the claimed invention was made to combine the teachings of CHEN of having a computer-implemented system for training a lightweight student network model for performing a task, the system comprising: a processor; a memory; and executable instructions stored in the memory for causing the processor to perform a method comprising: augmenting a pretrained network model with a task-specific prediction head to provide a teacher model, the pretrained network model being larger than the lightweight student network model; training the teacher model on the task using a task objective; training the lightweight student network model jointly using the task objective and a distillation objective, with the teachings of GUO of having the distillation objective comprising a similarity between predictions from the lightweight student model and predictions from the trained teacher model.
Wherein CHEN’s computer-implemented system wherein the distillation objective comprising a similarity between predictions from the lightweight student model and predictions from the trained teacher model.
The motivation behind this modification would have been to provide an enhanced system for training a lightweight student network model that provides computational efficiency and training precision, since both CHEN and GUO relate to training of a student model, wherein CHEN discloses distillation training is performed based on reusing the unlabeled pretraining data to distill the network to a student network that performs one or more specialized tasks wherein such semi-supervised contrastive learning improves accuracy and computational efficiency over previously known methods, and GUO discloses a lightweight model training solution, which may effectively prevent the underfitting phenomenon in the case where a knowledge distillation strategy is adopted to train the lightweight model, thereby improving training precision of the lightweight model and enhancing precision of the knowledge distillation of the lightweight model. Please see CHEN (US 20250086462 A1), Paragraph [0035], and GUO (US 20240070454 A1), Paragraph [0030].
Claims 14-17, 19, 22, 24, 26, and 27 are rejected under 35 U.S.C. 103 as being unpatentable over CHEN (US 20250086462 A1), hereinafter referenced as CHEN in view of GUO (US 20240070454 A1), hereinafter referenced as GUO, in further view of STEINER (US 20250117893 A1), hereinafter referenced as STEINER.
Regarding claim 14, CHEN in view of GUO teach the method of claim 13,
CHEN further teaches wherein the dataset comprises original data and corresponding labels (Fig. 12, Paragraph [0116] - CHEN discloses classification model 1204 obtains or otherwise receives a set of labeled, fine-tuning input data 1202.);
wherein the extended dataset comprises: the original data (Fig. 13, Paragraph [0126] - CHEN discloses when distillation training involves labeled training data, distillation loss 1318 may be combined with ground-truth labeled examples using a weighted combination. See also Equation [0005].);
CHEN in view of GUO fail to explicitly teach and synthetic data generated from the original data.
However, STEINER explicitly teaches and synthetic data generated from the original data (Fig. 1, Paragraph [0068] - STEINER discloses model trainer 108 can be or include or be implemented by a computing system configured to provide inputs to machine-learned model(s) 102 and evaluate outputs from machine-learned model(s) 102 to provide model update(s) 110. Paragraph [0070] - STEINER discloses model trainer 108 can access a training dataset that contains image(s) 104. Images 104 can include a mixture of image attributes. An example attribute includes magnification or zoom. Paragraph [0073] - STEINER discloses a mixture of magnifications can be synthesized. For example, a mixture of magnifications can be synthesized by transforming an original image (e.g., an image having a native magnification) to emulate various different recording magnifications.).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date the claimed invention was made to combine the teachings of CHEN in view of GUO of having a computer-implemented method for training a lightweight student network model for performing a task, the method comprising: augmenting a pretrained network model with a task-specific prediction head to provide a teacher model, the pretrained network model being larger than the lightweight student network model; training the teacher model on the task, said training the teacher model using a task objective; training the lightweight student network model jointly using at least the task objective and a distillation objective, the distillation objective comprising a similarity between predictions from the lightweight student model and predictions from the trained teacher model; with the teachings of STEINER of having and synthetic data generated from the original data.
Wherein CHEN’s computer-implemented method wherein having and synthetic data generated from the original data.
The motivation behind this modification would have been to provide an enhanced method for training a lightweight student network model that provides computational efficiency and improved accuracy of an image processing model, since both CHEN and STEINER relate to training a lightweight model, wherein CHEN discloses distillation training is performed based on reusing the unlabeled pretraining data to distill the network to a student network that performs one or more specialized tasks wherein such semi-supervised contrastive learning improves accuracy and computational efficiency over previously known methods, and STEINER relates to self-supervised training of machine-learned image processing models; benefits of the present disclosure can include improved accuracy or quality of the image processing model; reduced energy consumption; reduced training costs and training time; optimizing with limited or no downstream task data; greater data efficiency in downstream fine-tuning; and improved versatility of a pretrained image processing model. Please see CHEN (US 20250086462 A1), Paragraph [0035], and STEINER (US 20250117893 A1), Paragraph [0059-0063].
Regarding claim 15, CHEN and GUO in view of STEINER teach the method of claim 13,
CHEN further teaches further comprising: generating the extended dataset (Fig. 2A, Paragraph [0043] - CHEN discloses stochastic data augmentation module (shown generally at 203) that transforms any given data example (e.g., an input image x shown at 202) randomly resulting in two correlated views of the same example. Paragraph [0044] - CHEN further discloses three augmentations can be applied at 203: random cropping followed by resize back to the original size, random color distortions, and random Gaussian blur.);
wherein said generating the extended dataset (Fig. 2A, Paragraph [0043] - CHEN discloses stochastic data augmentation module (shown generally at 203) that transforms any given data example (e.g., an input image x shown at 202) randomly resulting in two correlated views of the same example. Paragraph [0044] - CHEN further discloses three augmentations can be applied at 203: random cropping followed by resize back to the original size, random color distortions, and random Gaussian blur.) comprises:
CHEN and GUO fail to explicitly teach generating synthetic data from the original data using a generative model.
However, STEINER explicitly teaches generating synthetic data from the original data (Fig. 1, Paragraph [0068] - STEINER discloses model trainer 108 can be or include or be implemented by a computing system configured to provide inputs to machine-learned model(s) 102 and evaluate outputs from machine-learned model(s) 102 to provide model update(s) 110. Paragraph [0070] - STEINER discloses model trainer 108 can access a training dataset that contains image(s) 104. Images 104 can include a mixture of image attributes. An example attribute includes magnification or zoom. Paragraph [0073] - STEINER discloses a mixture of magnifications can be synthesized. For example, a mixture of magnifications can be synthesized by transforming an original image (e.g., an image having a native magnification) to emulate various different recording magnifications. See also Paragraph [0302].)
using a generative model (Fig. 35, Paragraph [0302] - STEINER discloses mapping at ratios less than unity can be implemented by upsampling, oversampling, etc., including using machine-learned models to generate an output image at an output resolution based on an input image at a different input resolution. Various machine-learned models can include transformer-based models, convolutional neural networks, diffusion-based models, etc.).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date the claimed invention was made to combine the teachings of CHEN in view of GUO of having a computer-implemented method for training a lightweight student network model for performing a task, the method comprising: augmenting a pretrained network model with a task-specific prediction head to provide a teacher model, the pretrained network model being larger than the lightweight student network model; training the teacher model on the task, said training the teacher model using a task objective; training the lightweight student network model jointly using at least the task objective and a distillation objective, the distillation objective comprising a similarity between predictions from the lightweight student model and predictions from the trained teacher model; with the teachings of STEINER of having generating synthetic data from the original data using a generative model.
Wherein CHEN’s computer-implemented method wherein said generating the extended dataset comprises: generating synthetic data from the original data using a generative model.
The motivation behind this modification would have been to provide an enhanced method for training a lightweight student network model that provides computational efficiency and improved accuracy of an image processing model, since both CHEN and STEINER relate to training a lightweight model, wherein CHEN discloses distillation training is performed based on reusing the unlabeled pretraining data to distill the network to a student network that performs one or more specialized tasks wherein such semi-supervised contrastive learning improves accuracy and computational efficiency over previously known methods, and STEINER relates to self-supervised training of machine-learned image processing models; benefits of the present disclosure can include improved accuracy or quality of the image processing model; reduced energy consumption; reduced training costs and training time; optimizing with limited or no downstream task data; greater data efficiency in downstream fine-tuning; and improved versatility of a pretrained image processing model. Please see CHEN (US 20250086462 A1), Paragraph [0035], and STEINER (US 20250117893 A1), Paragraph [0059-0063].
Regarding claim 16, CHEN and GUO in view of STEINER teach the method of claim 15,
CHEN in view of GUO fail to explicitly teach wherein said generating the synthetic data uses a diffusion-based method that processes subsets of original samples to generate each of a plurality of synthetic samples.
However, STEINER explicitly teaches wherein said generating the synthetic data uses a diffusion-based method that processes subsets of original samples to generate each of a plurality of synthetic samples (Fig. 35, Paragraph [0302] - STEINER discloses at 3504, example method 3500 includes generating, from the reference histopathology image [wherein the reference histopathology image is original data], a plurality of image patches at a respectively plurality of emulated magnifications [wherein image patches at emulated magnifications are synthetic data], wherein the plurality of image patches conform to an input dimension of the image processing model. For instance, a slide image scanned at 20× magnification can be used to emulate one quarter of an image of the same slide scanned at 10× magnification. Accordingly, an image can provide a patch at native magnification by mapping pixels in the image to pixels in the patch at a unity ratio. Paragraph [0302] - STEINER discloses mapping at ratios less than unity can be implemented by upsampling, oversampling, etc., including using machine-learned models to generate an output image at an output resolution based on an input image at a different input resolution. Various machine-learned models can include transformer-based models, convolutional neural networks, diffusion-based models, etc.).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date the claimed invention was made to combine the teachings of CHEN in view of GUO of having a computer-implemented method for training a lightweight student network model for performing a task, the method comprising: augmenting a pretrained network model with a task-specific prediction head to provide a teacher model, the pretrained network model being larger than the lightweight student network model; training the teacher model on the task, said training the teacher model using a task objective; training the lightweight student network model jointly using at least the task objective and a distillation objective, the distillation objective comprising a similarity between predictions from the lightweight student model and predictions from the trained teacher model; with the teachings of STEINER of having wherein said generating the synthetic data uses a diffusion-based method that processes subsets of original samples to generate each of a plurality of synthetic samples.
Wherein CHEN’s computer-implemented method wherein said generating the synthetic data uses a diffusion-based method that processes subsets of original samples to generate each of a plurality of synthetic samples.
The motivation behind this modification would have been to provide an enhanced method for training a lightweight student network model that provides computational efficiency and improved accuracy of an image processing model, since both CHEN and STEINER relate to training a lightweight model, wherein CHEN discloses distillation training is performed based on reusing the unlabeled pretraining data to distill the network to a student network that performs one or more specialized tasks wherein such semi-supervised contrastive learning improves accuracy and computational efficiency over previously known methods, and STEINER relates to self-supervised training of machine-learned image processing models; benefits of the present disclosure can include improved accuracy or quality of the image processing model; reduced energy consumption; reduced training costs and training time; optimizing with limited or no downstream task data; greater data efficiency in downstream fine-tuning; and improved versatility of a pretrained image processing model. Please see CHEN (US 20250086462 A1), Paragraph [0035], and STEINER (US 20250117893 A1), Paragraph [0059-0063].
Regarding claim 17, CHEN and GUO in view of STEINER teach the method of claim 15,
CHEN further teaches wherein said generating the extended dataset is task-agnostic (Fig. 13, Paragraph [0120] - CHEN discloses unlabeled distillation input data 1302 was first used when pretraining a model in a task-agnostic fashion and then again reused after performing fine-tuning of the model to distill the fine-tuned model to a student specialized in one or more tasks.).
Regarding claim 19, CHEN in view of GUO teach the method of claim 1,
CHEN in view of GUO fail to explicitly teach wherein said training the teacher model and training the lightweight student network model do not include finetuning the pretrained network model.
However, STEINER explicitly teaches wherein said training the teacher model and training the lightweight student network model do not include finetuning the pretrained network model (Fig. 25, Paragraph [0061] - STEINER discloses a pretrained image processing model can perform a downstream task without the need for fine-tuning. Paragraph [0210] - STEINER further discloses fine-tuning can be omitted, for example, if a pre-trained model as satisfactory performance, if the model was already fine-tuned, or if other tuning approaches are preferred.).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date the claimed invention was made to combine the teachings of CHEN in view of GUO of having a computer-implemented method for training a lightweight student network model for performing a task, the method comprising: augmenting a pretrained network model with a task-specific prediction head to provide a teacher model, the pretrained network model being larger than the lightweight student network model; training the teacher model on the task, said training the teacher model using a task objective; training the lightweight student network model jointly using at least the task objective and a distillation objective, the distillation objective comprising a similarity between predictions from the lightweight student model and predictions from the trained teacher model; with the teachings of STEINER of having wherein said training the teacher model and training the lightweight student network model do not include finetuning the pretrained network model.
Wherein CHEN’s computer-implemented method wherein said training the teacher model and training the lightweight student network model do not include finetuning the pretrained network model.
The motivation behind this modification would have been to provide an enhanced method for training a lightweight student network model that provides computational efficiency and improved accuracy of an image processing model, since both CHEN and STEINER relate to training a lightweight model, wherein CHEN discloses distillation training is performed based on reusing the unlabeled pretraining data to distill the network to a student network that performs one or more specialized tasks wherein such semi-supervised contrastive learning improves accuracy and computational efficiency over previously known methods, and STEINER relates to self-supervised training of machine-learned image processing models; benefits of the present disclosure can include improved accuracy or quality of the image processing model; reduced energy consumption; reduced training costs and training time; optimizing with limited or no downstream task data; greater data efficiency in downstream fine-tuning; and improved versatility of a pretrained image processing model. Please see CHEN (US 20250086462 A1), Paragraph [0035], and STEINER (US 20250117893 A1), Paragraph [0059-0063].
Regarding claim 22, CHEN in view of GUO teach the method of claim 1,
CHEN further teaches wherein said probing uses a dataset including original data and corresponding labels (Fig. 11, Paragraph [0115] - CHEN discloses the computing system performs fine-tuning of the image classification model based on a set of labeled training data.);
wherein said training the lightweight student model uses an extended dataset (Fig. 13, Paragraph [0124] - CHEN discloses a fine-tuned classification model 1304 provides labels for training student network 1314 and distillation loss 1318 may be minimized. See also Equation [0004]. Paragraph [0126] - CHEN discloses when distillation training involves labeled training data, distillation loss 1318 may be combined with ground-truth labeled examples using a weighted combination. See also Equation [0005].),
the extended dataset comprising: the original data (Fig. 13, Paragraph [0126] - CHEN discloses when distillation training involves labeled training data, distillation loss 1318 may be combined with ground-truth labeled examples using a weighted combination. See also Equation [0005].);
Although CHEN further teaches wherein the original data comprises images (Fig. 13, Paragraph [0109] - CHEN discloses training data generally may include any type of visual and non-visual data including, but not limited to, images, video content, image frames of video content, audio data, textual data, geospatial data, sensor data, etc.).
CHEN in view of GUO fail to explicitly teach and synthetic data generated from the original data;
However, STEINER explicitly teaches and synthetic data generated from the original data (Fig. 1, Paragraph [0068] - STEINER discloses model trainer 108 can be or include or be implemented by a computing system configured to provide inputs to machine-learned model(s) 102 and evaluate outputs from machine-learned model(s) 102 to provide model update(s) 110. Paragraph [0070] - STEINER discloses model trainer 108 can access a training dataset that contains image(s) 104. Images 104 can include a mixture of image attributes. An example attribute includes magnification or zoom. Paragraph [0073] - STEINER discloses a mixture of magnifications can be synthesized. For example, a mixture of magnifications can be synthesized by transforming an original image (e.g., an image having a native magnification) to emulate various different recording magnifications. See also Paragraph [0302].);
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date the claimed invention was made to combine the teachings of CHEN in view of GUO of having a computer-implemented method for training a lightweight student network model for performing a task, the method comprising: augmenting a pretrained network model with a task-specific prediction head to provide a teacher model, the pretrained network model being larger than the lightweight student network model; training the teacher model on the task, said training the teacher model using a task objective; training the lightweight student network model jointly using at least the task objective and a distillation objective, the distillation objective comprising a similarity between predictions from the lightweight student model and predictions from the trained teacher model; with the teachings of STEINER of having and synthetic data generated from the original data.
Wherein CHEN’s computer-implemented method wherein the original data comprises images and synthetic data generated from the original data.
The motivation behind this modification would have been to provide an enhanced method for training a lightweight student network model that provides computational efficiency and improved accuracy of an image processing model, since both CHEN and STEINER relate to training a lightweight model, wherein CHEN discloses distillation training is performed based on reusing the unlabeled pretraining data to distill the network to a student network that performs one or more specialized tasks wherein such semi-supervised contrastive learning improves accuracy and computational efficiency over previously known methods, and STEINER relates to self-supervised training of machine-learned image processing models; benefits of the present disclosure can include improved accuracy or quality of the image processing model; reduced energy consumption; reduced training costs and training time; optimizing with limited or no downstream task data; greater data efficiency in downstream fine-tuning; and improved versatility of a pretrained image processing model. Please see CHEN (US 20250086462 A1), Paragraph [0035], and STEINER (US 20250117893 A1), Paragraph [0059-0063].
Regarding claim 24, CHEN in view of GUO teach the method of claim 1,
CHEN in view of GUO fail to explicitly teach wherein the pretrained network model and the lightweight student network model comprise vision transformers.
However, STEINER explicitly teaches wherein the pretrained network model (Fig. 25, Paragraph [0191] - STEINER discloses workbench 15 can implement a pre-training pipeline 17-2 to pre-train development model 16.)
and the lightweight student network model (Fig. 25, Paragraph [0205] - STEINER discloses obtain a lightweight model for running in resource-constrained environments, a smaller model can be a “student model” that learns to imitate development model 16 as a “teacher model.” In this manner, for instance, the investment in learning the parameters and configurations of development model 16 can be efficiently transferred to a smaller model for more efficient inference.)
comprise vision transformers (Fig. 1, Paragraph [0153] - STEINER discloses features described herein with respect to machine-learned model 1 (e.g., in the context of FIGS. 21 to 30) are to be understood as also describing various example attributes of implementations of machine-learned model 102, as machine-learned model 102 can be implemented using various configurations of machine-learned model 1. Paragraph [0089] - STEINER discloses implementations of machine-learned model 102 provided herein focused on utilization of a Vision Transformer (ViT) backbone (ViT-S and ViT-B).).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date the claimed invention was made to combine the teachings of CHEN in view of GUO of having a computer-implemented method for training a lightweight student network model for performing a task, the method comprising: augmenting a pretrained network model with a task-specific prediction head to provide a teacher model, the pretrained network model being larger than the lightweight student network model; training the teacher model on the task, said training the teacher model using a task objective; training the lightweight student network model jointly using at least the task objective and a distillation objective, the distillation objective comprising a similarity between predictions from the lightweight student model and predictions from the trained teacher model; with the teachings of STEINER of having wherein the pretrained network model and the lightweight student network model comprise vision transformers.
Wherein CHEN’s computer-implemented method wherein the pretrained network model and the lightweight student network model comprise vision transformers.
The motivation behind this modification would have been to provide an enhanced method for training a lightweight student network model that provides computational efficiency and improved accuracy of an image processing model, since both CHEN and STEINER relate to training a lightweight model, wherein CHEN discloses distillation training is performed based on reusing the unlabeled pretraining data to distill the network to a student network that performs one or more specialized tasks wherein such semi-supervised contrastive learning improves accuracy and computational efficiency over previously known methods, and STEINER relates to self-supervised training of machine-learned image processing models; benefits of the present disclosure can include improved accuracy or quality of the image processing model; reduced energy consumption; reduced training costs and training time; optimizing with limited or no downstream task data; greater data efficiency in downstream fine-tuning; and improved versatility of a pretrained image processing model. Please see CHEN (US 20250086462 A1), Paragraph [0035], and STEINER (US 20250117893 A1), Paragraph [0059-0063].
Regarding claim 26, CHEN teaches a processor-based system for training a lightweight student network model for performing a task (Fig. 14A, Paragraph [0127] - CHEN discloses FIG. 14A illustrates one example computing system that can be used to implement the present disclosure. Other computing systems can be used as well. Paragraph [0090] - CHEN further discloses user computing device 102 includes one or more processors 112 and a memory 114. The one or more processors 112 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, a FPGA, a controller, a microcontroller, etc.) and can be one processor or a plurality of processors that are operatively connected.), the system comprising:
a teacher model (Fig. 13, 1304 called classification model, Paragraph [0123]) for connecting to the lightweight student network model (Fig. 13, #1314 called student network, Paragraph [0123]) and having teacher model parameters (Fig. 13, Paragraph [0123] - CHEN discloses classification model 1304 is used to train a student network 1314 that is more specialized for a target task. For example, fine-tuned classification model 1304 is used when performing distillation training to distill the model to student network 1314 comprising a relatively smaller number of parameters relative to image classification model. As such, student network 1314 generally is lightweight and better suited to be deployed to client computing devices with limited local computing resources.), said teacher model comprising:
a backbone model that is larger than the lightweight student network model and having backbone model parameters (Fig. 13, Paragraph [0122] - CHEN discloses classification model 1304 includes a fine-tuned network 1306, such as a network (e.g., a fine-tuned base encoder neural network, large convolutional neural network, etc.) that was first pretrained with contrastive learning using unlabeled training data and later fine-tuned based on a relatively small set of labeled training data. Paragraph [0123] - CHEN further discloses fine-tuned classification model 1304 is used when performing distillation training to distill the model to student network 1314 comprising a relatively smaller number of parameters relative to image classification model.);
and a task-specific prediction head connected downstream of said backbone model and having prediction head parameters (Fig. 13, Paragraph [0122] - CHEN discloses classification head 1310 may receive and process one or more representations to generate classification output 1312, such as a classification prediction, detection prediction, recognition prediction, segmentation prediction, and/or other types of predictions and prediction tasks.);
using at least on the task objective and a distillation objective (Fig. 13, Paragraph [0123] - CHEN discloses classification model 1304 is used to train a student network 1314 that is more specialized for a target task. For example, fine-tuned classification model 1304 is used when performing distillation training to distill the model to student network 1314 comprising a relatively smaller number of parameters relative to image classification model.),
and a training loss determination module (Fig. 14A, #160 called model trainer, Paragraph [0100] - CHEN discloses training computing system 150 can include a model trainer 160 that trains the machine-learned models 120 and/or 140 stored at the user computing device 102 and/or the server computing system 130 using various training or learning techniques, such as, for example, backwards propagation of errors. For example, a loss function can be backpropagated through the model(s) to update one or more parameters of the model(s) (e.g., based on a gradient of the loss function). See also Paragraph [0115].) configured to:
determine a supervised task loss for the task objective over original data (Fig. 12, Paragraph [0117] - CHEN discloses fine-tuning of classification model 1204 may be performed by adjusting various parameters based on labeled fine-tuning input data 1202, for example, using a supervised cross-entropy loss or other type of loss function (not shown), allowing classification model 1204 to slightly adjust internal representations for one or more specific tasks.),
the original data having associated labels (Fig. 12, Paragraph [0116] - CHEN discloses classification model 1204 obtains or otherwise receives a set of labeled, fine-tuning input data 1202.);
and determine a distillation loss for the distillation objective over extended data (Fig. 13, Paragraph [0124] - CHEN discloses a fine-tuned classification model 1304 provides labels for training student network 1314 and distillation loss 1318 may be minimized. See also Equation [0004].),
ALTHOUGH CHEN further teaches the extended data comprising the original data and additional data generated using the original data (Fig. 13, Paragraph [0126] - CHEN discloses when distillation training involves labeled training data, distillation loss 1318 may be combined with ground-truth labeled examples using a weighted combination. See also Equation [0005].).
CHEN fails to explicitly teach the distillation objective comprising a similarity between predictions from the lightweight student model and predictions from the teacher model;
However, GUO explicitly teaches the distillation objective comprising a similarity between predictions from the lightweight student model and predictions from the teacher model (Fig. 2, Paragraph [0025] - GUO discloses distillation based on relationship, that is, considering the differences between the teacher model and the student model in metrics such as similarity for different samples, thereby guiding the training of the student model.);
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date the claimed invention was made to combine the teachings of CHEN of having processor-based system for training a lightweight student network model for performing a task, the system comprising: a teacher model for connecting to the lightweight student network model and having teacher model parameters, said teacher model comprising: a backbone model that is larger than the lightweight student network model and having backbone model parameters; and a task-specific prediction head connected downstream of said backbone model and having prediction head parameters; using at least on the task objective and a distillation objective, and a training loss determination module configured to: determine a supervised task loss for the task objective over original data, the original data having associated labels; and determine a distillation loss for the distillation objective over extended data, the extended data comprising the original data and additional data generated using the original data, with the teachings of GUO of having the distillation objective comprising a similarity between predictions from the lightweight student model and predictions from the teacher model.
Wherein CHEN’s processor-based system wherein having the distillation objective comprising a similarity between predictions from the lightweight student model and predictions from the teacher model.
The motivation behind this modification would have been to provide an enhanced system for training a lightweight student network model that provides computational efficiency and training precision, since both CHEN and GUO relate to training of a student model, wherein CHEN discloses distillation training is performed based on reusing the unlabeled pretraining data to distill the network to a student network that performs one or more specialized tasks wherein such semi-supervised contrastive learning improves accuracy and computational efficiency over previously known methods, and GUO discloses a lightweight model training solution, which may effectively prevent the underfitting phenomenon in the case where a knowledge distillation strategy is adopted to train the lightweight model, thereby improving training precision of the lightweight model and enhancing precision of the knowledge distillation of the lightweight model. Please see CHEN (US 20250086462 A1), Paragraph [0035], and GUO (US 20240070454 A1), Paragraph [0030].
CHEN in view of GUO fail to explicitly teach a model optimization module configured to: optimize the parameters of the prediction head, while the backbone model parameters are frozen, based on a task objective; and after said optimizing the parameters of the prediction head, optimize parameters of the lightweight student network model, while the teacher model parameters are frozen,
However, STEINER explicitly teaches a model optimization module (Fig. 25, #19 called computational optimization toolkit, Paragraph [0205]) configured to: optimize the parameters of the prediction head,
while the backbone model parameters are frozen, based on a task objective (Fig. 25, Paragraph [0152] - STEINER further discloses various portions of the machine-learned model can be “frozen” for certain training stages. For example, parameters associated with an embedding space can be “frozen” during fine-tuning (e.g., to retain information learned from a broader domain(s) than present in the fine-tuning dataset(s)). See also Paragraph [0198].);
and after said optimizing the parameters of the prediction head, optimize parameters of the lightweight student network model (Fig. 25, Paragraph [0205] - STEINER discloses tools for distillation 19-3 can provide for the training of lighter-weight models based on the knowledge encoded in development model 16. For instance, development model 16 can be a highly performant, large machine-learned model optimized using model development platform 12. To obtain a lightweight model for running in resource-constrained environments, a smaller model can be a “student model” that learns to imitate development model 16 as a “teacher model.”),
while the teacher model parameters are frozen (Fig. 25, Paragraph [0152] - STEINER further discloses various portions of the machine-learned model can be “frozen” for certain training stages. For example, parameters associated with an embedding space can be “frozen” during fine-tuning (e.g., to retain information learned from a broader domain(s) than present in the fine-tuning dataset(s)).),
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date the claimed invention was made to combine the teachings of CHEN in view of GUO of having a processor-based system for training a lightweight student network model for performing a task, the system comprising: a teacher model for connecting to the lightweight student network model and having teacher model parameters, said teacher model comprising: a backbone model that is larger than the lightweight student network model and having backbone model parameters; and a task-specific prediction head connected downstream of said backbone model and having prediction head parameters; using at least on the task objective and a distillation objective, the distillation objective comprising a similarity between predictions from the lightweight student model and predictions from the teacher model; and a training loss determination module configured to: determine a supervised task loss for the task objective over original data, the original data having associated labels; and determine a distillation loss for the distillation objective over extended data, the extended data comprising the original data and additional data generated using the original data, with the teachings of STEINER of having a model optimization module configured to: optimize the parameters of the prediction head, while the backbone model parameters are frozen, based on a task objective; and after said optimizing the parameters of the prediction head, optimize parameters of the lightweight student network model, while the teacher model parameters are frozen.
Wherein CHEN’s processor-based system wherein having a model optimization module configured to: optimize the parameters of the prediction head, while the backbone model parameters are frozen, based on a task objective; and after said optimizing the parameters of the prediction head, optimize parameters of the lightweight student network model, while the teacher model parameters are frozen.
The motivation behind this modification would have been to provide an enhanced system for training a lightweight student network model that provides computational efficiency and improved accuracy of an image processing model, since both CHEN and STEINER relate to training a lightweight model, wherein CHEN discloses distillation training is performed based on reusing the unlabeled pretraining data to distill the network to a student network that performs one or more specialized tasks wherein such semi-supervised contrastive learning improves accuracy and computational efficiency over previously known methods, and STEINER relates to self-supervised training of machine-learned image processing models; benefits of the present disclosure can include improved accuracy or quality of the image processing model; reduced energy consumption; reduced training costs and training time; optimizing with limited or no downstream task data; greater data efficiency in downstream fine-tuning; and improved versatility of a pretrained image processing model. Please see CHEN (US 20250086462 A1), Paragraph [0035], and STEINER (US 20250117893 A1), Paragraph [0059-0063].
Regarding claim 27, CHEN and GUO in view of STEINER teach the system of claim 26,
CHEN further teaches further comprising an extended dataset generation module configured to generate the additional data from the original data (Fig. 2A, Paragraph [0043] - CHEN discloses stochastic data augmentation module (shown generally at 203) that transforms any given data example (e.g., an input image x shown at 202) randomly resulting in two correlated views of the same example. Paragraph [0044] - CHEN further discloses three augmentations can be applied at 203: random cropping followed by resize back to the original size, random color distortions, and random Gaussian blur.),
CHEN in view of GUO fail to explicitly teach said extended dataset generation module comprising a synthetic data generation module for generating synthetic data from the original data.
However, STEINER explicitly teaches said extended dataset generation module comprising a synthetic data generation module for generating synthetic data from the original data (Fig. 35, Paragraph [0070] - STEINER discloses model trainer 108 can access a training dataset that contains image(s) 104. Images 104 can include a mixture of image attributes. An example attribute includes magnification or zoom. Paragraph [0073] - STEINER discloses a mixture of magnifications can be synthesized. For example, a mixture of magnifications can be synthesized by transforming an original image (e.g., an image having a native magnification) to emulate various different recording magnifications. Paragraph [0302] - STEINER discloses mapping at ratios less than unity can be implemented by upsampling, oversampling, etc., including using machine-learned models to generate an output image at an output resolution based on an input image at a different input resolution. Various machine-learned models can include transformer-based models, convolutional neural networks, diffusion-based models, etc.).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date the claimed invention was made to combine the teachings of CHEN and GUO in view of STEINER of having a processor-based system for training a lightweight student network model for performing a task, the system comprising: a teacher model for connecting to the lightweight student network model and having teacher model parameters, said teacher model comprising: a backbone model that is larger than the lightweight student network model and having backbone model parameters; and a task-specific prediction head connected downstream of said backbone model and having prediction head parameters; a model optimization module configured to: optimize the parameters of the prediction head, while the backbone model parameters are frozen, based on a task objective; and after said optimizing the parameters of the prediction head, optimize parameters of the lightweight student network model, while the teacher model parameters are frozen, using at least on the task objective and a distillation objective,, with the teachings of STEINER of having said extended dataset generation module comprising a synthetic data generation module for generating synthetic data from the original data.
Wherein CHEN’s processor-based system wherein having said extended dataset generation module comprising a synthetic data generation module for generating synthetic data from the original data.
The motivation behind this modification would have been to provide an enhanced system for training a lightweight student network model that provides computational efficiency and improved accuracy of an image processing model, since both CHEN and STEINER relate to training a lightweight model, wherein CHEN discloses distillation training is performed based on reusing the unlabeled pretraining data to distill the network to a student network that performs one or more specialized tasks wherein such semi-supervised contrastive learning improves accuracy and computational efficiency over previously known methods, and STEINER relates to self-supervised training of machine-learned image processing models; benefits of the present disclosure can include improved accuracy or quality of the image processing model; reduced energy consumption; reduced training costs and training time; optimizing with limited or no downstream task data; greater data efficiency in downstream fine-tuning; and improved versatility of a pretrained image processing model. Please see CHEN (US 20250086462 A1), Paragraph [0035], and STEINER (US 20250117893 A1), Paragraph [0059-0063].
Claims 28-31 are rejected under 35 U.S.C. 103 as being unpatentable over MALACH (US 20250371854 A1), hereinafter referenced as MALACH in view of CHEN (US 20250086462 A1), hereinafter referenced as CHEN, in further view of GUO (US 20240070454 A1), hereinafter referenced as GUO.
Regarding claim 28, MALACH teaches a computer-implemented method for performing a task of navigating or controlling an autonomous device (Fig. 3C, Paragraph [0095] - MALACH discloses the functionality associated with the blocks as discussed herein with reference to FIGS. 3B and 3C may be executed, for instance, via processing circuitry associated with any suitable computing system, e.g., the safety system 200, and/or a computing device, as described in greater detail below. The processing circuitry may implement the aspects as described herein as part of an ADAS and/or AV system of the vehicle 100.), the method comprising:
receiving a input image (Fig. 3C, Paragraph [0110] - MALACH discloses the processor may be configured to receive one or more images including a representation of a feature of interest (block 312).)
from an input device (Fig. 2, #104 called image acquisition device, Paragraph [0043]) of the autonomous device (Fig. 2, Paragraph [0043] - MALACH discloses one or more processors 102 of the safety system 200 may include processors 214A, 214B, 216, and/or 218, one or more image acquisition devices 104 such as, e.g., one or more vehicle cameras or any other suitable sensor configured to perform image acquisition over any suitable range of wavelengths (e.g., RADAR, LiDAR, etc.). Paragraph [0095] - MALACH discloses functionality associated with the blocks as discussed herein with reference to FIGS. 3B and 3C may be executed, for instance, via processing circuitry associated with any suitable computing system, e.g., the safety system 200, and/or a computing device.);
processing the input image by a lightweight student network model (Fig. 3C, Paragraph [0113] - MALACH discloses one or more images may be provided as input to both a trained supervisory neural network (a teacher network) and one or more student neural networks (block 314).)
and generating a prediction (Fig. 3C, Paragraph [0177] - MALACH discloses conditional samplers as teachers may be used for knowledge distillation, for example, by generating students which may approximate the Bayes optimal prediction.);
Although MALACH further teaches and processing the generated prediction to perform the task (Fig. 3B, Paragraph [0098] - MALACH discloses one or more student machine learning trained models may be deployed (block 376) to a suitable computing system. For example, deployment may include implementing a student machine learning trained model within a safety system 200 or other suitable computing system, e.g. one identified with a vehicle. The student machine learning trained models, once deployed, may be utilized to perform vehicle functions such as, for example, object classification, navigation action determination, etc.);
MALACH fails to explicitly teach wherein the lightweight student network model having been trained by: augmenting a pretrained network model with a task-specific prediction head to provide a teacher model, the pretrained network model being larger than the lightweight student network model; training the teacher model on the task, said training the teacher model using a task objective; and training the lightweight student network model jointly using at least the task objective and a distillation objective,
However, CHEN explicitly teaches wherein the lightweight student network model having been trained by: augmenting a pretrained network model (Fig. 13, #1304 called classification model, Paragraph [0122]) with a task-specific prediction head to provide a teacher model (Fig. 13, Paragraph [0122] - CHEN discloses classification model 1304 also includes one or more projection head layer(s) 1308, for example, originally from a projection head neural network pretrained with contrastive learning using unlabeled training data, where the specific projection head layer(s) were preserved after the pretraining and later fine-tuned based on the set of label training data.),
the pretrained network model (Fig. 13, #1304 called classification model, Paragraph [0122]) being larger than the lightweight student network model (Fig. 13, Paragraph [0123] - CHEN discloses classification model 1304 is used to train a student network 1314 that is more specialized for a target task. For example, fine-tuned classification model 1304 is used when performing distillation training to distill the model to student network 1314 comprising a relatively smaller number of parameters relative to image classification model. As such, student network 1314 generally is lightweight and better suited to be deployed to client computing devices with limited local computing resources.);
training the teacher model on the task (Fig. 13, Paragraph [0122] - CHEN discloses classification model 1304 also includes one or more projection head layer(s) 1308, for example, originally from a projection head neural network pretrained with contrastive learning using unlabeled training data, where the specific projection head layer(s) were preserved after the pretraining and later fine-tuned based on the set of label training data.),
said training the teacher model using a task objective (Fig. 13, Paragraph [0122] - CHEN discloses classification head 1310 may receive and process one or more representations to generate classification output 1312, such as a classification prediction, detection prediction, recognition prediction, segmentation prediction, and/or other types of predictions and prediction tasks.);
and training the lightweight student network model jointly using at least the task objective and a distillation objective (Fig. 13, Paragraph [0123] - CHEN discloses classification model 1304 is used to train a student network 1314 that is more specialized for a target task. For example, fine-tuned classification model 1304 is used when performing distillation training to distill the model to student network 1314 comprising a relatively smaller number of parameters relative to image classification model.),
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date the claimed invention was made to combine the teachings of MALACH of having computer-implemented method for performing a task of navigating or controlling an autonomous device, the method comprising: receiving a input image from an input device of the autonomous device; processing the input image by a lightweight student network model and generating a prediction; and processing the generated prediction to perform the task; with the teachings of CHEN of having wherein the lightweight student network model having been trained by: augmenting a pretrained network model with a task-specific prediction head to provide a teacher model, the pretrained network model being larger than the lightweight student network model; training the teacher model on the task, said training the teacher model using a task objective; and training the lightweight student network model jointly using at least the task objective and a distillation objective.
Wherein MALACH’s computer-implemented method wherein the lightweight student network model having been trained by: augmenting a pretrained network model with a task-specific prediction head to provide a teacher model, the pretrained network model being larger than the lightweight student network model; training the teacher model on the task, said training the teacher model using a task objective; and training the lightweight student network model jointly using at least the task objective and a distillation objective.
The motivation behind this modification would have been to provide an enhanced system for training a lightweight student network model that provides improved performance and computational efficiency, since both MALACH and CHEN relate to training of a student model, wherein MALACH discloses techniques for improving performance of student trained model and to achieve those improvements more efficiently than by typical means (e.g., proceeding through an available dataset, adjustment the model in training, etc.) and CHEN discloses distillation training is performed based on reusing the unlabeled pretraining data to distill the network to a student network that performs one or more specialized tasks wherein such semi-supervised contrastive learning improves accuracy and computational efficiency over previously known methods. Please see MALACH (US 20250371854 A1), Paragraph [0231], and CHEN (US 20250086462 A1), Paragraph [0035].
MALACH in view of CHEN fail to explicitly teach the distillation objective comprising a similarity between predictions from the lightweight student model and predictions from the trained teacher model.
However, GUO explicitly teaches the distillation objective comprising a similarity between predictions from the lightweight student model and predictions from the trained teacher model (Fig. 2, Paragraph [0025] - GUO discloses distillation based on relationship, that is, considering the differences between the teacher model and the student model in metrics such as similarity for different samples, thereby guiding the training of the student model.).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date the claimed invention was made to combine the teachings of MALACH in view of CHEN of having computer-implemented method for performing a task of navigating or controlling an autonomous device, the method comprising: receiving a input image from an input device of the autonomous device; processing the input image by a lightweight student network model and generating a prediction; and processing the generated prediction to perform the task; the lightweight student network model having been trained by: augmenting a pretrained network model with a task-specific prediction head to provide a teacher model, the pretrained network model being larger than the lightweight student network model; training the teacher model on the task, said training the teacher model using a task objective; and training the lightweight student network model jointly using at least the task objective and a distillation objective, with the teachings of GUO of having the distillation objective comprising a similarity between predictions from the lightweight student model and predictions from the trained teacher model.
Wherein MALACH’s computer-implemented method wherein having the distillation objective comprising a similarity between predictions from the lightweight student model and predictions from the trained teacher model.
The motivation behind this modification would have been to provide an enhanced system for training a lightweight student network model that provides improved performance and training precision, since both MALACH and GUO relate to training of a student model, wherein MALACH discloses techniques for improving performance of student trained model and to achieve those improvements more efficiently than by typical means (e.g., proceeding through an available dataset, adjustment the model in training, etc.) and GUO discloses a lightweight model training solution, which may effectively prevent the underfitting phenomenon in the case where a knowledge distillation strategy is adopted to train the lightweight model, thereby improving training precision of the lightweight model and enhancing precision of the knowledge distillation of the lightweight model. Please see MALACH (US 20250371854 A1), Paragraph [0231], and GUO (US 20240070454 A1), Paragraph [0030].
Regarding claim 29, MALACH and CHEN in view of GUO teach the method of claim 28,
MALACH fails to explicitly teach wherein the task further comprises an image processing task that comprises one or more of image classification, object detection, and/or semantic segmentation.
However, CHEN explicitly teaches wherein the task further comprises an image processing task that comprises one or more of image classification, object detection, and/or semantic segmentation (Fig. 13, Paragraph [0121] - CHEN discloses classification model 1304 may be an image classification model or any other type of classification model. etc. Paragraph [0122] - CHEN further discloses classification head 1310 may receive and process one or more representations to generate classification output 1312, such as a classification prediction [wherein classification prediction is image classification], detection prediction [wherein detection prediction is object detection], recognition prediction, segmentation prediction [wherein segmentation prediction is semantic segmentation], and/or other types of predictions and prediction tasks.).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date the claimed invention was made to combine the teachings of MALACH and CHEN in view of GUO of having computer-implemented method for performing a task of navigating or controlling an autonomous device, the method comprising: receiving a input image from an input device of the autonomous device; processing the input image by a lightweight student network model and generating a prediction; and processing the generated prediction to perform the task; wherein the lightweight student network model having been trained by: augmenting a pretrained network model with a task-specific prediction head to provide a teacher model, the pretrained network model being larger than the lightweight student network model; training the teacher model on the task, said training the teacher model using a task objective; and training the lightweight student network model jointly using at least the task objective and a distillation objective, with the teachings of CHEN of having wherein the task further comprises an image processing task that comprises one or more of image classification, object detection, and/or semantic segmentation.
Wherein MALACH’s computer-implemented method wherein the task further comprises an image processing task that comprises one or more of image classification, object detection, and/or semantic segmentation.
The motivation behind this modification would have been to provide an enhanced system for training a lightweight student network model that provides improved performance and computational efficiency, since both MALACH and CHEN relate to training of a student model, wherein MALACH discloses techniques for improving performance of student trained model and to achieve those improvements more efficiently than by typical means (e.g., proceeding through an available dataset, adjustment the model in training, etc.) and CHEN discloses distillation training is performed based on reusing the unlabeled pretraining data to distill the network to a student network that performs one or more specialized tasks wherein such semi-supervised contrastive learning improves accuracy and computational efficiency over previously known methods. Please see MALACH (US 20250371854 A1), Paragraph [0231], and CHEN (US 20250086462 A1), Paragraph [0035].
Regarding claim 30, MALACH teaches an autonomous device (Fig. 1, #100 called vehicle, Paragraph [0037] - MALACH discloses vehicle 100 may include any type of vehicle (e.g., a road vehicle) and may be an autonomous vehicle (AV). Paragraph [0038] - MALACH discloses safety system 200 may be implemented with vehicle 100 as part of any suitable type of autonomous or driving assistance control system.), comprising:
a memory (Fig. 2, #202 called one or more memories, Paragraph [0043]) for storing a lightweight student network model (Fig. 2, Paragraph [0054] - MALACH discloses relevant memory accessed by the one or more processors 214A, 214B, 216, 218 (e.g. the one or more memories 202) may also store one or more databases and image processing software, as well as a trained system, such as a neural network, or a deep neural network, for example, that may be utilized to perform the tasks in accordance with any of the aspects as discussed herein. See also Paragraph [0056].);
an input device (Fig. 2, #104 called image acquisition device, Paragraph [0043]) for receiving an input image (Fig. 2, Paragraph [0043] - MALACH discloses one or more processors 102 of the safety system 200 may include processors 214A, 214B, 216, and/or 218, one or more image acquisition devices 104 such as, e.g., one or more vehicle cameras or any other suitable sensor configured to perform image acquisition over any suitable range of wavelengths (e.g., RADAR, LiDAR, etc.).);
a processor (Fig. 1, #102 called processor, Paragraph [0095]) for processing the input image by the lightweight student network model (Fig. 3C, Paragraph [0113] - MALACH discloses one or more images may be provided as input to both a trained supervisory neural network (a teacher network) and one or more student neural networks (block 314).)
and generating a prediction (Fig. 3C, Paragraph [0177] - MALACH discloses conditional samplers as teachers may be used for knowledge distillation, for example, by generating students which may approximate the Bayes optimal prediction.),
Although MALACH further teaches and for processing the generated prediction to perform a task of navigating or controlling the autonomous device (Fig. 3B, Paragraph [0098] - MALACH discloses one or more student machine learning trained models may be deployed (block 376) to a suitable computing system. For example, deployment may include implementing a student machine learning trained model within a safety system 200 or other suitable computing system, e.g. one identified with a vehicle. The student machine learning trained models, once deployed, may be utilized to perform vehicle functions such as, for example, object classification, navigation action determination, etc.);
MALACH fails to explicitly teach wherein the lightweight student network model having been trained by: augmenting a pretrained network model with a task-specific prediction head to provide a teacher model, the pretrained network model being larger than the lightweight student network model; training the teacher model on the task, said training the teacher model using a task objective; and training the lightweight student network model jointly using at least the task objective and a distillation objective,
However, CHEN explicitly teaches wherein the lightweight student network model having been trained by: augmenting a pretrained network model (Fig. 13, #1304 called classification model, Paragraph [0122]) with a task-specific prediction head to provide a teacher model (Fig. 13, Paragraph [0122] - CHEN discloses classification model 1304 also includes one or more projection head layer(s) 1308, for example, originally from a projection head neural network pretrained with contrastive learning using unlabeled training data, where the specific projection head layer(s) were preserved after the pretraining and later fine-tuned based on the set of label training data.),
the pretrained network model (Fig. 13, #1304 called classification model, Paragraph [0122]) being larger than the lightweight student network model (Fig. 13, Paragraph [0123] - CHEN discloses classification model 1304 is used to train a student network 1314 that is more specialized for a target task. For example, fine-tuned classification model 1304 is used when performing distillation training to distill the model to student network 1314 comprising a relatively smaller number of parameters relative to image classification model. As such, student network 1314 generally is lightweight and better suited to be deployed to client computing devices with limited local computing resources.);
training the teacher model on the task (Fig. 13, Paragraph [0122] - CHEN discloses classification model 1304 also includes one or more projection head layer(s) 1308, for example, originally from a projection head neural network pretrained with contrastive learning using unlabeled training data, where the specific projection head layer(s) were preserved after the pretraining and later fine-tuned based on the set of label training data.),
said training the teacher model using a task objective (Fig. 13, Paragraph [0122] - CHEN discloses classification head 1310 may receive and process one or more representations to generate classification output 1312, such as a classification prediction, detection prediction, recognition prediction, segmentation prediction, and/or other types of predictions and prediction tasks.);
and training the lightweight student network model jointly using at least the task objective and a distillation objective (Fig. 13, Paragraph [0123] - CHEN discloses classification model 1304 is used to train a student network 1314 that is more specialized for a target task. For example, fine-tuned classification model 1304 is used when performing distillation training to distill the model to student network 1314 comprising a relatively smaller number of parameters relative to image classification model.),
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date the claimed invention was made to combine the teachings of MALACH of having an autonomous device, comprising: a memory for storing a lightweight student network model; an input device for receiving an input image; a processor for processing the input image by the lightweight student network model and generating a prediction, and for processing the generated prediction to perform a task of navigating or controlling the autonomous device; with the teachings of CHEN of having wherein the lightweight student network model having been trained by: augmenting a pretrained network model with a task-specific prediction head to provide a teacher model, the pretrained network model being larger than the lightweight student network model; training the teacher model on the task, said training the teacher model using a task objective; and training the lightweight student network model jointly using at least the task objective and a distillation objective.
Wherein MALACH’s autonomous device wherein the lightweight student network model having been trained by: augmenting a pretrained network model with a task-specific prediction head to provide a teacher model, the pretrained network model being larger than the lightweight student network model; training the teacher model on the task, said training the teacher model using a task objective; and training the lightweight student network model jointly using at least the task objective and a distillation objective.
The motivation behind this modification would have been to provide an enhanced system for training a lightweight student network model that provides improved performance and computational efficiency, since both MALACH and CHEN relate to training of a student model, wherein MALACH discloses techniques for improving performance of student trained model and to achieve those improvements more efficiently than by typical means (e.g., proceeding through an available dataset, adjustment the model in training, etc.) and CHEN discloses distillation training is performed based on reusing the unlabeled pretraining data to distill the network to a student network that performs one or more specialized tasks wherein such semi-supervised contrastive learning improves accuracy and computational efficiency over previously known methods. Please see MALACH (US 20250371854 A1), Paragraph [0231], and CHEN (US 20250086462 A1), Paragraph [0035].
MALACH in view of CHEN fail to explicitly teach the distillation objective comprising a similarity between predictions from the lightweight student model and predictions from the trained teacher model.
However, GUO explicitly teaches the distillation objective comprising a similarity between predictions from the lightweight student model and predictions from the trained teacher model (Fig. 2, Paragraph [0025] - GUO discloses distillation based on relationship, that is, considering the differences between the teacher model and the student model in metrics such as similarity for different samples, thereby guiding the training of the student model.).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date the claimed invention was made to combine the teachings of MALACH in view of CHEN of having an autonomous device, comprising: a memory for storing a lightweight student network model; an input device for receiving an input image; a processor for processing the input image by the lightweight student network model and generating a prediction, and for processing the generated prediction to perform a task of navigating or controlling the autonomous device; wherein the lightweight student network model having been trained by: augmenting a pretrained network model with a task-specific prediction head to provide a teacher model, the pretrained network model being larger than the lightweight student network model; training the teacher model on the task, said training the teacher model using a task objective; and training the lightweight student network model jointly using at least the task objective and a distillation objective, with the teachings of GUO of having the distillation objective comprising a similarity between predictions from the lightweight student model and predictions from the trained teacher model.
Wherein MALACH’s autonomous device wherein having the distillation objective comprising a similarity between predictions from the lightweight student model and predictions from the trained teacher model.
The motivation behind this modification would have been to provide an enhanced system for training a lightweight student network model that provides improved performance and training precision, since both MALACH and GUO relate to training of a student model, wherein MALACH discloses techniques for improving performance of student trained model and to achieve those improvements more efficiently than by typical means (e.g., proceeding through an available dataset, adjustment the model in training, etc.) and GUO discloses a lightweight model training solution, which may effectively prevent the underfitting phenomenon in the case where a knowledge distillation strategy is adopted to train the lightweight model, thereby improving training precision of the lightweight model and enhancing precision of the knowledge distillation of the lightweight model. Please see MALACH (US 20250371854 A1), Paragraph [0231], and GUO (US 20240070454 A1), Paragraph [0030].
Regarding claim 31, MALACH and CHEN in view of GUO teach the method of claim 30,
MALACH fails to explicitly teach wherein the task further comprises an image processing task that comprises one or more of image classification, object detection, and/or semantic segmentation.
However, CHEN explicitly teaches wherein the task further comprises an image processing task that comprises one or more of image classification, object detection, and/or semantic segmentation (Fig. 13, Paragraph [0121] - CHEN discloses classification model 1304 may be an image classification model or any other type of classification model. etc. Paragraph [0122] - CHEN further discloses classification head 1310 may receive and process one or more representations to generate classification output 1312, such as a classification prediction [wherein classification prediction is image classification], detection prediction [wherein detection prediction is object detection], recognition prediction, segmentation prediction [wherein segmentation prediction is semantic segmentation], and/or other types of predictions and prediction tasks.).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date the claimed invention was made to combine the teachings of MALACH and CHEN in view of GUO of having an autonomous device, comprising: a memory for storing a lightweight student network model; an input device for receiving an input image; a processor for processing the input image by the lightweight student network model and generating a prediction, and for processing the generated prediction to perform a task of navigating or controlling the autonomous device; wherein the lightweight student network model having been trained by: augmenting a pretrained network model with a task-specific prediction head to provide a teacher model, the pretrained network model being larger than the lightweight student network model; training the teacher model on the task, said training the teacher model using a task objective; and training the lightweight student network model jointly using at least the task objective and a distillation objective, with the teachings of CHEN of having wherein the task further comprises an image processing task that comprises one or more of image classification, object detection, and/or semantic segmentation.
Wherein MALACH’s autonomous device wherein the task further comprises an image processing task that comprises one or more of image classification, object detection, and/or semantic segmentation.
The motivation behind this modification would have been to provide an enhanced system for training a lightweight student network model that provides improved performance and computational efficiency, since both MALACH and CHEN relate to training of a student model, wherein MALACH discloses techniques for improving performance of student trained model and to achieve those improvements more efficiently than by typical means (e.g., proceeding through an available dataset, adjustment the model in training, etc.) and CHEN discloses distillation training is performed based on reusing the unlabeled pretraining data to distill the network to a student network that performs one or more specialized tasks wherein such semi-supervised contrastive learning improves accuracy and computational efficiency over previously known methods. Please see MALACH (US 20250371854 A1), Paragraph [0231], and CHEN (US 20250086462 A1), Paragraph [0035].
Conclusion
Listed below are the prior arts made of record and not relied upon but are considered pertinent to applicant’s disclosure.
VIJAYARAGHAVAN et al. (US 20260030511 A1) - A method, computer system, and a computer program product for data-free knowledge amalgamation are provided. Multiple pre-trained teacher machine learning models are obtained. Each is trained on a respective different set of training data. Pseudo-data samples that mimic original training data of the teacher models are generated. A block-wise amalgamation with a self-regulative strategy to integrate knowledge from the multiple teacher models is implemented by inputting the pseudo-data samples into the teacher models and into a student machine learning model. The implementing also includes aligning intermediate representations of the student model with a unified representation capturing relevant features from the teacher models..… Fig. 1, Abstract.
QIN et al. (US 20260024363 A1) - A semantic segmentation model training method and apparatus, an electronic device and a storage medium are provided. The semantic segmentation model training method includes: acquiring a teacher semantic segmentation model that is pre-trained, the teacher semantic segmentation model including a first teacher network and a second teacher network, the first teacher network having structural characteristics of low depth and high width, and the second teacher network having structural characteristics of high depth and low width; processing a sample image based on the teacher semantic segmentation model to obtain a first segmentation map and a second segmentation map; and training a student semantic segmentation model that is lightweight according to the sample image, the first segmentation map and the second segmentation map, so as to obtain a target semantic segmentation model..… Fig. 1, Abstract.
YAO et al. (US 20250252318 A1) - A neural network can be trained through knowledge distillation. A support neural network is generated based on a target neural network. The support neural network is a teacher model, and the target neural network is a student model. The support neural network may have same layers as the target neural networks. Some or all layers of the support neural network may be connected to facilitate data transfer between these layers. The support neural network and target neural network are merged into a merged network. The merged network is trained. At least one layer in the support neural network is connected to a layer in the target neural network to facilitate data transfer from the target neural network to the support neural network during the training. After the training, the target neural network is separated from the merged network and can be used to perform machine learning tasks...… Fig. 1, Abstract.
LIU et al. (US 20210142164 A1) - Systems and methods are provided that employ knowledge distillation under a multi-task learning setting. In some embodiments, the systems and methods are implemented with a larger teacher model and a smaller student model, each of which comprise one or more shared layers and a plurality of task layers for performing multiple tasks. During training of the teacher model, its shared layers are initialized, and then the teacher model is multi-task refined. The teacher model predicts teacher logits. During training of the student model, its shared layers are initialized. Knowledge distillation is employed to transfer knowledge from the teacher model to the student model by the student model updating its shared layers and task layers, for example, according to the teacher logits of the teacher model. Other features are also provided...… Fig. 1, Abstract.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to BEZAWIT N SHIMELES whose telephone number is (571)272-7663. The examiner can normally be reached M-F 7:30am-5pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Chineyere Wills-Burns can be reached at (571) 272-9752. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/BEZAWIT NOLAWI SHIMELES/Examiner, Art Unit 2673
/CHINEYERE WILLS-BURNS/Supervisory Patent Examiner, Art Unit 2673