DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
This office action is in responsive to communication(s): original application filed on 12/22/2022, said application claims a priority filing date of 07/29/2022. Claims 1-26 are pending. Claims 1, 16, and 25 are independent.
Drawings
The drawings are objected to as failing to comply with 37 CFR 1.84(p)(5) because they include the following reference character(s) not mentioned in the description: 220, 220A, 220B, and 220C in FIG. 2C. Corrected drawing sheets in compliance with 37 CFR 1.121(d), or amendment to the specification to add the reference character(s) in the description in compliance with 37 CFR 1.121(b) are required in reply to the Office action to avoid abandonment of the application. Any amended replacement drawing sheet should include all of the figures appearing on the immediate prior version of the sheet, even if only one figure is being amended. Each drawing sheet submitted after the filing date of an application must be labeled in the top margin as either “Replacement Sheet” or “New Sheet” pursuant to 37 CFR 1.121(d). If the changes are not accepted by the examiner, the applicant will be notified and informed of any required corrective action in the next Office action. The objection to the drawings will not be held in abeyance.
The drawings are objected to as failing to comply with 37 CFR 1.84(p)(4) because (1) reference characters "402E" in FIG. 4 and "404E" in ¶ [0072] have both been used to designate "model portions"; and (2) reference characters "402A" in ¶ [0072] and "404A" in FIG.4 have both been used to designate "compressed model portions". Corrected drawing sheets in compliance with 37 CFR 1.121(d) are required in reply to the Office action to avoid abandonment of the application. Any amended replacement drawing sheet should include all of the figures appearing on the immediate prior version of the sheet, even if only one figure is being amended. Each drawing sheet submitted after the filing date of an application must be labeled in the top margin as either “Replacement Sheet” or “New Sheet” pursuant to 37 CFR 1.121(d). If the changes are not accepted by the examiner, the applicant will be notified and informed of any required corrective action in the next Office action. The objection to the drawings will not be held in abeyance.
The drawings are objected to as failing to comply with 37 CFR 1.84(p)(4) because (1) reference character “404E” has been used to designate both "model portions" in ¶ [0072] and "compressed model portions" in FIG. 4 and ¶ [0072]; and (2) reference character “402A” has been used to designate both "model portions" in FIG. 4 and ¶ [0072] and "compressed model portions" in ¶ [0072]. Corrected drawing sheets in compliance with 37 CFR 1.121(d) are required in reply to the Office action to avoid abandonment of the application. Any amended replacement drawing sheet should include all of the figures appearing on the immediate prior version of the sheet, even if only one figure is being amended. Each drawing sheet submitted after the filing date of an application must be labeled in the top margin as either “Replacement Sheet” or “New Sheet” pursuant to 37 CFR 1.121(d). If the changes are not accepted by the examiner, the applicant will be notified and informed of any required corrective action in the next Office action. The objection to the drawings will not be held in abeyance.
Specification
The disclosure is objected to because of the following informalities:
in ¶ [0072], "… The set of candidate compression schemes can include schemes for compressing model portions 402A, 402C, and 404E. The compression schemes can be applied to obtain compressed model portions 402A, 404C, and 404E. As depicted, each of the compressed model portions 402A, 404C, and 404E can be a different size due to the different compression schemes …" appears to be "… The set of candidate compression schemes can include schemes for compressing model portions 402A, 402C, and 402E. The compression schemes can be applied to obtain compressed model portions 404A, 404C, and 404E. As depicted, each of the compressed model portions 404A, 404C, and 404E can be a different size due to the different compression schemes …".
Appropriate correction is required.
Claim Objections
Claims 1, 3-7, 15-16, and 18-22 are objected to because of the following informalities:
in Claim 1, lines 6-7; and Claim 16, lines 9-10, "… respectively select one or more candidate compression schemes from the one or more sets of compression schemes …" appears to be "… respectively select one or more candidate compression schemes from the one or more respective sets of compression schemes …";
in Claim 3, lines 1-4; and Claim 18, lines 1-3, "… wherein training the compressed machine-learned model via distillation of the machine-learned model comprises training, by the computing system, the one or more compressed model portions via distillation of the one or more corresponding portions …" appears to be "… wherein training the compressed machine-learned model via the distillation of the machine-learned model comprises training, by the computing system, the one or more compressed model portions via the distillation of the one or more corresponding portions …";
in Claim 4, lines 2-9; and Claim 19, lines 1-8, "… wherein training the one or more compressed model portions via distillation of the one or more corresponding portions … training … the one or more compressed model portions via distillation of the one or more corresponding portions …" appears to be "… wherein training the one or more compressed model portions via the distillation of the one or more corresponding portions … training … the one or more compressed model portions via the distillation of the one or more corresponding portions …";
in Claim 5, lines 1-2; and Claim 20, lines 1-2, "… wherein training the one or more compressed model portions via distillation comprises …" appears to be "… wherein training the one or more compressed model portions via the distillation comprises …";
in Claim 6, lines 5-6; and Claim 21, lines 5-6, "… respectively select the one or more candidate compression schemes from the one or more sets of compression schemes" appears to be "… respectively select the one or more candidate compression schemes from the one or more respective sets of compression schemes";
in Claim 7, lines 1-2; and Claim 22, lines 1-2, "… wherein the cost function evaluates changes in the accuracy metric and the performance metric using …" appears to be "… wherein the cost function evaluates the changes in the accuracy metric and the performance metric using …";
in Claim 15, lines 1-2, " The computer-implemented method of claim 14, in which the computation task is to generate classification data …" appears to be " The computer-implemented method of claim 14, wherein the computation task is to generate classification data …".
Appropriate correction is required.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 5, 12-15, 20, and 23-26 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claims 5 and 20 recite the limitation "… for each of the one or more compressed model portions: processing (…) an input with a compressed model portion to obtain a first output; processing (…) the input with the corresponding model portion to obtain a second output; and training (…) the compressed model portion based on a loss function …" in lines 2-9 and 2-7 respectively, which rendering these claims indefinite because "… for each compressed model portion of the one or more compressed model portions, mapping (…) values from a set of parameters of a corresponding model portion of the machine-learned model to a corresponding set of parameters of the compressed model portion …" is also recited in their respective based claim, and (1) it is unclear whether "a compressed model portion" recited here is the same as or different to "the compressed model portion" (i.e., " each compressed model portion ") recited in their respective based claim and (2) if they are different, which instance of "compressed model portion" is referred by "the compressed model portion" recited here. For examination purpose, "… for said each compressed model portion of the one or more compressed model portions: processing (…) an input with the compressed model portion to obtain a first output; processing (…) the input with the corresponding model portion to obtain a second output; and training (…) the compressed model portion based on a loss function …" is considered.
Claims 12 and 23 recite the limitation "… obtaining … data descriptive of selection of the one or more respective sets of compression schemes for the one or more model portions of a plurality of model portions of a machine-learned model" in lines 4-6, which rendering these claims indefinite because "… obtaining (…) data descriptive of one or more respective sets of compression schemes for one or more model portions of a plurality of model portions of a machine-learned model …" is also recited in their respective based claim and it is unclear (1) whether these two instances of "data descriptive" are the same or different; (2) whether these two instances of " a plurality of model portions" are the same or different; and (3) whether these two instances of "a machine-learned model" are the same or different. Clarification is required.
Claims 13 and 24 recite the limitation "… wherein a portion of a machine-learned model comprises … one or more layers of the machine-learned model" in lines 2-3 and 1-3 respectively, which rendering these claims indefinite because "… obtaining (…) data descriptive of one or more respective sets of compression schemes for one or more model portions of a plurality of model portions of a machine-learned model …" is also recited in their respective based claim, and (1) it is unclear whether "a machine-learned model" recited here is the same as or different to "a machine-learned model" recited in their respective based claim; and (2) if they are different, which instance of "a machine-learned model" is referred by "the machine-learned model" recited here. Clarification is required .
Claims 13 and 24 recite the limitation "... and/or ..." in line 2, which rendering these claims indefinite because it is unclear what is included or excluded by the claim language. For examination purpose, "A and/or B" will be considered as "A or B or both" or "at least one of A and B".
Claim 14 recites the limitation "A computer-implemented method of employing a compressed machine-learned model obtained by a method according to claim 1 to perform a computational task of processing input data …" in lines 1-2, which rendering the claim indefinite because "A computer-implemented method for portion-specific compression and optimization of machine-learned models, comprising … to obtain a compressed machine-learned model … " is also recited in its based claim, and it is unclear (1) whether "a computer-implemented method" recited here, "a method" recited here, and "a computer-implemented method" recited in its based claim are the same or different; and (2) whether "a compressed machine-learned model" recited here and "a compressed machine-learned model" recited in its based claim are the same or different. For examination purpose, "The computer-implemented method of claim 1, further comprising employing the compressed machine-learned model to perform a computational task of processing input data …" is considered.
Claim 15 is rejected for fully incorporating the deficiency of their respective base claims.
Claim 25 recites the limitation "... obtaining data descriptive of a user selection of one or more candidate compression schemes from one or more respective sets of compression schemes for compression of one or more respective model portions of a plurality of model portions of an uncompressed machine- learned model; and applying the one or more compression schemes to one or more respective model portions of the plurality of model portions of the uncompressed machine-learned model to obtain a compressed machine-learned model" in lines 4-10, which rendering the claim indefinite because (1) it is unclear what is referred by "the one or more compression schemes"; and (2) it is unclear whether two instances of "one or more respective model portions" are the same or different. Clarification is required.
Claim 26 is rejected for fully incorporating the deficiency of their respective base claims.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claims 1-26 are rejected under 35 U.S.C. 102(a)(1)/(a)(2) as being anticipated by MOONS et al. (US 2022/0156508 A1, pub. date: 03/19/2022; filed on 11/16/2021), hereinafter MOONS.
Independent Claims 1 and 16
MOONS discloses a computer-implemented method for portion-specific compression and optimization of machine-learned models (MOONS, ¶ [0033]: selecting a neural network architecture suitable for a hardware configuration; using an accuracy predictor to select, from a search space, a neural network comprising a first plurality of the blockwise knowledge distillation trained search blocks; building the accuracy predictor using blockwise knowledge distillation trained search blocks that were trained from the search space; a search for identifying knowledge distillation trained neural network blocks from the search space based on predicted accuracy and any number and combination of cost functions for the search blocks; fine-tuning a neural network made of knowledge distillation trained neural network blocks from the search space selected based on the search and using to generate a distilled neural network; fine-tuning the neural network use weights initialized from the knowledge distillation training of the neural network blocks from the search space; fine-tuning the neural network use knowledge distillation to fine-tuning the neural network; ¶¶ [0038]-[0045]: hand-optimizing efficient on-platform neural networks is prohibitively time and resource consuming, requiring a trained engineer to select a neural network(s) based on experience, from among trillions of architectural options; neural architecture search methods (NAS) are used to help alleviate the costs associated with hand-optimize efficient op-platform neural networks; provide improvements on the high cost of hand-optimizing efficient on-platform neural networks, and improvements on the limitations of existing NAS methods, by automating the design process for a wide and diverse set of potential neural network architectures and hardware platforms and by making the design process less resource intensive; reduce the resource requirements of designing energy efficient on-platform neural networks both in terms of man-hours and compute costs; these reductions in resource requirements may be achieved by (1) a novel way to build accuracy models for a diverse search space of candidate neural network architectures with a variety of cell-types, wherein (a) blockwise knowledge distillation may be implemented to build accuracy models from a diverse neural network architectural search space; and (b) using blockwise knowledge distillation may allow for cheap modeling of the accuracy of neural networks with varying micro-architectures (network-depths, kernel-sizes and expansion-rates), and also across varying macro-architectures (cell-types, attention-mechanisms, activation functions and channel-widths), which is not a feature of existing NAS methods; (2) a quick evolutionary search-phase extracting a front of architectures in terms of accuracy and some on-target efficiency metric (e.g., number of operations, latency, energy consumption, etc.), such as a Pareto-optimal front, wherein (a)the evolutionary search may be performed using the prior accuracy model together with hardware measurements in the loop and may be repeated quickly many times, amortizing the resource-costs of building the accuracy model; and (b) a brief search phase may find latency-accuracy Pareto-optimal neural network architectures for any use-case or hardware platform by running a 2D-optimization algorithm using the accuracy model together with hardware measurements in the loop, which may be quickly rerun whenever anything changes to a use-case, hardware platform, or platform software version; and (3) the way the accuracy model is built in (1) may allow fine-tuning of any neural network in a search space quickly, up to full accuracy; scalable for use in multiple use-case, for multiple hardware configurations, or multiple platform software configurations; the upfront costs of building an accuracy model by performing blockwise knowledge distillation, which requires training partial neural networks (blocks) may be amortized by the ability to reuse the accuracy model multiple times for various circumstances; adapt to a resource budget of any number of processors (e.g., GPUs), as opposed to existing NAS methods that can only scale up their compute to certain practical maximum numbers of GPUs before training can become unstable; the neural network designed using an accuracy model, evolutionary search, and finetuning may also be applicable to downstream tasks; the neural network designed using an accuracy model, evolutionary search, and finetuning may also be applicable to designing, compressing, improving, and/or selecting hardware for other neural networks; ¶¶ [0096]-[0097] with FIG. 9: given the accuracy model and the collection of fine-tuned blocks, a search algorithm may be implemented to identify search blocks and/or neural networks in the search space that achieve a particular balance between predicted accuracy and a cost for execution of the fine-tuned blocks and/or neural networks; the search algorithm may be an evolutionary algorithm executed to find Pareto-optimal search blocks and/or neural networks in the search space that achieve a certain model accuracy, up to maximum model accuracy, and a certain a target cost function, such as a minimum target cost function; a two-dimensional criterion may be used for the search, in which predicted accuracy is balanced with hardware latency; the blocks of the search space, represented here by points, may be plotted based on the predicted accuracy for each block using the accuracy predictor and based on measured and/or predicted hardware latency for implementing the blocks; the dashed line may represent a frontier at which the search may identify as values for which blocks would best suit the criteria; the points plotted closest to the line may represent blocks which the search may identify as best suit the criteria; ¶ [0117] with FIG. 16: applying a distilled neural network architecture from a search-space for refining and compressing a trained neural network; to accomplish this, the trained neural network 1600 may be used as the reference neural network in the search space, and as the reference neural network for training the search blocks with knowledge distillation; refining the neural network 1600 may involve designing a distilled neural network 1602 as a scenario specific version of the neural network; the parameters for the search blocks in the search space and/or the search for search blocks to generate a distilled neural network 1602 may be configured for a specific scenario, such as use-case, hardware configuration, software configuration, etc.; as such, distilling a neural network 1600 based on the reference neural network may result in a scenario specific version of the reference neural network, which may be a distilled neural network 1602 configured to perform better for the scenario; refining the neural network 1600 may involve designing a distilled neural network 1602 as a compressed version of the neural network 1600; the parameters for the search blocks in the search space may be set to be smaller than the corresponding blocks in the reference neural network; as such, distilling a neural network 1600 based on the reference neural network may result in a compressed version of the reference neural network 1600, which may be a smaller and more efficient distilled neural network 1602), comprising:
obtaining, by a computing system comprising one or more computing devices (MOONS, ¶ [0069] with FIG. 6: method 600 may be performed by a processor in a computing system; ¶ [0098] with FIG. 10: method 1000 may be performed by a processor in a computing system; ¶ [0118] with FIG. 17: a variety of commercially available computing systems and computing devices, such as a server 1700), data descriptive of one or more respective sets of compression schemes for one or more model portions of a plurality of model portions of a machine-learned model (MOONS, ¶¶ [0057]-[0066] with FIGS. 3A-B: the complexity of these computations may be reduced by reducing the number of weights that contribute to the output activation, which may be accomplished by setting the values of select weights to zero; the complexity of these computations may also be reduced by using the same set of weights in the calculation of every output of every processing node in a layer; by using convolution, the neural network layer may compute a weighted sum for each output activation using only a small "neighborhood" of inputs (e.g., by setting all other weights beyond the neighborhood to zero, etc.), and share the same set of weights (or filter) for every output; the use of convolution in multiple layers allows the neural network to employ a very deep hierarchy of layers; the normalization functionality component 306, 316 may be configured to control the input distribution across layers to speed up training and the improve accuracy of the outputs or activations; the pooling functionality components 308, 318 may be configured to reduce the dimensionality of a feature map generated by the convolution functionality component 302, 312 and/or otherwise allow the convolutional neural network 300 to resist small shifts and distortions in values; ¶ [0068] with FIG. 5: a search-space may be set to provide every potential neural network architecture as a neural network of N blocks 502 (e.g., block 1 502a, block 2 502b, block 3 502c, block 4 502d, block 5 502e ), which may be referred to herein as search blocks, where N is a positive integer; every block 502 B_n may be any of X_n block implementations, with varying cell type ("cell type") (e.g., style of convolutions), attention mechanisms ("attention"), kernel sizes ("kernel"), number of layers per block, activation functions ("activation"), expansion rates ("expand"), network width ("width scale"), network depth ("depth"), output channels ("ch"), etc.; structural restrictions on the network architectures may be limited, such as being limited to spatial dimensions of the tensors input and output from the blocks matching those of the reference model; ¶¶ [0069]-[0071] in FIG. 6: a method of defining a search-space for distilling neural network architectures; in block 602, define a reference neural network for a search-space; the reference neural network may be used to define constraints of the neural network architectures included in the search-space; receive parameters for defining the reference neural network from a user input or from a memory accessible to the processor; in block 604, define varying parameters for the search space; the varying parameters may be used to define the neural network architectures included in the search-space, which may include varying cell-type (e.g., style of convolutions), attention mechanisms, kernel sizes, number of layers per block, activation functions, expansion rates, network width, network depth, etc.; any number and combination of the varying parameters and the constraints of the reference neural network may define a block or neural network architecture in the search space; receive varying parameters from a user input or from a memory accessible to the processor; ¶ [0087]: the means to build an accuracy model poses no limitations on the variety of the search space, as opposed to existing NAS methods, which either rely on weight sharing or have a limited search space due to GPU-memory limitations; allow for mixing different cell-types, activation functions, quantization levels and/or attention mechanisms, while still being able to model their accuracy reliably; allow for building an accuracy model for a single search space with different attention mechanisms, activation functions, cell-types, channels, quantization settings; allow for adding quantized blocks into the search space, which may build an accuracy model for quantized networks directly; the accuracy predictor 706 may be understood as a coarse sensitivity model that may indicate which blocks require complex implementations in order to build neural networks with high accuracy; ¶¶ [0098]-[0099] with 1002 in FIG. 10: a method of searching for a neural network architecture from a search-space for distilling neural network architectures; in block 1002, set search parameters; search parameters may include a predicted accuracy and a cost function for executing search blocks; the cost function may be scenario-agnostic, such to as a number of operations or a number of parameters in a neural network; the accuracy predictor may be configured to account for scenario specific parameters for the search; the cost function may be scenario-aware, wherein the scenarios may include hardware configurations, software versions, input data parameters, a number of operations or a number of parameters in a neural network, on-device latency, throughput, energy, etc.; set the search parameters based on programmed parameters retrieved from a memory; set the search parameters based on a user input; ¶ [0117] with FIG. 16: the parameters for the search blocks in the search space may be set to be smaller than the corresponding blocks in the reference neural network; as such, distilling a neural network 1600 based on the reference neural network may result in a compressed version of the reference neural network 1600, which may be a smaller and more efficient distilled neural network 1602);
evaluating, by the computing system, a cost function to respectively select one or more candidate compression schemes from the one or more sets of compression schemes (MOONS, ¶ [0004]: selecting a second plurality of the blockwise knowledge distillation trained search blocks based on criteria of predicted accuracy using the accuracy predictor and a cost function for implementing the second plurality of the blockwise knowledge distillation trained search blocks; ¶ [0011]: selecting the neural network of the search space based on a search of the blockwise knowledge distillation trained search blocks using a criterion of predicted accuracy using the accuracy predictor and a cost function for implementing blockwise knowledge distillation trained search blocks of the neural network; ¶ [0033]: a search for identifying knowledge distillation trained neural network blocks from the search space based on predicted accuracy and any number and combination of cost functions for the search blocks; ¶ [0092] with 808 in FIG. 8: in block 808, select a sub-set of neural networks from the search space as targets for building an accuracy model; select the sub-set of neural networks from the search space based on a programmed algorithm, heuristic, technique, criteria, etc.; select the sub-set of neural networks from the search space based on a user input; each of neural networks of the sub-set may include any combination of trained search blocks; the neural networks of the sub-set may include any combination of trained neural networks of the search space; ¶¶ [0100]-[0102] with 1004-1008 in FIG. 10 and ¶¶ [0096]-[0097] with FIG. 9: in block 1004, determine search parameter values for search blocks; calculate cost function values for implementing the search blocks; measure cost function values for implementing the search blocks; retrieve cost function values for implementing the search blocks from the memory accessible to the processor; receive cost function values for implementing the search blocks via user input; the cost function may be scenario-agnostic; the cost function may be scenario-aware, wherein the scenarios may include hardware configurations, software versions, input data parameters, a number of operations or a number of parameters in a neural network, on-device latency, throughput, energy, etc.; in block 1006, determine search blocks suited to the criteria; the search algorithm may be an evolutionary search algorithm; the search algorithm may be configured to identify any number and combination of search blocks, such as N search blocks, suited to the criteria; the configuration of the search algorithm may control whether search parameter values for search blocks may be interpreted as suited to the criteria; the search algorithm may be a Y-dimensional search algorithm, where Y is any integer greater than 1, executed to find search blocks that achieve a certain model accuracy, up to maximum model accuracy, and any number and combination of certain target cost functions, such as a minimum target cost function; e.g., the search criteria may be a two-dimensional search criteria balancing predicted accuracy and cost for implementing search blocks; the search criteria may be to identify Pareto-optimal search blocks for inclusion in a neural network; in block 1008, select the search blocks suited to the criteria; select any number and combination of search blocks suited to the criteria, such as N search blocks suited to the criteria; the processor may select the search blocks best suited to the criteria; select search blocks suited to the criteria within selection parameters; the selection parameters may include a function of the cost function for any number and combination of the search blocks suited to the criteria; the selection parameters may include a function of the predicted accuracy for any number and combination of the search blocks suited to the criteria); and
respectively applying, by the computing system, the one or more candidate compression schemes to the one or more model portions to obtain a compressed machine-learned model comprising one or more compressed model portions that correspond to the one or more model portions (MOONS, ¶¶ [0072]-[0087] with FIGS. 7A-7F: structures and functions for generating an accuracy model using blockwise knowledge distillation for distilling neural network architectures; a trained reference neural network 700 (e.g., non-shaded blocks 1, 2, 3, ... , N) may be used to implement blockwise knowledge distillation to any number and combination of blocks (e.g., shaded blocks 1, 2, 3, ... , N, for which shading, size, and labeling may indicated same of different characteristics between blocks) defined in the search space; the individual blocks of the trained reference neural network 700 may be used to train individual blocks of the neural network architectures 702 (e.g., rows of shaded blocks 702a, 702b, 702c, 702d, 702e, such as likesized and shaded blocks) defined in the search space (e.g., using non-shaded blocks 1, 2, 3, ... , N to train like-labeled shaded blocks 1, 2, 3, ... , N); the individual blocks of the neural network architectures 702 may be trained such that a loss function, such as mean square error (MSE), per-channel Noise-To-Signal ratio (NSR), or any other relevant loss function, between the outputs of the blocks of the trained reference neural network 700 and the individual blocks of the neural network architectures 702 is reduced, such as to within a loss function threshold; e.g., the loss function defined between the blocks of the trained reference neural network 700 and the individual blocks of the neural network architectures 702 may converge to within the loss function threshold while training the individual blocks of the neural network architectures 702; using the trained reference neural network 700, the parameters of all possible blocks B_xn in the search space may be trained; this may allow for (A) extracting quality metrics useful in building supervised regression models for accuracy and (B) initializing their weights for building unsupervised accuracy models, or for quick fine-tuning of models in the search space; training the parameters of all possible blocks B_xn in the search space may be implemented using blockwise knowledge distillation; this process may reduce MSE, per-channel Noise-To-Signal ratio (NSR), or any other relevant loss function of the output features D_xn relative to the trained reference neural network 700; training the parameters of all possible blocks B_xn in the search space may be implemented using conventional knowledge distillation; a new neural network may be constructed where as few as 1 out of the N blocks of the trained reference neural network 700 is replaced by block B_xn; the new neural network is then used to train the weights W_xn of B_xn through knowledge distillation; the target parameters may be developed by fine-tuning a number of sampled neural networks 704 (e.g., rows of shaded blocks 704a, 704b, 704c, 704d, such as variably shaded and/or sized blocks) from the search space, an example of which is illustrated in FIG. 7C (e.g., in which blocks having shading, sizing, and labels may correspond to like shaded, sized, and labeled blocks in FIG. 7A); finetuning may involve training the sampled neural networks 704 using their initializations yielded from the blockwise knowledge distillation using end-to-end knowledge distillation; using the initializations yielded from the blockwise knowledge distillation may reduce the resources needed to fine-tune the sampled neural networks 704 compared to training from scratch to match from scratch training accuracy; these fine-tuned neural networks 704 may also be referred to herein as distilled neural networks; these fine-tuned neural networks 704 may then be used to build an accuracy predictor 706, an example of which is illustrated in FIG. 7F such as in which blocks having shading, sizing, and labels may correspond to like shaded, sized, and labeled blocks in FIG. 7C, and may indicate the loss function, such as MSE, for which the blocks have been trained to reduce; allow for mixing different cell-types, activation functions, quantization levels and/or attention mechanisms, while still being able to model their accuracy reliably; allow for building an accuracy model for a single search space with different attention mechanisms, activation functions, cell-types, channels, quantization settings; allow for adding quantized blocks into the search space, which may build an accuracy model for quantized networks directly; the accuracy predictor 706 may be understood as a coarse sensitivity model that may indicate which blocks require complex implementations in order to build neural networks with high accuracy; ¶¶ [0088]-[0097] with FIG. 8: building multiple neural network architectures for various use-cases; in block 802, train any number and combination of search blocks from the search space using knowledge distillation with a trained reference neural network; in block 804, extract quality features and initialize weights for each search block through knowledge distillation; in block 806, store the quality features and initialize weights for each trained search block and/or neural network architecture; in block 808, select a sub-set of neural networks from the search space as targets for building an accuracy model; in block 810, train fine-tune the sub-set of neural networks as targets for building an accuracy model using knowledge distillation; in block 812, extract targets for building an accuracy model using the fine-tuned neural networks; in block 814, generate an accuracy predictor for neural networks from the search space using the quality features and targets; ¶ [0103] with FIG. 11: any neural network 1102 from the search space (e.g., non-shaded blocks 1, 2, 3, ..., Nin neural networks 702 in FIG. 7A) may be fine-tuned using a trained neural network 1104 having blocks trained using knowledge distillation such as shaded blocks 1, 2, 3, ... , N, for which blocks having shading, sizing, and labels may correspond to like shaded, sized, and labeled blocks in FIG. 7F, neural networks 704 in FIG. 7C, or accuracy predictor 706 in FIG. 7F; the neural network 1102 from the search space and initialized with weights from the knowledge distillation of the search blocks may be fine-tuned using by using end-to-end knowledge distillation using the trained reference neural network 1104; ¶¶ [0104]-[0107] with FIG. 12: fine-tuning a neural network architecture from a search-space for distilling neural network architectures; in block 1202, select a neural network from the search space; the neural network from the search space may include any number and combination of search blocks; select the neural network based on the selection of search blocks suited to the criteria; in block 1204, initialize the neural network from the search space using weights from the training of the search blocks using knowledge distillation; in block 1206, fine-tune the neural network having search blocks initialized with the weight from the knowledge distillation by using knowledge distillation with a trained reference neural network; ¶ [0117] with FIG. 16: a trained neural network 1600 may be modified to use trained search blocks (e.g., shaded blocks 1, 2, 3, 5, 7, 8, 9 in neural network 1602) from the search space to refine and compress the trained neural network 1600; refining the trained neural network may retain whole the structure of the trained neural network, replacing the blocks (e.g., non-shaded blocks 1, 2, 3, 5, 7, 8, 9 in neural network 1600) within the trained neural network with trained blocks (e.g., shaded blocks 1, 2, 3, 5, 7, 8, 9 in neural network 1602) from the search space, which may be more efficient on hardware).
MOONS further discloses a computing system (MOONS , ¶ [0069] with FIG. 6: method 600 may be performed by a processor in a computing system; ¶[0088] with FIG. 8: method 800 may be performed by a processor in a computing system; ¶ [0098] with FIG. 10: method 1000 may be performed by a processor in a computing system; ¶ [0104] with FIG. 12: method 1200 may be performed by a processor in a computing system; ¶ [0118] with 1700 in FIG. 17: a variety of commercially available computing systems and computing devices, such as a server 1700) comprising: one or more processors (MOONS, ¶¶ [0118]-[0119] with 1701 in FIG. 17: a server 1700 typically includes a processor 1701; the processor 1701 may be any programmable microprocessor, microcomputer or multiple processor chip or chip); and one or more non-transitory computer-readable media (MOONS, ¶¶ [0118]-[0119] with 1702 and 1703 in FIG. 17: a server 1700 typically includes a processor 1701 coupled to volatile memory 1702 and a large capacity nonvolatile memory, such as a disk drive 1703; non-transitory processor-readable medium, such as a disk drive 1703) that store instructions that, when executed by the one or more processors, cause the computing system to perform operations described above (MOONS, ¶ [0119] with FIG.17: the processor 1701 may be any programmable microprocessor, microcomputer or multiple processor chip or chips that may be configured by software instructions (applications) to perform a variety of functions; software applications may be stored on non-transitory processor-readable medium, such as a disk drive 1703, before the instructions are accessed and loaded into the processor).
Claims 2 and 17
MOONS discloses all the elements as stated in Claims 1 and 16 respectively and further discloses training, by the computing system, the compressed machine-learned model via distillation of the machine-learned model (MOONS, ¶ [0033]: selecting a neural network architecture suitable for a hardware configuration; using an accuracy predictor to select, from a search space, a neural network comprising a first plurality of the blockwise knowledge distillation trained search blocks; building the accuracy predictor using blockwise knowledge distillation trained search blocks that were trained from the search space; a search for identifying knowledge distillation trained neural network blocks from the search space based on predicted accuracy and any number and combination of cost functions for the search blocks; fine-tuning a neural network made of knowledge distillation trained neural network blocks from the search space selected based on the search and using to generate a distilled neural network; fine-tuning the neural network use weights initialized from the knowledge distillation training of the neural network blocks from the search space; fine-tuning the neural network use knowledge distillation to fine-tuning the neural network; ¶¶ [0072]-[0087] with FIGS. 7A-7F: structures and functions for generating an accuracy model using blockwise knowledge distillation for distilling neural network architectures; a trained reference neural network 700 (e.g., non-shaded blocks 1, 2, 3, ... , N) may be used to implement blockwise knowledge distillation to any number and combination of blocks (e.g., shaded blocks 1, 2, 3, ... , N, for which shading, size, and labeling may indicated same of different characteristics between blocks) defined in the search space; the individual blocks of the trained reference neural network 700 may be used to train individual blocks of the neural network architectures 702 (e.g., rows of shaded blocks 702a, 702b, 702c, 702d, 702e, such as likesized and shaded blocks) defined in the search space (e.g., using non-shaded blocks 1, 2, 3, ... , N to train like-labeled shaded blocks 1, 2, 3, ... , N); the individual blocks of the neural network architectures 702 may be trained such that a loss function, such as mean square error (MSE), per-channel Noise-To-Signal ratio (NSR), or any other relevant loss function, between the outputs of the blocks of the trained reference neural network 700 and the individual blocks of the neural network architectures 702 is reduced, such as to within a loss function threshold; e.g., the loss function defined between the blocks of the trained reference neural network 700 and the individual blocks of the neural network architectures 702 may converge to within the loss function threshold while training the individual blocks of the neural network architectures 702; using the trained reference neural network 700, the parameters of all possible blocks B_xn in the search space may be trained; this may allow for (A) extracting quality metrics useful in building supervised regression models for accuracy and (B) initializing their weights for building unsupervised accuracy models, or for quick fine-tuning of models in the search space; training the parameters of all possible blocks B_xn in the search space may be implemented using blockwise knowledge distillation; this process may reduce MSE, per-channel Noise-To-Signal ratio (NSR), or any other relevant loss function of the output features D xn relative to the trained reference neural network 700; training the parameters of all possible blocks B_xn in the search space may be implemented using conventional knowledge distillation; a new neural network may be constructed where as few as 1 out of the N blocks of the trained reference neural network 700 is replaced by block B_xn; the new neural network is then used to train the weights W_xn of B_xn through knowledge distillation; the target parameters may be developed by fine-tuning a number of sampled neural networks 704 (e.g., rows of shaded blocks 704a, 704b, 704c, 704d, such as variably shaded and/or sized blocks) from the search space, an example of which is illustrated in FIG. 7C (e.g., in which blocks having shading, sizing, and labels may correspond to like shaded, sized, and labeled blocks in FIG. 7A); finetuning may involve training the sampled neural networks 704 using their initializations yielded from the blockwise knowledge distillation using end-to-end knowledge distillation; using the initializations yielded from the blockwise knowledge distillation may reduce the resources needed to fine-tune the sampled neural networks 704 compared to training from scratch to match from scratch training accuracy; these fine-tuned neural networks 704 may also be referred to herein as distilled neural networks; these fine-tuned neural networks 704 may then be used to build an accuracy predictor 706, an example of which is illustrated in FIG. 7F such as in which blocks having shading, sizing, and labels may correspond to like shaded, sized, and labeled blocks in FIG. 7C, and may indicate the loss function, such as MSE, for which the blocks have been trained to reduce; allow for mixing different cell-types, activation functions, quantization levels and/or attention mechanisms, while still being able to model their accuracy reliably; allow for building an accuracy model for a single search space with different attention mechanisms, activation functions, cell-types, channels, quantization settings; allow for adding quantized blocks into the search space, which may build an accuracy model for quantized networks directly; the accuracy predictor 706 may be understood as a coarse sensitivity model that may indicate which blocks require complex implementations in order to build neural networks with high accuracy; ¶¶ [0088]-[0097] with FIG. 8: building multiple neural network architectures for various use-cases; in block 802, train any number and combination of search blocks from the search space using knowledge distillation with a trained reference neural network; in block 804, extract quality features and initialize weights for each search block through knowledge distillation; in block 806, store the quality features and initialize weights for each trained search block and/or neural network architecture; in block 808, select a sub-set of neural networks from the search space as targets for building an accuracy model; in block 810, train fine-tune the sub-set of neural networks as targets for building an accuracy model using knowledge distillation; in block 812, extract targets for building an accuracy model using the fine-tuned neural networks; in block 814, generate an accuracy predictor for neural networks from the search space using the quality features and targets; ¶ [0103] with FIG. 11: any neural network 1102 from the search space (e.g., non-shaded blocks 1, 2, 3, ..., Nin neural networks 702 in FIG. 7A) may be fine-tuned using a trained neural network 1104 having blocks trained using knowledge distillation such as shaded blocks 1, 2, 3, ... , N, for which blocks having shading, sizing, and labels may correspond to like shaded, sized, and labeled blocks in FIG. 7F, neural networks 704 in FIG. 7C, or accuracy predictor 706 in FIG. 7F; the neural network 1102 from the search space and initialized with weights from the knowledge distillation of the search blocks may be fine-tuned using by using end-to-end knowledge distillation using the trained reference neural network 1104; ¶¶ [0104]-[0107] with FIG. 12: fine-tuning a neural network architecture from a search-space for distilling neural network architectures; in block 1202, select a neural network from the search space; the neural network from the search space may include any number and combination of search blocks; select the neural network based on the selection of search blocks suited to the criteria; in block 1204, initialize the neural network from the search space using weights from the training of the search blocks using knowledge distillation; in block 1206, fine-tune the neural network having search blocks initialized with the weight from the knowledge distillation by using knowledge distillation with a trained reference neural network; ¶ [0117] with FIG. 16: applying a distilled neural network architecture from a search-space for refining and compressing a trained neural network; a trained neural network 1600 may be modified to use trained search blocks (e.g., shaded blocks 1, 2, 3, 5, 7, 8, 9 in neural network 1602) from the search space to refine and compress the trained neural network 1600; refining the trained neural network may retain whole the structure of the trained neural network, replacing the blocks (e.g., non-shaded blocks 1, 2, 3, 5, 7, 8, 9 in neural network 1600) within the trained neural network with trained blocks (e.g., shaded blocks 1, 2, 3, 5, 7, 8, 9 in neural network 1602) from the search space, which may be more efficient on hardware; to accomplish this, the trained neural network 1600 may be used as the reference neural network in the search space, and as the reference neural network for training the search blocks with knowledge distillation; refining the neural network 1600 may involve designing a distilled neural network 1602 as a scenario specific version of the neural network; the parameters for the search blocks in the search space and/or the search for search blocks to generate a distilled neural network 1602 may be configured for a specific scenario, such as use-case, hardware configuration, software configuration, etc.; as such, distilling a neural network 1600 based on the reference neural network may result in a scenario specific version of the reference neural network, which may be a distilled neural network 1602 configured to perform better for the scenario; refining the neural network 1600 may involve designing a distilled neural network 1602 as a compressed version of the neural network 1600; the parameters for the search blocks in the search space may be set to be smaller than the corresponding blocks in the reference neural network; as such, distilling a neural network 1600 based on the reference neural network may result in a compressed version of the reference neural network 1600, which may be a smaller and more efficient distilled neural network 1602).
Claims 3 and 18
MOONS discloses all the elements as stated in Claims 2 and 17 respectively and further discloses wherein training the compressed machine-learned model via distillation of the machine-learned model comprises training, by the computing system, the one or more compressed model portions via distillation of the one or more corresponding portions of the machine-learned model (MOONS, ¶ [0033]: selecting a neural network architecture suitable for a hardware configuration; using an accuracy predictor to select, from a search space, a neural network comprising a first plurality of the blockwise knowledge distillation trained search blocks; building the accuracy predictor using blockwise knowledge distillation trained search blocks that were trained from the search space; a search for identifying knowledge distillation trained neural network blocks from the search space based on predicted accuracy and any number and combination of cost functions for the search blocks; fine-tuning a neural network made of knowledge distillation trained neural network blocks from the search space selected based on the search and using to generate a distilled neural network; fine-tuning the neural network use weights initialized from the knowledge distillation training of the neural network blocks from the search space; fine-tuning the neural network use knowledge distillation to fine-tuning the neural network; ¶¶ [0072]-[0087] with FIGS. 7A-7F: structures and functions for generating an accuracy model using blockwise knowledge distillation for distilling neural network architectures; a trained reference neural network 700 (e.g., non-shaded blocks 1, 2, 3, ... , N) may be used to implement blockwise knowledge distillation to any number and combination of blocks (e.g., shaded blocks 1, 2, 3, ... , N, for which shading, size, and labeling may indicated same of different characteristics between blocks) defined in the search space; the individual blocks of the trained reference neural network 700 may be used to train individual blocks of the neural network architectures 702 (e.g., rows of shaded blocks 702a, 702b, 702c, 702d, 702e, such as likesized and shaded blocks) defined in the search space (e.g., using non-shaded blocks 1, 2, 3, ... , N to train like-labeled shaded blocks 1, 2, 3, ... , N); the individual blocks of the neural network architectures 702 may be trained such that a loss function, such as mean square error (MSE), per-channel Noise-To-Signal ratio (NSR), or any other relevant loss function, between the outputs of the blocks of the trained reference neural network 700 and the individual blocks of the neural network architectures 702 is reduced, such as to within a loss function threshold; e.g., the loss function defined between the blocks of the trained reference neural network 700 and the individual blocks of the neural network architectures 702 may converge to within the loss function threshold while training the individual blocks of the neural network architectures 702; using the trained reference neural network 700, the parameters of all possible blocks B_xn in the search space may be trained; this may allow for (A) extracting quality metrics useful in building supervised regression models for accuracy and (B) initializing their weights for building unsupervised accuracy models, or for quick fine-tuning of models in the search space; training the parameters of all possible blocks B_xn in the search space may be implemented using blockwise knowledge distillation; this process may reduce MSE, per-channel Noise-To-Signal ratio (NSR), or any other relevant loss function of the output features D xn relative to the trained reference neural network 700; training the parameters of all possible blocks B_xn in the search space may be implemented using conventional knowledge distillation; a new neural network may be constructed where as few as 1 out of the N blocks of the trained reference neural network 700 is replaced by block B_xn; the new neural network is then used to train the weights W_xn of B_xn through knowledge distillation; the target parameters may be developed by fine-tuning a number of sampled neural networks 704 (e.g., rows of shaded blocks 704a, 704b, 704c, 704d, such as variably shaded and/or sized blocks) from the search space, an example of which is illustrated in FIG. 7C (e.g., in which blocks having shading, sizing, and labels may correspond to like shaded, sized, and labeled blocks in FIG. 7A); finetuning may involve training the sampled neural networks 704 using their initializations yielded from the blockwise knowledge distillation using end-to-end knowledge distillation; using the initializations yielded from the blockwise knowledge distillation may reduce the resources needed to fine-tune the sampled neural networks 704 compared to training from scratch to match from scratch training accuracy; these fine-tuned neural networks 704 may also be referred to herein as distilled neural networks; these fine-tuned neural networks 704 may then be used to build an accuracy predictor 706, an example of which is illustrated in FIG. 7F such as in which blocks having shading, sizing, and labels may correspond to like shaded, sized, and labeled blocks in FIG. 7C, and may indicate the loss function, such as MSE, for which the blocks have been trained to reduce; allow for mixing different cell-types, activation functions, quantization levels and/or attention mechanisms, while still being able to model their accuracy reliably; allow for building an accuracy model for a single search space with different attention mechanisms, activation functions, cell-types, channels, quantization settings; allow for adding quantized blocks into the search space, which may build an accuracy model for quantized networks directly; the accuracy predictor 706 may be understood as a coarse sensitivity model that may indicate which blocks require complex implementations in order to build neural networks with high accuracy; ¶¶ [0088]-[0097] with FIG. 8: building multiple neural network architectures for various use-cases; in block 802, train any number and combination of search blocks from the search space using knowledge distillation with a trained reference neural network; in block 804, extract quality features and initialize weights for each search block through knowledge distillation; in block 806, store the quality features and initialize weights for each trained search block and/or neural network architecture; in block 808, select a sub-set of neural networks from the search space as targets for building an accuracy model; in block 810, train fine-tune the sub-set of neural networks as targets for building an accuracy model using knowledge distillation; in block 812, extract targets for building an accuracy model using the fine-tuned neural networks; in block 814, generate an accuracy predictor for neural networks from the search space using the quality features and targets; ¶ [0103] with FIG. 11: any neural network 1102 from the search space (e.g., non-shaded blocks 1, 2, 3, ..., Nin neural networks 702 in FIG. 7A) may be fine-tuned using a trained neural network 1104 having blocks trained using knowledge distillation such as shaded blocks 1, 2, 3, ... , N, for which blocks having shading, sizing, and labels may correspond to like shaded, sized, and labeled blocks in FIG. 7F, neural networks 704 in FIG. 7C, or accuracy predictor 706 in FIG. 7F; the neural network 1102 from the search space and initialized with weights from the knowledge distillation of the search blocks may be fine-tuned using by using end-to-end knowledge distillation using the trained reference neural network 1104; ¶¶ [0104]-[0107] with FIG. 12: fine-tuning a neural network architecture from a search-space for distilling neural network architectures; in block 1202, select a neural network from the search space; the neural network from the search space may include any number and combination of search blocks; select the neural network based on the selection of search blocks suited to the criteria; in block 1204, initialize the neural network from the search space using weights from the training of the search blocks using knowledge distillation; in block 1206, fine-tune the neural network having search blocks initialized with the weight from the knowledge distillation by using knowledge distillation with a trained reference neural network; ¶ [0117] with FIG. 16: applying a distilled neural network architecture from a search-space for refining and compressing a trained neural network; a trained neural network 1600 may be modified to use trained search blocks (e.g., shaded blocks 1, 2, 3, 5, 7, 8, 9 in neural network 1602) from the search space to refine and compress the trained neural network 1600; refining the trained neural network may retain whole the structure of the trained neural network, replacing the blocks (e.g., non-shaded blocks 1, 2, 3, 5, 7, 8, 9 in neural network 1600) within the trained neural network with trained blocks (e.g., shaded blocks 1, 2, 3, 5, 7, 8, 9 in neural network 1602) from the search space, which may be more efficient on hardware; to accomplish this, the trained neural network 1600 may be used as the reference neural network in the search space, and as the reference neural network for training the search blocks with knowledge distillation; refining the neural network 1600 may involve designing a distilled neural network 1602 as a scenario specific version of the neural network; the parameters for the search blocks in the search space and/or the search for search blocks to generate a distilled neural network 1602 may be configured for a specific scenario, such as use-case, hardware configuration, software configuration, etc.; as such, distilling a neural network 1600 based on the reference neural network may result in a scenario specific version of the reference neural network, which may be a distilled neural network 1602 configured to perform better for the scenario; refining the neural network 1600 may involve designing a distilled neural network 1602 as a compressed version of the neural network 1600; the parameters for the search blocks in the search space may be set to be smaller than the corresponding blocks in the reference neural network; as such, distilling a neural network 1600 based on the reference neural network may result in a compressed version of the reference neural network 1600, which may be a smaller and more efficient distilled neural network 1602).
Claims 4 and 19
MOONS discloses all the elements as stated in Claims 2 and 17 respectively and further discloses wherein training the one or more compressed model portions via distillation of the one or more corresponding portions of the machine-learned model comprises: for each compressed model portion of the one or more compressed model portions, mapping, by the computing system, values from a set of parameters of a corresponding model portion of the machine-learned model to a corresponding set of parameters of the compressed model portion; and training, by the computing system, the one or more compressed model portions via distillation of the one or more corresponding portions of the machine-learned model (MOONS, ¶¶ [0057]-[0066] with FIGS. 3A-B: the complexity of these computations may be reduced by reducing the number of weights that contribute to the output activation, which may be accomplished by setting the values of select weights to zero; the complexity of these computations may also be reduced by using the same set of weights in the calculation of every output of every processing node in a layer; by using convolution, the neural network layer may compute a weighted sum for each output activation using only a small "neighborhood" of inputs (e.g., by setting all other weights beyond the neighborhood to zero, etc.), and share the same set of weights (or filter) for every output; a set of weights is called a filter or kernel; a filter (or kernel) may also be a two- or three-dimensional matrix of weight parameters; implement a filter via a multidimensional array, map, table or any other information structure known in the art; the use of convolution in multiple layers allows the neural network to employ a very deep hierarchy of layers; the convolution functionality component 302, 312 may be configured to generate a matrix of output activations called a feature map; the normalization functionality component 306, 316 may be configured to control the input distribution across layers to speed up training and the improve accuracy of the outputs or activations; the pooling functionality components 308, 318 may be configured to reduce the dimensionality of a feature map generated by the convolution functionality component 302, 312 and/or otherwise allow the convolutional neural network 300 to resist small shifts and distortions in values; the inputs to the first layer 301 may be structured as a set of three-dimensional input feature maps 352 that form a channel of input feature maps; the neural network has a batch size of N three-dimensional feature maps 352 with height H and width W each having C number of channels of input feature maps (illustrated as two-dimensional maps in C channels), and M three-dimensional filters 354 including C filters for each channel (also illustrated as two-dimensional filters for C channels); applying the 1 to M filters 354 to the 1 to N three-dimensional feature maps 352 results in N output feature maps 356 that include M channels of width F and height E; each channel may be convolved with a three-dimensional filter 354; the results of these convolutions may be summed across all the channels to generate the output activations of the first layer 301 in the form of a channel of output feature maps 356; additional three-dimensional filters may be applied to the input feature maps 352 to create additional output channels, and multiple input feature maps 352 may be processed together as a batch to improve the reuse of the filter weights; the results of the output channel ( e.g., set of output feature maps 356) may be fed to the second layer 311 in the convolutional neural network 300 for further processing; ¶ [0068] with FIG. 5: a search-space may be set to provide every potential neural network architecture as a neural network of N blocks 502 (e.g., block 1 502a, block 2 502b, block 3 502c, block 4 502d, block 5 502e ), which may be referred to herein as search blocks, where N is a positive integer; every block 502 B_n may be any of X_n block implementations, with varying cell type ("cell type") (e.g., style of convolutions), attention mechanisms ("attention"), kernel sizes ("kernel"), number of layers per block, activation functions ("activation"), expansion rates ("expand"), network width ("width scale"), network depth ("depth"), output channels ("ch"), etc.; structural restrictions on the network architectures may be limited, such as being limited to spatial dimensions of the tensors input and output from the blocks matching those of the reference model; ¶ [0077]: the weights W xn of B_xn may be updated using conventional knowledge distillation, by minimizing a loss function defined by the tasks ground truth and the output classifier of the trained reference neural network 700; ¶¶ [0072]-[0087] with FIGS. 7A-7F: structures and functions for generating an accuracy model using blockwise knowledge distillation for distilling neural network architectures; a trained reference neural network 700 (e.g., non-shaded blocks 1, 2, 3, ... , N) may be used to implement blockwise knowledge distillation to any number and combination of blocks (e.g., shaded blocks 1, 2, 3, ... , N, for which shading, size, and labeling may indicated same of different characteristics between blocks) defined in the search space; the individual blocks of the trained reference neural network 700 may be used to train individual blocks of the neural network architectures 702 (e.g., rows of shaded blocks 702a, 702b, 702c, 702d, 702e, such as likesized and shaded blocks) defined in the search space (e.g., using non-shaded blocks 1, 2, 3, ... , N to train like-labeled shaded blocks 1, 2, 3, ... , N); the individual blocks of the neural network architectures 702 may be trained such that a loss function, such as mean square error (MSE), per-channel Noise-To-Signal ratio (NSR), or any other relevant loss function, between the outputs of the blocks of the trained reference neural network 700 and the individual blocks of the neural network architectures 702 is reduced, such as to within a loss function threshold; e.g., the loss function defined between the blocks of the trained reference neural network 700 and the individual blocks of the neural network architectures 702 may converge to within the loss function threshold while training the individual blocks of the neural network architectures 702; to implement the blockwise knowledge distillation a reference neural network 700 may be selected, trained, and split it into N blocks BT_n; this block BT_n may transform input feature map DT_(n-1) into DT_n using a transfer function FT_n(WT_n), where WT_n are the parameters of the block; using the trained reference neural network 700, the parameters of all possible blocks B_xn in the search space may be trained; this may allow for (A) extracting quality metrics useful in building supervised regression models for accuracy and (B) initializing their weights for building unsupervised accuracy models, or for quick fine-tuning of models in the search space; training the parameters of all possible blocks B_xn in the search space may be implemented using blockwise knowledge distillation; the parameters W_xn of all the blocks B_xn, which define the transfer function F_xn(W_xn), may be trained using a blockwise knowledge distillation scheme to approximate the reference function F_n as closely as possible; this is done in a block-wise way by using stochastic gradient descent, requiring gradient back-propagation only from DT_n to DT_(n-1) through B_xn; this process may reduce MSE, per-channel Noise-To-Signal ratio (NSR), or any other relevant loss function of the output features D xn relative to the trained reference neural network 700; DT_n, using the trained reference neural networks' DT_(n-1) input feature maps; this blockwise knowledge distillation may converge the loss function quickly, such as after 1 full epoch of training; training the parameters of all possible blocks B_xn in the search space may be implemented using conventional knowledge distillation; the weights W_xn of B_xn may be updated using conventional knowledge distillation, by minimizing a loss function defined by the tasks ground truth and the output classifier of the trained reference neural network 700; a new neural network may be constructed where as few as 1 out of the N blocks of the trained reference neural network 700 is replaced by block B_xn; the new neural network is then used to train the weights W_xn of B_xn through knowledge distillation; the knowledge distillation, blockwise or conventional, may be used as a means to measure the quality of a block B_n,x in the search space and to initialize its weights for later fine-tuning; the MSE may be used as a metric indicating the quality of the block B_n,x; later, the block weights Wn,x from blockwise knowledge distillation may serve as the weight-initialization for quickly fine-tuning any neural network architecture sampled from the search-space, such as to full accuracy; the built accuracy predictor may be trained using parameters including features ( e.g., quality metrics) and targets (e.g., accuracy); the target parameters may be developed by fine-tuning a number of sampled neural networks 704 (e.g., rows of shaded blocks 704a, 704b, 704c, 704d, such as variably shaded and/or sized blocks) from the search space, an example of which is illustrated in FIG. 7C (e.g., in which blocks having shading, sizing, and labels may correspond to like shaded, sized, and labeled blocks in FIG. 7A); finetuning may involve training the sampled neural networks 704 using their initializations yielded from the blockwise knowledge distillation using end-to-end knowledge distillation; using the initializations yielded from the blockwise knowledge distillation may reduce the resources needed to fine-tune the sampled neural networks 704 compared to training from scratch to match from scratch training accuracy; these fine-tuned neural networks 704 may also be referred to herein as distilled neural networks; these fine-tuned neural networks 704 may then be used to build an accuracy predictor 706, an example of which is illustrated in FIG. 7F such as in which blocks having shading, sizing, and labels may correspond to like shaded, sized, and labeled blocks in FIG. 7C, and may indicate the loss function, such as MSE, for which the blocks have been trained to reduce; a number of blockwise features (e.g., features illustrated in FIG. 7B) may be used to build a ranking of architectures in the search space ( e.g., targets illustrated in FIG. 7D), an example of which is illustrated in FIG. 7E; the NSR of the feature maps may be between a trained reference neural network 700 and fine-tuned neural network 704; this may be understood as a measure of the distance between F_n and F_xn; the closer F_xn is to F_n, the higher the quality of block B_xn and the lower the average NSR or MSE; training/validation loss of a fine-tuned neural network 704 M_xn may be used as an accuracy target to fit the accuracy predictor 706; allow for mixing different cell-types, activation functions, quantization levels and/or attention mechanisms, while still being able to model their accuracy reliably; allow for building an accuracy model for a single search space with different attention mechanisms, activation functions, cell-types, channels, quantization settings; allow for adding quantized blocks into the search space, which may build an accuracy model for quantized networks directly; the accuracy predictor 706 may be understood as a coarse sensitivity model that may indicate which blocks require complex implementations in order to build neural networks with high accuracy; ¶¶ [0088]-[0097] with FIG. 8: building multiple neural network architectures for various use-cases; in block 802, train any number and combination of search blocks from the search space using knowledge distillation with a trained reference neural network; the search blocks may be trained using blockwise knowledge distillation in which the search blocks are associated with blocks of the trained reference neural network and trained using the associated blocks to converge a loss function of each search block with a loss function of an associated block; the search blocks may be trained using conventional knowledge distillation in which a neural network architecture having search blocks may be trained using the trained reference neural network to converge a loss function of the neural network architecture with a loss function of the trained reference neural network; e.g., the loss function of the blocks and/or the neural networks may converge while training the search blocks to an accuracy level, up to full accuracy; e.g., convergence of the loss function of the blocks and/or the neural networks may be to within a loss function threshold; in block 804, extract quality features and initialize weights for each search block through knowledge distillation; in block 806, store the quality features and initialize weights for each trained search block and/or neural network architecture; in block 808, select a sub-set of neural networks from the search space as targets for building an accuracy model; in block 810, train fine-tune the sub-set of neural networks as targets for building an accuracy model using knowledge distillation; in block 812, extract targets for building an accuracy model using the fine-tuned neural networks; in block 814, generate an accuracy predictor for neural networks from the search space using the quality features and targets; ¶ [0103] with FIG. 11: any neural network 1102 from the search space (e.g., non-shaded blocks 1, 2, 3, ..., Nin neural networks 702 in FIG. 7A) may be fine-tuned using a trained neural network 1104 having blocks trained using knowledge distillation such as shaded blocks 1, 2, 3, ... , N, for which blocks having shading, sizing, and labels may correspond to like shaded, sized, and labeled blocks in FIG. 7F, neural networks 704 in FIG. 7C, or accuracy predictor 706 in FIG. 7F; the neural network 1102 from the search space and initialized with weights from the knowledge distillation of the search blocks may be fine-tuned using by using end-to-end knowledge distillation using the trained reference neural network 1104; ¶¶ [0104]-[0107] with FIG. 12: fine-tuning a neural network architecture from a search-space for distilling neural network architectures; in block 1202, select a neural network from the search space; the neural network from the search space may include any number and combination of search blocks; select the neural network based on the selection of search blocks suited to the criteria; in block 1204, initialize the neural network from the search space using weights from the training of the search blocks using knowledge distillation; in block 1206, fine-tune the neural network having search blocks initialized with the weight from the knowledge distillation by using knowledge distillation with a trained reference neural network; ¶ [0117] with FIG. 16: applying a distilled neural network architecture from a search-space for refining and compressing a trained neural network; a trained neural network 1600 may be modified to use trained search blocks (e.g., shaded blocks 1, 2, 3, 5, 7, 8, 9 in neural network 1602) from the search space to refine and compress the trained neural network 1600; refining the trained neural network may retain whole the structure of the trained neural network, replacing the blocks (e.g., non-shaded blocks 1, 2, 3, 5, 7, 8, 9 in neural network 1600) within the trained neural network with trained blocks (e.g., shaded blocks 1, 2, 3, 5, 7, 8, 9 in neural network 1602) from the search space, which may be more efficient on hardware; to accomplish this, the trained neural network 1600 may be used as the reference neural network in the search space, and as the reference neural network for training the search blocks with knowledge distillation; refining the neural network 1600 may involve designing a distilled neural network 1602 as a scenario specific version of the neural network; the parameters for the search blocks in the search space and/or the search for search blocks to generate a distilled neural network 1602 may be configured for a specific scenario, such as use-case, hardware configuration, software configuration, etc.; as such, distilling a neural network 1600 based on the reference neural network may result in a scenario specific version of the reference neural network, which may be a distilled neural network 1602 configured to perform better for the scenario; refining the neural network 1600 may involve designing a distilled neural network 1602 as a compressed version of the neural network 1600; the parameters for the search blocks in the search space may be set to be smaller than the corresponding blocks in the reference neural network; as such, distilling a neural network 1600 based on the reference neural network may result in a compressed version of the reference neural network 1600, which may be a smaller and more efficient distilled neural network 1602).
Claims 5 and 20
MOONS discloses all the elements as stated in Claims 4 and 18 respectively and further discloses wherein training the one or more compressed model portions via distillation comprises, for each of the one or more compressed model portions: processing, by the computing system, an input with a compressed model portion to obtain a first output; processing, by the computing system, the input with the corresponding model portion to obtain a second output; and training, by the computing system, the compressed model portion based on a loss function that evaluates a difference between the second output and the first output (MOONS, ¶ [0051]: the difference between the expected/desired output and the output generated by the neural network 100 is referred to as loss (L); ¶¶ [0072]-[0087] with FIGS. 7A-7F: structures and functions for generating an accuracy model using blockwise knowledge distillation for distilling neural network architectures; a trained reference neural network 700 (e.g., non-shaded blocks 1, 2, 3, ... , N) may be used to implement blockwise knowledge distillation to any number and combination of blocks (e.g., shaded blocks 1, 2, 3, ... , N, for which shading, size, and labeling may indicated same of different characteristics between blocks) defined in the search space; the individual blocks of the trained reference neural network 700 may be used to train individual blocks of the neural network architectures 702 (e.g., rows of shaded blocks 702a, 702b, 702c, 702d, 702e, such as likesized and shaded blocks) defined in the search space (e.g., using non-shaded blocks 1, 2, 3, ... , N to train like-labeled shaded blocks 1, 2, 3, ... , N); the individual blocks of the neural network architectures 702 may be trained such that a loss function, such as mean square error (MSE), per-channel Noise-To-Signal ratio (NSR), or any other relevant loss function, between the outputs of the blocks of the trained reference neural network 700 and the individual blocks of the neural network architectures 702 is reduced, such as to within a loss function threshold; e.g., the loss function defined between the blocks of the trained reference neural network 700 and the individual blocks of the neural network architectures 702 may converge to within the loss function threshold while training the individual blocks of the neural network architectures 702; to implement the blockwise knowledge distillation a reference neural network 700 may be selected, trained, and split it into N blocks BT_n; this block BT_n may transform input feature map DT_(n-1) into DT_n using a transfer function FT_n(WT_n), where WT_n are the parameters of the block; using the trained reference neural network 700, the parameters of all possible blocks B_xn in the search space may be trained; this may allow for (A) extracting quality metrics useful in building supervised regression models for accuracy and (B) initializing their weights for building unsupervised accuracy models, or for quick fine-tuning of models in the search space; training the parameters of all possible blocks B_xn in the search space may be implemented using blockwise knowledge distillation; the parameters W_xn of all the blocks B_xn, which define the transfer function F_xn(W_xn), may be trained using a blockwise knowledge distillation scheme to approximate the reference function F_n as closely as possible; this is done in a block-wise way by using stochastic gradient descent, requiring gradient back-propagation only from DT_n to DT_(n-1) through B_xn; this process may reduce MSE, per-channel Noise-To-Signal ratio (NSR), or any other relevant loss function of the output features D xn relative to the trained reference neural network 700; DT_n, using the trained reference neural networks' DT_(n-1) input feature maps; this blockwise knowledge distillation may converge the loss function quickly, such as after 1 full epoch of training; training the parameters of all possible blocks B_xn in the search space may be implemented using conventional knowledge distillation; the weights W_xn of B_xn may be updated using conventional knowledge distillation, by minimizing a loss function defined by the tasks ground truth and the output classifier of the trained reference neural network 700; a new neural network may be constructed where as few as 1 out of the N blocks of the trained reference neural network 700 is replaced by block B_xn; the new neural network is then used to train the weights W_xn of B_xn through knowledge distillation; the knowledge distillation, blockwise or conventional, may be used as a means to measure the quality of a block B_n,x in the search space and to initialize its weights for later fine-tuning; the MSE may be used as a metric indicating the quality of the block B_n,x; later, the block weights Wn,x from blockwise knowledge distillation may serve as the weight-initialization for quickly fine-tuning any neural network architecture sampled from the search-space, such as to full accuracy; the built accuracy predictor may be trained using parameters including features ( e.g., quality metrics) and targets (e.g., accuracy); the target parameters may be developed by fine-tuning a number of sampled neural networks 704 (e.g., rows of shaded blocks 704a, 704b, 704c, 704d, such as variably shaded and/or sized blocks) from the search space, an example of which is illustrated in FIG. 7C (e.g., in which blocks having shading, sizing, and labels may correspond to like shaded, sized, and labeled blocks in FIG. 7A); finetuning may involve training the sampled neural networks 704 using their initializations yielded from the blockwise knowledge distillation using end-to-end knowledge distillation; using the initializations yielded from the blockwise knowledge distillation may reduce the resources needed to fine-tune the sampled neural networks 704 compared to training from scratch to match from scratch training accuracy; these fine-tuned neural networks 704 may also be referred to herein as distilled neural networks; these fine-tuned neural networks 704 may then be used to build an accuracy predictor 706, an example of which is illustrated in FIG. 7F such as in which blocks having shading, sizing, and labels may correspond to like shaded, sized, and labeled blocks in FIG. 7C, and may indicate the loss function, such as MSE, for which the blocks have been trained to reduce; a number of blockwise features (e.g., features illustrated in FIG. 7B) may be used to build a ranking of architectures in the search space ( e.g., targets illustrated in FIG. 7D), an example of which is illustrated in FIG. 7E; the NSR of the feature maps may be between a trained reference neural network 700 and fine-tuned neural network 704; this may be understood as a measure of the distance between F_n and F_xn; the closer F_xn is to F_n, the higher the quality of block B_xn and the lower the average NSR or MSE; training/validation loss of a fine-tuned neural network 704 M_xn may be used as an accuracy target to fit the accuracy predictor 706; allow for mixing different cell-types, activation functions, quantization levels and/or attention mechanisms, while still being able to model their accuracy reliably; allow for building an accuracy model for a single search space with different attention mechanisms, activation functions, cell-types, channels, quantization settings; allow for adding quantized blocks into the search space, which may build an accuracy model for quantized networks directly; the accuracy predictor 706 may be understood as a coarse sensitivity model that may indicate which blocks require complex implementations in order to build neural networks with high accuracy; ¶¶ [0088]-[0097] with FIG. 8: building multiple neural network architectures for various use-cases; in block 802, train any number and combination of search blocks from the search space using knowledge distillation with a trained reference neural network; the search blocks may be trained using blockwise knowledge distillation in which the search blocks are associated with blocks of the trained reference neural network and trained using the associated blocks to converge a loss function of each search block with a loss function of an associated block; the search blocks may be trained using conventional knowledge distillation in which a neural network architecture having search blocks may be trained using the trained reference neural network to converge a loss function of the neural network architecture with a loss function of the trained reference neural network; e.g., the loss function of the blocks and/or the neural networks may converge while training the search blocks to an accuracy level, up to full accuracy; e.g., convergence of the loss function of the blocks and/or the neural networks may be to within a loss function threshold; in block 804, extract quality features and initialize weights for each search block through knowledge distillation; extracting a quality feature may include determining an NSR of the feature maps between a trained reference block and/or neural network and a trained search block and/or neural network architecture; in block 806, store the quality features and initialize weights for each trained search block and/or neural network architecture; in block 808, select a sub-set of neural networks from the search space as targets for building an accuracy model; in block 810, train fine-tune the sub-set of neural networks as targets for building an accuracy model using knowledge distillation; in block 812, extract targets for building an accuracy model using the fine-tuned neural networks; the targets may be accuracy measurements of the fine-tuned neural networks; accuracy may be measured using the NSR of the feature maps between a trained reference neural network and fine-tuned neural network where the lower the average NSR of the blocks the higher the accuracy; accuracy may be measured using SNR of the feature maps between a trained reference neural network and fine-tuned neural network where the lower the average SNR of the blocks the higher the accuracy; in block 814, generate an accuracy predictor for neural networks from the search space using the quality features and targets; ¶ [0103] with FIG. 11: any neural network 1102 from the search space (e.g., non-shaded blocks 1, 2, 3, ..., Nin neural networks 702 in FIG. 7A) may be fine-tuned using a trained neural network 1104 having blocks trained using knowledge distillation such as shaded blocks 1, 2, 3, ... , N, for which blocks having shading, sizing, and labels may correspond to like shaded, sized, and labeled blocks in FIG. 7F, neural networks 704 in FIG. 7C, or accuracy predictor 706 in FIG. 7F; the neural network 1102 from the search space and initialized with weights from the knowledge distillation of the search blocks may be fine-tuned using by using end-to-end knowledge distillation using the trained reference neural network 1104; ¶¶ [0104]-[0107] with FIG. 12: fine-tuning a neural network architecture from a search-space for distilling neural network architectures; in block 1202, select a neural network from the search space; the neural network from the search space may include any number and combination of search blocks; select the neural network based on the selection of search blocks suited to the criteria; in block 1204, initialize the neural network from the search space using weights from the training of the search blocks using knowledge distillation; in block 1206, fine-tune the neural network having search blocks initialized with the weight from the knowledge distillation by using knowledge distillation with a trained reference neural network; ¶ [0117] with FIG. 16: applying a distilled neural network architecture from a search-space for refining and compressing a trained neural network; a trained neural network 1600 may be modified to use trained search blocks (e.g., shaded blocks 1, 2, 3, 5, 7, 8, 9 in neural network 1602) from the search space to refine and compress the trained neural network 1600; refining the trained neural network may retain whole the structure of the trained neural network, replacing the blocks (e.g., non-shaded blocks 1, 2, 3, 5, 7, 8, 9 in neural network 1600) within the trained neural network with trained blocks (e.g., shaded blocks 1, 2, 3, 5, 7, 8, 9 in neural network 1602) from the search space, which may be more efficient on hardware; to accomplish this, the trained neural network 1600 may be used as the reference neural network in the search space, and as the reference neural network for training the search blocks with knowledge distillation; refining the neural network 1600 may involve designing a distilled neural network 1602 as a scenario specific version of the neural network; the parameters for the search blocks in the search space and/or the search for search blocks to generate a distilled neural network 1602 may be configured for a specific scenario, such as use-case, hardware configuration, software configuration, etc.; as such, distilling a neural network 1600 based on the reference neural network may result in a scenario specific version of the reference neural network, which may be a distilled neural network 1602 configured to perform better for the scenario; refining the neural network 1600 may involve designing a distilled neural network 1602 as a compressed version of the neural network 1600; the parameters for the search blocks in the search space may be set to be smaller than the corresponding blocks in the reference neural network; as such, distilling a neural network 1600 based on the reference neural network may result in a compressed version of the reference neural network 1600, which may be a smaller and more efficient distilled neural network 1602).
Claims 6 and 21
MOONS discloses all the elements as stated in Claims 1 and 16 respectively and further discloses wherein evaluating the cost function to respectively select the one or more candidate compression schemes comprises evaluating, by the computing system, a cost function that evaluates changes in an accuracy metric and a performance metric associated with compression of a model portion using a candidate compression scheme to respectively select the one or more candidate compression schemes from the one or more sets of compression schemes (MOONS, ¶ [0004]: selecting a second plurality of the blockwise knowledge distillation trained search blocks based on criteria of predicted accuracy using the accuracy predictor and a cost function for implementing the second plurality of the blockwise knowledge distillation trained search blocks; ¶ [0011]: selecting the neural network of the search space based on a search of the blockwise knowledge distillation trained search blocks using a criterion of predicted accuracy using the accuracy predictor and a cost function for implementing blockwise knowledge distillation trained search blocks of the neural network; ¶ [0033]: a search for identifying knowledge distillation trained neural network blocks from the search space based on predicted accuracy and any number and combination of cost functions for the search blocks; ¶¶ [0039]-[0045]: reduce the resource requirements of designing energy efficient on-platform neural networks both in terms of man-hours and compute costs; a quick evolutionary search-phase extracting a front of architectures in terms of accuracy and some on-target efficiency metric (e.g., number of operations, latency, energy consumption, etc.), such as a Pareto-optimal front; the evolutionary search may be performed using the prior accuracy model together with hardware measurements in the loop and may be repeated quickly many times, amortizing the resource-costs of building the accuracy model; a brief search phase may find latency-accuracy Pareto-optimal neural network architectures for any use-case or hardware platform by running a 2D-optimization algorithm using the accuracy model together with hardware measurements in the loop, which may be quickly rerun whenever anything changes to a use-case, hardware platform, or platform software version; the upfront costs of building an accuracy model by performing blockwise knowledge distillation, which requires training partial neural networks (blocks) may be amortized by the ability to reuse the accuracy model multiple times for various circumstances; using a built accuracy model, the cost of a search scales linearly with the number of different use-cases, multiple hardware configurations, or multiple platform software configurations to which the accuracy model is applied to find efficient neural networks; the neural network designed using an accuracy model, evolutionary search, and finetuning may also be applicable to designing, compressing, improving, and/or selecting hardware for other neural networks; ¶¶ [0096]-[0097] with FIG. 9: given the accuracy model and the collection of fine-tuned blocks, a search algorithm may be implemented to identify search blocks and/or neural networks in the search space that achieve a particular balance between predicted accuracy and a cost for execution of the fine-tuned blocks and/or neural networks; the search algorithm may be an evolutionary algorithm executed to find Pareto-optimal search blocks and/or neural networks in the search space that achieve a certain model accuracy, up to maximum model accuracy, and a certain a target cost function, such as a minimum target cost function; a two-dimensional criterion may be used for the search, in which predicted accuracy is balanced with hardware latency; the blocks of the search space, represented here by points, may be plotted based on the predicted accuracy for each block using the accuracy predictor and based on measured and/or predicted hardware latency for implementing the blocks; the dashed line may represent a frontier at which the search may identify as values for which blocks would best suit the criteria; the points plotted closest to the line may represent blocks which the search may identify as best suit the criteria; ¶¶ [0099]-[0102] with FIG. 10: in block 1002, set search parameters; search parameters may include a predicted accuracy and a cost function for executing search blocks; the cost function may be scenario-agnostic, such to as a number of operations or a number of parameters in a neural network; the accuracy predictor may be configured to account for scenario specific parameters for the search; the cost function may be scenario-aware, wherein the scenarios may include hardware configurations, software versions, input data parameters, a number of operations or a number of parameters in a neural network, on-device latency, throughput, energy, etc.; set the search parameters based on programmed parameters retrieved from a memory; set the search parameters based on a user input; in block 1004, determine search parameter values for search blocks; calculate cost function values for implementing the search blocks; measure cost function values for implementing the search blocks; retrieve cost function values for implementing the search blocks from the memory accessible to the processor; receive cost function values for implementing the search blocks via user input; the cost function may be scenario-agnostic; the cost function may be scenario-aware, wherein the scenarios may include hardware configurations, software versions, input data parameters, a number of operations or a number of parameters in a neural network, on-device latency, throughput, energy, etc.; in block 1006, determine search blocks suited to the criteria; the search algorithm may be an evolutionary search algorithm; the search algorithm may be configured to identify any number and combination of search blocks, such as N search blocks, suited to the criteria; the configuration of the search algorithm may control whether search parameter values for search blocks may be interpreted as suited to the criteria; the search algorithm may be a Y-dimensional search algorithm, where Y is any integer greater than 1, executed to find search blocks that achieve a certain model accuracy, up to maximum model accuracy, and any number and combination of certain target cost functions, such as a minimum target cost function; e.g., the search criteria may be a two-dimensional search criteria balancing predicted accuracy and cost for implementing search blocks; the search criteria may be to identify Pareto-optimal search blocks for inclusion in a neural network; in block 1008, select the search blocks suited to the criteria; select any number and combination of search blocks suited to the criteria, such as N search blocks suited to the criteria; the processor may select the search blocks best suited to the criteria; select search blocks suited to the criteria within selection parameters; the selection parameters may include a function of the cost function for any number and combination of the search blocks suited to the criteria; the selection parameters may include a function of the predicted accuracy for any number and combination of the search blocks suited to the criteria; ¶¶ [0110]-[0116] with FIGS. 13A-C and 14: implement a sampling algorithm 1302, which may be implemented for a search for suitable search blocks; the sampling algorithm may take a predicted accuracy value 1304, 1334, 1354 from the accuracy predictor 1306, 1336, 1356 (e.g., original task accuracy predictor, target task accuracy predictor, ImageNet accuracy predictor) as a proxy metric for accuracy of the process implementing the downstream task; the sampling algorithm 1302 may use the proxy information to update the search for search blocks to generate a distilled neural network; the sampling algorithm may also take cost function metric 1308, 1358 (e.g., measured latency, predicted latency); to determine a cost function value, the distilled neural network may be embedded into a larger neural network, which may include parts 1310, 1360 (e.g., device, target hardware (HW)) for implementing the downstream task; the cost function metric may be measured for the larger neural network including the distilled neural network implemented on a hardware or hardware simulator; the sampling algorithm may use the cost function metric to update the search for search blocks to be used for a neural network configured to contribute to the downstream task; the search for suitable search blocks may be executed based on a predicted accuracy of a task generated by an accuracy predictor and a measured latency 1308 of sampled search blocks measured at a device 1310; the sampling algorithm may use the predicted accuracy 1354 and a predicted latency 1358 of using sampled neural network algorithms on target hardware 1360 for implementing the downstream task to search for suitable search blocks; the sampling algorithm may take a target performance value 1404 (e.g., Target Dataset Performance, which may include accuracy) from the target performance predictor 1406 and output end-to-end model encoding; use the target performance information to update the search for search blocks to update the search for search blocks to generate a distilled neural network; the sampling algorithm may also take cost function metric 1408 (e.g., predicted latency); to determine a cost function value, the distilled neural network may be embedded into a larger neural network, which may include parts 1410 (e.g., Target HW) for implementing the downstream task; the cost function metric may be measured for the larger neural network including the distilled neural network implemented on a hardware or hardware simulator.; the sampling algorithm may use the cost function metric to update the search for search blocks to be used for a neural network configured to contribute to the downstream task).
Claim 7
MOONS discloses all the elements as stated in Claim 6 and further discloses wherein the cost function evaluates changes in the accuracy metric and the performance metric using a combinatorial search space (MOONS, ¶ [0033]: a search for identifying knowledge distillation trained neural network blocks from the search space based on predicted accuracy and any number and combination of cost functions for the search blocks; ¶¶ [0039]-[0045]: reduce the resource requirements of designing energy efficient on-platform neural networks both in terms of man-hours and compute costs; a quick evolutionary search-phase extracting a front of architectures in terms of accuracy and some on-target efficiency metric (e.g., number of operations, latency, energy consumption, etc.), such as a Pareto-optimal front; the evolutionary search may be performed using the prior accuracy model together with hardware measurements in the loop and may be repeated quickly many times, amortizing the resource-costs of building the accuracy model; a brief search phase may find latency-accuracy Pareto-optimal neural network architectures for any use-case or hardware platform by running a 2D-optimization algorithm using the accuracy model together with hardware measurements in the loop, which may be quickly rerun whenever anything changes to a use-case, hardware platform, or platform software version; the upfront costs of building an accuracy model by performing blockwise knowledge distillation, which requires training partial neural networks (blocks) may be amortized by the ability to reuse the accuracy model multiple times for various circumstances; using a built accuracy model, the cost of a search scales linearly with the number of different use-cases, multiple hardware configurations, or multiple platform software configurations to which the accuracy model is applied to find efficient neural networks; the neural network designed using an accuracy model, evolutionary search, and finetuning may also be applicable to designing, compressing, improving, and/or selecting hardware for other neural networks; ¶ [0071]: the varying parameters may be used to define the neural network architectures included in the search-space, which may include varying cell-type (e.g., style of convolutions), attention mechanisms, kernel sizes, number of layers per block, activation functions, expansion rates, network width, network depth, etc.; any number and combination of the varying parameters and the constraints of the reference neural network may define a block or neural network architecture in the search space; ¶ [0087]: allow for mixing different cell-types, activation functions, quantization levels and/or attention mechanisms, while still being able to model their accuracy reliably; allow for building an accuracy model for a single search space with different attention mechanisms, activation functions, cell-types, channels, quantization settings; allow for adding quantized blocks into the search space, which may build an accuracy model for quantized networks directly; the accuracy predictor 706 may be understood as a coarse sensitivity model that may indicate which blocks require complex implementations in order to build neural networks with high accuracy; ¶ [0089]: train any number and combination of search blocks from the search space using knowledge distillation with a trained reference neural network; ¶ [0092]: each of neural networks of the sub-set may include any combination of trained search blocks; the neural networks of the sub-set may include any combination of trained neural networks of the search space; ¶¶ [0096]-[0097] with FIG. 9: given the accuracy model and the collection of fine-tuned blocks, a search algorithm may be implemented to identify search blocks and/or neural networks in the search space that achieve a particular balance between predicted accuracy and a cost for execution of the fine-tuned blocks and/or neural networks; the search algorithm may be an evolutionary algorithm executed to find Pareto-optimal search blocks and/or neural networks in the search space that achieve a certain model accuracy, up to maximum model accuracy, and a certain a target cost function, such as a minimum target cost function; a two-dimensional criterion may be used for the search, in which predicted accuracy is balanced with hardware latency; the blocks of the search space, represented here by points, may be plotted based on the predicted accuracy for each block using the accuracy predictor and based on measured and/or predicted hardware latency for implementing the blocks; the dashed line may represent a frontier at which the search may identify as values for which blocks would best suit the criteria; the points plotted closest to the line may represent blocks which the search may identify as best suit the criteria; ¶¶ [0099]-[0102] with FIG. 10: in block 1002, set search parameters; search parameters may include a predicted accuracy and a cost function for executing search blocks; the cost function may be scenario-agnostic, such to as a number of operations or a number of parameters in a neural network; the accuracy predictor may be configured to account for scenario specific parameters for the search; the cost function may be scenario-aware, wherein the scenarios may include hardware configurations, software versions, input data parameters, a number of operations or a number of parameters in a neural network, on-device latency, throughput, energy, etc.; set the search parameters based on programmed parameters retrieved from a memory; set the search parameters based on a user input; in block 1004, determine search parameter values for search blocks; calculate cost function values for implementing the search blocks; measure cost function values for implementing the search blocks; retrieve cost function values for implementing the search blocks from the memory accessible to the processor; receive cost function values for implementing the search blocks via user input; the cost function may be scenario-agnostic; the cost function may be scenario-aware, wherein the scenarios may include hardware configurations, software versions, input data parameters, a number of operations or a number of parameters in a neural network, on-device latency, throughput, energy, etc.; in block 1006, determine search blocks suited to the criteria; the search algorithm may be an evolutionary search algorithm; the search algorithm may be configured to identify any number and combination of search blocks, such as N search blocks, suited to the criteria; the configuration of the search algorithm may control whether search parameter values for search blocks may be interpreted as suited to the criteria; the search algorithm may be a Y-dimensional search algorithm, where Y is any integer greater than 1, executed to find search blocks that achieve a certain model accuracy, up to maximum model accuracy, and any number and combination of certain target cost functions, such as a minimum target cost function; e.g., the search criteria may be a two-dimensional search criteria balancing predicted accuracy and cost for implementing search blocks; the search criteria may be to identify Pareto-optimal search blocks for inclusion in a neural network; in block 1008, select the search blocks suited to the criteria; select any number and combination of search blocks suited to the criteria, such as N search blocks suited to the criteria; the processor may select the search blocks best suited to the criteria; select search blocks suited to the criteria within selection parameters; the selection parameters may include a function of the cost function for any number and combination of the search blocks suited to the criteria; the selection parameters may include a function of the predicted accuracy for any number and combination of the search blocks suited to the criteria; ¶¶ [0110]-[0116] with FIGS. 13A-C and 14: implement a sampling algorithm 1302, which may be implemented for a search for suitable search blocks; the sampling algorithm may take a predicted accuracy value 1304, 1334, 1354 from the accuracy predictor 1306, 1336, 1356 (e.g., original task accuracy predictor, target task accuracy predictor, ImageNet accuracy predictor) as a proxy metric for accuracy of the process implementing the downstream task; the sampling algorithm 1302 may use the proxy information to update the search for search blocks to generate a distilled neural network; the sampling algorithm may also take cost function metric 1308, 1358 (e.g., measured latency, predicted latency); to determine a cost function value, the distilled neural network may be embedded into a larger neural network, which may include parts 1310, 1360 (e.g., device, target hardware (HW)) for implementing the downstream task; the cost function metric may be measured for the larger neural network including the distilled neural network implemented on a hardware or hardware simulator; the sampling algorithm may use the cost function metric to update the search for search blocks to be used for a neural network configured to contribute to the downstream task; the search for suitable search blocks may be executed based on a predicted accuracy of a task generated by an accuracy predictor and a measured latency 1308 of sampled search blocks measured at a device 1310; the sampling algorithm may use the predicted accuracy 1354 and a predicted latency 1358 of using sampled neural network algorithms on target hardware 1360 for implementing the downstream task to search for suitable search blocks; the sampling algorithm may take a target performance value 1404 (e.g., Target Dataset Performance, which may include accuracy) from the target performance predictor 1406 and output end-to-end model encoding; use the target performance information to update the search for search blocks to update the search for search blocks to generate a distilled neural network; the sampling algorithm may also take cost function metric 1408 (e.g., predicted latency); to determine a cost function value, the distilled neural network may be embedded into a larger neural network, which may include parts 1410 (e.g., Target HW) for implementing the downstream task; the cost function metric may be measured for the larger neural network including the distilled neural network implemented on a hardware or hardware simulator.; the sampling algorithm may use the cost function metric to update the search for search blocks to be used for a neural network configured to contribute to the downstream task).
Claim 8
MOONS discloses all the elements as stated in Claim 7 and further discloses wherein the performance metric is indicative of whether, following the compression of the model portion, the compressed machine learned model meets a suitability criterion for implementation using a specific data processing system having lesser computational capacity than the computing system (MOONS, ¶¶ [0096]-[0097] with FIG. 9: ¶¶ [0096]-[0097] with FIG. 9: given the accuracy model and the collection of fine-tuned blocks, a search algorithm may be implemented to identify search blocks and/or neural networks in the search space that achieve a particular balance between predicted accuracy and a cost for execution of the fine-tuned blocks and/or neural networks; the search algorithm may be an evolutionary algorithm executed to find Pareto-optimal search blocks and/or neural networks in the search space that achieve a certain model accuracy, up to maximum model accuracy, and a certain a target cost function, such as a minimum target cost function; a two-dimensional criterion may be used for the search, in which predicted accuracy is balanced with hardware latency; the blocks of the search space, represented here by points, may be plotted based on the predicted accuracy for each block using the accuracy predictor and based on measured and/or predicted hardware latency for implementing the blocks; the dashed line may represent a frontier at which the search may identify as values for which blocks would best suit the criteria; the points plotted closest to the line may represent blocks which the search may identify as best suit the criteria; ¶¶ [0101]-[0102] with 1006-1008 in FIG. 10: In block 1006, determine search blocks suited to the criteria; the search algorithm may be an evolutionary search algorithm; the search algorithm may be configured to identify any number and combination of search blocks, such as N search blocks, suited to the criteria; the configuration of the search algorithm may control whether search parameter values for search blocks may be interpreted as suited to the criteria; the search algorithm may be a Y-dimensional search algorithm, where Y is any integer greater than 1, executed to find search blocks that achieve a certain model accuracy, up to maximum model accuracy, and any number and combination of certain target cost functions, such as a minimum target cost function; e.g., the search criteria may be a two-dimensional search criteria balancing predicted accuracy and cost for implementing search blocks; the search criteria may be to identify Pareto-optimal search blocks for inclusion in a neural network; in block 1008, select the search blocks suited to the criteria; select any number and combination of search blocks suited to the criteria, such as N search blocks suited to the criteria; the processor may select the search blocks best suited to the criteria; select search blocks suited to the criteria within selection parameters; the selection parameters may include a function of the cost function for any number and combination of the search blocks suited to the criteria; the selection parameters may include a function of the predicted accuracy for any number and combination of the search blocks suited to the criteria; ¶ [0105] in 1202 in FIG. 12: in block 1202, select a neural network from the search space; the neural network from the search space may include any number and combination of search blocks; select the neural network based on the selection of search blocks suited to the criteria).
Claim 9
MOONS discloses all the elements as stated in Claim 8 and further discloses wherein the specific data processing system is a mobile or wearable computing device (MOONS, ¶ [0002]: energy-efficiency of artificial intelligence (AI) workloads, such as running neural networks for visual recognition, is key for mobile and automotive hardware platforms; ¶ [0034]: mobile device, cellular telephones, smartphones, portable computing devices, personal or mobile multi-media players, personal data assistants (PDA's), laptop computers, tablet computers, smartbooks; IoT devices, palm-top computers, wireless electronic mail receivers, multimedia Internet enabled cellular telephones, connected vehicles, wireless gaming controllers; ¶¶ [0039]-[0045]: reduce the resource requirements of designing energy efficient on-platform neural networks both in terms of man-hours and compute costs; a quick evolutionary search-phase extracting a front of architectures in terms of accuracy and some on-target efficiency metric (e.g., number of operations, latency, energy consumption, etc.), such as a Pareto-optimal front; ¶¶[0096] and [0099]-[0100]: the cost function may be scenario-aware, wherein the scenarios may include hardware configurations, software versions, input data parameters, a number of operations or a number of parameters in a neural network, on-device latency, throughput, energy, etc.; ¶ [0117] with FIG. 16: a trained neural network 1600 may be modified to use trained search blocks (e.g., shaded blocks 1, 2, 3, 5, 7, 8, 9 in neural network 1602) from the search space to refine and compress the trained neural network 1600; refining the trained neural network may retain whole the structure of the trained neural network, replacing the blocks (e.g., non-shaded blocks 1, 2, 3, 5, 7, 8, 9 in neural network 1600) within the trained neural network with trained blocks (e.g., shaded blocks 1, 2, 3, 5, 7, 8, 9 in neural network 1602) from the search space, which may be more efficient on hardware; the parameters for the search blocks in the search space may be set to be smaller than the corresponding blocks in the reference neural network; as such, distilling a neural network 1600 based on the reference neural network may result in a compressed version of the reference neural network 1600, which may be a smaller and more efficient distilled neural network 1602).
Claim 10
MOONS discloses all the elements as stated in Claim 1 and further discloses wherein the machine learned model is adapted to perform a computational task based on input data which is selected from image data, speech data or sensor data (MOONS, ¶ [0044]: the designed neural network may be an image classification neural network and may be applicable for use in computer vision tasks, such as object detection, semantic segmentation models, super-resolution models, video classification, video segmentation, etc.; ¶¶ [0047] and [0049] with FIG. 1A: determining whether an image contains a specific item (e.g., dog, cat, etc.); the neural network 100 may be configured to receive pixels of an image (i.e., input values) in the first layer, and generate outputs indicating the presence of different low-level features (e.g., lines, edges, etc.) in the image; in training of a neural network for image recognition, lines may be combined into shapes, shapes may be combined into sets of shapes, etc., and at the output layer, the neural network 100 may generate a probability value that indicates whether a particular object is present in the image; ¶ [0109]: a distilled neural network may be configured for image recognition and a downstream task may include object detection, semantic image segmentation, video classification, etc.; ¶ [0114] with FIG. 13C: the original task may be image classification on the ImageNet dataset and a downstream task may be object detection on the common objects in context (COCO) dataset).
Claim 11
MOONS discloses all the elements as stated in Claim 10 and further discloses wherein the computational task is to generate classification data indicative of which one of a predetermined plurality of categories matches content of the input data (MOONS, ¶ [0044]: the designed neural network may be an image classification neural network and may be applicable for use in computer vision tasks, such as object detection, semantic segmentation models, super-resolution models, video classification, video segmentation, etc.; ¶¶ [0047] and [0049] with FIG. 1A: determining whether an image contains a specific item (e.g., dog, cat, etc.); the neural network 100 may be configured to receive pixels of an image (i.e., input values) in the first layer, and generate outputs indicating the presence of different low-level features (e.g., lines, edges, etc.) in the image; in training of a neural network for image recognition, lines may be combined into shapes, shapes may be combined into sets of shapes, etc., and at the output layer, the neural network 100 may generate a probability value that indicates whether a particular object is present in the image; ¶ [0068]: structural restrictions on the network architectures may be limited, such as being limited to spatial dimensions of the tensors input and output from the blocks matching those of the reference model; ¶ [0077]: the weights W xn of B_xn may be updated using conventional knowledge distillation, by minimizing a loss function defined by the tasks ground truth and the output classifier of the trained reference neural network 700; ¶ [0109] with FIG. 15: a neural network 1500 having a distilled neural network 1502 (e.g., backbone) feeding ( e.g., via a process 1504, such as feature fusion) into a neural network 1506, 1508 (e.g., classifier and/or regressor) for implementing a downstream task; a distilled neural network may be configured for image recognition and a downstream task may include object detection, semantic image segmentation, video classification, etc.; ¶ [0114] with FIG. 13C: the original task may be image classification on the ImageNet dataset and a downstream task may be object detection on the common objects in context (COCO) dataset).
Claims 12 and 23
MOONS discloses all the elements as stated in Claims 1 and 16 respectively and further discloses wherein obtaining the data descriptive of the one or more respective sets of compression schemes comprises: obtaining, by the computing system from a user device, data descriptive of selection of the one or more respective sets of compression schemes for the one or more model portions of a plurality of model portions of a machine-learned model (MOONS, ¶¶ [0069]-[0071] in FIG. 6: a method of defining a search-space for distilling neural network architectures; in block 602, define a reference neural network for a search-space; the reference neural network may be used to define constraints of the neural network architectures included in the search-space; receive parameters for defining the reference neural network from a user input or from a memory accessible to the processor; in block 604, define varying parameters for the search space; the varying parameters may be used to define the neural network architectures included in the search-space, which may include varying cell-type (e.g., style of convolutions), attention mechanisms, kernel sizes, number of layers per block, activation functions, expansion rates, network width, network depth, etc.; any number and combination of the varying parameters and the constraints of the reference neural network may define a block or neural network architecture in the search space; receive varying parameters from a user input or from a memory accessible to the processor; ¶ [0092] with 808 in FIG. 8: in block 808, select a sub-set of neural networks from the search space as targets for building an accuracy model; select the sub-set of neural networks from the search space based on a programmed algorithm, heuristic, technique, criteria, etc.; select the sub-set of neural networks from the search space based on a user input; each of neural networks of the sub-set may include any combination of trained search blocks; the neural networks of the sub-set may include any combination of trained neural networks of the search space; ¶¶ [0099]-[0100] with 1002-1004 in FIG. 10: in block 1002, set search parameters; search parameters may include a predicted accuracy and a cost function for executing search blocks; the cost function may be scenario-agnostic, such to as a number of operations or a number of parameters in a neural network; the accuracy predictor may be configured to account for scenario specific parameters for the search; the cost function may be scenario-aware, wherein the scenarios may include hardware configurations, software versions, input data parameters, a number of operations or a number of parameters in a neural network, on-device latency, throughput, energy, etc.; set the search parameters based on programmed parameters retrieved from a memory; set the search parameters based on a user input; in block 1004, determine search parameter values for search blocks; calculate cost function values for implementing the search blocks; measure cost function values for implementing the search blocks; retrieve cost function values for implementing the search blocks from the memory accessible to the processor; receive cost function values for implementing the search blocks via user input; ¶ [0117] with FIG. 16: applying a distilled neural network architecture from a search-space for refining and compressing a trained neural network; a trained neural network 1600 may be modified to use trained search blocks (e.g., shaded blocks 1, 2, 3, 5, 7, 8, 9 in neural network 1602) from the search space to refine and compress the trained neural network 1600; refining the trained neural network may retain whole the structure of the trained neural network, replacing the blocks (e.g., non-shaded blocks 1, 2, 3, 5, 7, 8, 9 in neural network 1600) within the trained neural network with trained blocks (e.g., shaded blocks 1, 2, 3, 5, 7, 8, 9 in neural network 1602) from the search space, which may be more efficient on hardware; to accomplish this, the trained neural network 1600 may be used as the reference neural network in the search space, and as the reference neural network for training the search blocks with knowledge distillation; refining the neural network 1600 may involve designing a distilled neural network 1602 as a scenario specific version of the neural network; the parameters for the search blocks in the search space and/or the search for search blocks to generate a distilled neural network 1602 may be configured for a specific scenario, such as use-case, hardware configuration, software configuration, etc.; as such, distilling a neural network 1600 based on the reference neural network may result in a scenario specific version of the reference neural network, which may be a distilled neural network 1602 configured to perform better for the scenario; refining the neural network 1600 may involve designing a distilled neural network 1602 as a compressed version of the neural network 1600; the parameters for the search blocks in the search space may be set to be smaller than the corresponding blocks in the reference neural network; as such, distilling a neural network 1600 based on the reference neural network may result in a compressed version of the reference neural network 1600, which may be a smaller and more efficient distilled neural network 1602).
Claims 13 and 24
MOONS discloses all the elements as stated in Claims 1 and 16 respectively and further discloses wherein a portion of a machine-learned model comprises one or more tensors and/or one or more layers of the machine-learned model (MOONS, ¶ [0068] with FIG. 5: a search-space may be set to provide every potential neural network architecture as a neural network of N blocks 502 (e.g., block 1 502a, block 2 502b, block 3 502c, block 4 502d, block 5 502e ), which may be referred to herein as search blocks, where N is a positive integer; every block 502 B_n may be any of X_n block implementations, with varying cell type ("cell type") (e.g., style of convolutions), attention mechanisms ("attention"), kernel sizes ("kernel"), number of layers per block, activation functions ("activation"), expansion rates ("expand"), network width ("width scale"), network depth ("depth"), output channels ("ch"), etc.; structural restrictions on the network architectures may be limited, such as being limited to spatial dimensions of the tensors input and output from the blocks matching those of the reference model).
Claim 14
MOONS discloses all the elements as stated in Claim 1 and further discloses employing a compressed machine-learned model to perform a computational task of processing input data to form output data, wherein the input data is selected from image data, speech data or sensor data (MOONS, ¶ [0044]: the designed neural network may be an image classification neural network and may be applicable for use in computer vision tasks, such as object detection, semantic segmentation models, super-resolution models, video classification, video segmentation, etc.; ¶¶ [0047] and [0049] with FIG. 1A: determining whether an image contains a specific item (e.g., dog, cat, etc.); the neural network 100 may be configured to receive pixels of an image (i.e., input values) in the first layer, and generate outputs indicating the presence of different low-level features (e.g., lines, edges, etc.) in the image; in training of a neural network for image recognition, lines may be combined into shapes, shapes may be combined into sets of shapes, etc., and at the output layer, the neural network 100 may generate a probability value that indicates whether a particular object is present in the image; ¶ [0109]: a distilled neural network may be configured for image recognition and a downstream task may include object detection, semantic image segmentation, video classification, etc.; ¶ [0114] with FIG. 13C: the original task may be image classification on the ImageNet dataset and a downstream task may be object detection on the common objects in context (COCO) dataset)
Claim 15
MOONS discloses all the elements as stated in Claim 14 and further discloses in which the computation task is to generate classification data indicative of which one of a predetermined plurality of categories matches content of the input data (MOONS, ¶ [0044]: the designed neural network may be an image classification neural network and may be applicable for use in computer vision tasks, such as object detection, semantic segmentation models, super-resolution models, video classification, video segmentation, etc.; ¶¶ [0047] and [0049] with FIG. 1A: determining whether an image contains a specific item (e.g., dog, cat, etc.); the neural network 100 may be configured to receive pixels of an image (i.e., input values) in the first layer, and generate outputs indicating the presence of different low-level features (e.g., lines, edges, etc.) in the image; in training of a neural network for image recognition, lines may be combined into shapes, shapes may be combined into sets of shapes, etc., and at the output layer, the neural network 100 may generate a probability value that indicates whether a particular object is present in the image; ¶ [0068]: structural restrictions on the network architectures may be limited, such as being limited to spatial dimensions of the tensors input and output from the blocks matching those of the reference model; ¶ [0077]: the weights W xn of B_xn may be updated using conventional knowledge distillation, by minimizing a loss function defined by the tasks ground truth and the output classifier of the trained reference neural network 700; ¶ [0109] with FIG. 15: a neural network 1500 having a distilled neural network 1502 (e.g., backbone) feeding ( e.g., via a process 1504, such as feature fusion) into a neural network 1506, 1508 (e.g., classifier and/or regressor) for implementing a downstream task; a distilled neural network may be configured for image recognition and a downstream task may include object detection, semantic image segmentation, video classification, etc.; ¶ [0114] with FIG. 13C: the original task may be image classification on the ImageNet dataset and a downstream task may be object detection on the common objects in context (COCO) dataset).
Claim 22
MOONS discloses all the elements as stated in Claim 21 and further discloses wherein the cost function evaluates changes in the accuracy metric and the performance metric using a topology search space or a portion-wise search space (MOONS, ¶ [0033]: a search for identifying knowledge distillation trained neural network blocks from the search space based on predicted accuracy and any number and combination of cost functions for the search blocks; ¶¶ [0039]-[0045]: reduce the resource requirements of designing energy efficient on-platform neural networks both in terms of man-hours and compute costs; a quick evolutionary search-phase extracting a front of architectures in terms of accuracy and some on-target efficiency metric (e.g., number of operations, latency, energy consumption, etc.), such as a Pareto-optimal front; the evolutionary search may be performed using the prior accuracy model together with hardware measurements in the loop and may be repeated quickly many times, amortizing the resource-costs of building the accuracy model; a brief search phase may find latency-accuracy Pareto-optimal neural network architectures for any use-case or hardware platform by running a 2D-optimization algorithm using the accuracy model together with hardware measurements in the loop, which may be quickly rerun whenever anything changes to a use-case, hardware platform, or platform software version; the upfront costs of building an accuracy model by performing blockwise knowledge distillation, which requires training partial neural networks (blocks) may be amortized by the ability to reuse the accuracy model multiple times for various circumstances; using a built accuracy model, the cost of a search scales linearly with the number of different use-cases, multiple hardware configurations, or multiple platform software configurations to which the accuracy model is applied to find efficient neural networks; the neural network designed using an accuracy model, evolutionary search, and finetuning may also be applicable to designing, compressing, improving, and/or selecting hardware for other neural networks; ¶ [0071]: the varying parameters may be used to define the neural network architectures included in the search-space, which may include varying cell-type (e.g., style of convolutions), attention mechanisms, kernel sizes, number of layers per block, activation functions, expansion rates, network width, network depth, etc.; any number and combination of the varying parameters and the constraints of the reference neural network may define a block or neural network architecture in the search space; ¶ [0082]: the NSR of the feature maps may be between a trained reference neural network 700 and fine-tuned neural network 70; this may be understood as a measure of the distance between F_n and F_xn; the closer F_xn is to F_n, the higher the quality of block B_xn and the lower the average NSR or MSE; ¶ [0085]: the accuracy model may be built using a graph convolutional neural network combining accuracy features and graph features; ¶ [0087]: allow for mixing different cell-types, activation functions, quantization levels and/or attention mechanisms, while still being able to model their accuracy reliably; allow for building an accuracy model for a single search space with different attention mechanisms, activation functions, cell-types, channels, quantization settings; allow for adding quantized blocks into the search space, which may build an accuracy model for quantized networks directly; the accuracy predictor 706 may be understood as a coarse sensitivity model that may indicate which blocks require complex implementations in order to build neural networks with high accuracy; ¶ [0089]: train any number and combination of search blocks from the search space using knowledge distillation with a trained reference neural network; ¶ [0092]: each of neural networks of the sub-set may include any combination of trained search blocks; the neural networks of the sub-set may include any combination of trained neural networks of the search space; ¶ [0095]: the accuracy model may be built using a graph convolutional neural network combining accuracy features and graph features; the accuracy predictor may be used for any neural network in the search space; ¶¶ [0096]-[0097] with FIG. 9: given the accuracy model and the collection of fine-tuned blocks, a search algorithm may be implemented to identify search blocks and/or neural networks in the search space that achieve a particular balance between predicted accuracy and a cost for execution of the fine-tuned blocks and/or neural networks; the search algorithm may be an evolutionary algorithm executed to find Pareto-optimal search blocks and/or neural networks in the search space that achieve a certain model accuracy, up to maximum model accuracy, and a certain a target cost function, such as a minimum target cost function; a two-dimensional criterion may be used for the search, in which predicted accuracy is balanced with hardware latency; the blocks of the search space, represented here by points, may be plotted based on the predicted accuracy for each block using the accuracy predictor and based on measured and/or predicted hardware latency for implementing the blocks; the dashed line may represent a frontier at which the search may identify as values for which blocks would best suit the criteria; the points plotted closest to the line may represent blocks which the search may identify as best suit the criteria; ¶¶ [0099]-[0102] with FIG. 10: in block 1002, set search parameters; search parameters may include a predicted accuracy and a cost function for executing search blocks; the cost function may be scenario-agnostic, such to as a number of operations or a number of parameters in a neural network; the accuracy predictor may be configured to account for scenario specific parameters for the search; the cost function may be scenario-aware, wherein the scenarios may include hardware configurations, software versions, input data parameters, a number of operations or a number of parameters in a neural network, on-device latency, throughput, energy, etc.; set the search parameters based on programmed parameters retrieved from a memory; set the search parameters based on a user input; in block 1004, determine search parameter values for search blocks; calculate cost function values for implementing the search blocks; measure cost function values for implementing the search blocks; retrieve cost function values for implementing the search blocks from the memory accessible to the processor; receive cost function values for implementing the search blocks via user input; the cost function may be scenario-agnostic; the cost function may be scenario-aware, wherein the scenarios may include hardware configurations, software versions, input data parameters, a number of operations or a number of parameters in a neural network, on-device latency, throughput, energy, etc.; in block 1006, determine search blocks suited to the criteria; the search algorithm may be an evolutionary search algorithm; the search algorithm may be configured to identify any number and combination of search blocks, such as N search blocks, suited to the criteria; the configuration of the search algorithm may control whether search parameter values for search blocks may be interpreted as suited to the criteria; the search algorithm may be a Y-dimensional search algorithm, where Y is any integer greater than 1, executed to find search blocks that achieve a certain model accuracy, up to maximum model accuracy, and any number and combination of certain target cost functions, such as a minimum target cost function; e.g., the search criteria may be a two-dimensional search criteria balancing predicted accuracy and cost for implementing search blocks; the search criteria may be to identify Pareto-optimal search blocks for inclusion in a neural network; in block 1008, select the search blocks suited to the criteria; select any number and combination of search blocks suited to the criteria, such as N search blocks suited to the criteria; the processor may select the search blocks best suited to the criteria; select search blocks suited to the criteria within selection parameters; the selection parameters may include a function of the cost function for any number and combination of the search blocks suited to the criteria; the selection parameters may include a function of the predicted accuracy for any number and combination of the search blocks suited to the criteria; ¶¶ [0110]-[0116] with FIGS. 13A-C and 14: implement a sampling algorithm 1302, which may be implemented for a search for suitable search blocks; the sampling algorithm may take a predicted accuracy value 1304, 1334, 1354 from the accuracy predictor 1306, 1336, 1356 (e.g., original task accuracy predictor, target task accuracy predictor, ImageNet accuracy predictor) as a proxy metric for accuracy of the process implementing the downstream task; the sampling algorithm 1302 may use the proxy information to update the search for search blocks to generate a distilled neural network; the sampling algorithm may also take cost function metric 1308, 1358 (e.g., measured latency, predicted latency); to determine a cost function value, the distilled neural network may be embedded into a larger neural network, which may include parts 1310, 1360 (e.g., device, target hardware (HW)) for implementing the downstream task; the cost function metric may be measured for the larger neural network including the distilled neural network implemented on a hardware or hardware simulator; the sampling algorithm may use the cost function metric to update the search for search blocks to be used for a neural network configured to contribute to the downstream task; the search for suitable search blocks may be executed based on a predicted accuracy of a task generated by an accuracy predictor and a measured latency 1308 of sampled search blocks measured at a device 1310; the sampling algorithm may use the predicted accuracy 1354 and a predicted latency 1358 of using sampled neural network algorithms on target hardware 1360 for implementing the downstream task to search for suitable search blocks; the sampling algorithm may take a target performance value 1404 (e.g., Target Dataset Performance, which may include accuracy) from the target performance predictor 1406 and output end-to-end model encoding; use the target performance information to update the search for search blocks to update the search for search blocks to generate a distilled neural network; the sampling algorithm may also take cost function metric 1408 (e.g., predicted latency); to determine a cost function value, the distilled neural network may be embedded into a larger neural network, which may include parts 1410 (e.g., Target HW) for implementing the downstream task; the cost function metric may be measured for the larger neural network including the distilled neural network implemented on a hardware or hardware simulator.; the sampling algorithm may use the cost function metric to update the search for search blocks to be used for a neural network configured to contribute to the downstream task).
Independent Claim 25
MOONS discloses one or more non-transitory computer-readable media (MOONS, ¶¶ [0118]-[0119] with 1702 and 1703 in FIG. 17: a server 1700 typically includes a processor 1701 coupled to volatile memory 1702 and a large capacity nonvolatile memory, such as a disk drive 1703; non-transitory processor-readable medium, such as a disk drive 1703) that store instructions (MOONS, ¶ [0119] with FIG. 17: software applications may be stored on non-transitory processor-readable medium, such as a disk drive 1703, before the instructions are accessed and loaded into the processor) that, when executed by one or more processors (MOONS, ¶¶ [0118]-[0119] with 1701 in FIG. 17: a server 1700 typically includes a processor 1701; the processor 1701 may be any programmable microprocessor, microcomputer or multiple processor chip or chip), cause the one or more processors to perform operations (MOONS, ¶ [0119] with FIG.17: the processor 1701 may be any programmable microprocessor, microcomputer or multiple processor chip or chips that may be configured by software instructions (applications) to perform a variety of functions), the operations comprising:
obtaining data descriptive of a user selection of one or more candidate compression schemes from one or more respective sets of compression schemes for compression of one or more respective model portions of a plurality of model portions of an uncompressed machine-learned model (MOONS, ¶¶ [0057]-[0066] with FIGS. 3A-B: the complexity of these computations may be reduced by reducing the number of weights that contribute to the output activation, which may be accomplished by setting the values of select weights to zero; the complexity of these computations may also be reduced by using the same set of weights in the calculation of every output of every processing node in a layer; by using convolution, the neural network layer may compute a weighted sum for each output activation using only a small "neighborhood" of inputs (e.g., by setting all other weights beyond the neighborhood to zero, etc.), and share the same set of weights (or filter) for every output; the use of convolution in multiple layers allows the neural network to employ a very deep hierarchy of layers; the normalization functionality component 306, 316 may be configured to control the input distribution across layers to speed up training and the improve accuracy of the outputs or activations; the pooling functionality components 308, 318 may be configured to reduce the dimensionality of a feature map generated by the convolution functionality component 302, 312 and/or otherwise allow the convolutional neural network 300 to resist small shifts and distortions in values; ¶ [0068] with FIG. 5: a search-space may be set to provide every potential neural network architecture as a neural network of N blocks 502 (e.g., block 1 502a, block 2 502b, block 3 502c, block 4 502d, block 5 502e ), which may be referred to herein as search blocks, where N is a positive integer; every block 502 B_n may be any of X_n block implementations, with varying cell type ("cell type") (e.g., style of convolutions), attention mechanisms ("attention"), kernel sizes ("kernel"), number of layers per block, activation functions ("activation"), expansion rates ("expand"), network width ("width scale"), network depth ("depth"), output channels ("ch"), etc.; structural restrictions on the network architectures may be limited, such as being limited to spatial dimensions of the tensors input and output from the blocks matching those of the reference model; ¶¶ [0069]-[0071] in FIG. 6: a method of defining a search-space for distilling neural network architectures; in block 602, define a reference neural network for a search-space; the reference neural network may be used to define constraints of the neural network architectures included in the search-space; receive parameters for defining the reference neural network from a user input or from a memory accessible to the processor; in block 604, define varying parameters for the search space; the varying parameters may be used to define the neural network architectures included in the search-space, which may include varying cell-type (e.g., style of convolutions), attention mechanisms, kernel sizes, number of layers per block, activation functions, expansion rates, network width, network depth, etc.; any number and combination of the varying parameters and the constraints of the reference neural network may define a block or neural network architecture in the search space; receive varying parameters from a user input or from a memory accessible to the processor; ¶ [0087]: the means to build an accuracy model poses no limitations on the variety of the search space, as opposed to existing NAS methods, which either rely on weight sharing or have a limited search space due to GPU-memory limitations; allow for mixing different cell-types, activation functions, quantization levels and/or attention mechanisms, while still being able to model their accuracy reliably; allow for building an accuracy model for a single search space with different attention mechanisms, activation functions, cell-types, channels, quantization settings; allow for adding quantized blocks into the search space, which may build an accuracy model for quantized networks directly; the accuracy predictor 706 may be understood as a coarse sensitivity model that may indicate which blocks require complex implementations in order to build neural networks with high accuracy; ¶ [0092] with 808 in FIG. 8: in block 808, select a sub-set of neural networks from the search space as targets for building an accuracy model; select the sub-set of neural networks from the search space based on a programmed algorithm, heuristic, technique, criteria, etc.; select the sub-set of neural networks from the search space based on a user input; each of neural networks of the sub-set may include any combination of trained search blocks; the neural networks of the sub-set may include any combination of trained neural networks of the search space; ¶¶ [0098]-[0100] with 1002-1004 in FIG. 10: a method of searching for a neural network architecture from a search-space for distilling neural network architectures; in block 1002, set search parameters; search parameters may include a predicted accuracy and a cost function for executing search blocks; the cost function may be scenario-agnostic, such to as a number of operations or a number of parameters in a neural network; the accuracy predictor may be configured to account for scenario specific parameters for the search; the cost function may be scenario-aware, wherein the scenarios may include hardware configurations, software versions, input data parameters, a number of operations or a number of parameters in a neural network, on-device latency, throughput, energy, etc.; set the search parameters based on programmed parameters retrieved from a memory; set the search parameters based on a user input; in block 1004, determine search parameter values for search blocks; calculate cost function values for implementing the search blocks; measure cost function values for implementing the search blocks; retrieve cost function values for implementing the search blocks from the memory accessible to the processor; receive cost function values for implementing the search blocks via user input; ¶ [0117] with FIG. 16: the parameters for the search blocks in the search space may be set to be smaller than the corresponding blocks in the reference neural network; as such, distilling a neural network 1600 based on the reference neural network may result in a compressed version of the reference neural network 1600, which may be a smaller and more efficient distilled neural network 1602); and
applying the one or more compression schemes to one or more respective model portions of the plurality of model portions of the uncompressed machine-learned model to obtain a compressed machine-learned model (MOONS, ¶¶ [0072]-[0087] with FIGS. 7A-7F: structures and functions for generating an accuracy model using blockwise knowledge distillation for distilling neural network architectures; a trained reference neural network 700 (e.g., non-shaded blocks 1, 2, 3, ... , N) may be used to implement blockwise knowledge distillation to any number and combination of blocks (e.g., shaded blocks 1, 2, 3, ... , N, for which shading, size, and labeling may indicated same of different characteristics between blocks) defined in the search space; the individual blocks of the trained reference neural network 700 may be used to train individual blocks of the neural network architectures 702 (e.g., rows of shaded blocks 702a, 702b, 702c, 702d, 702e, such as likesized and shaded blocks) defined in the search space (e.g., using non-shaded blocks 1, 2, 3, ... , N to train like-labeled shaded blocks 1, 2, 3, ... , N); the individual blocks of the neural network architectures 702 may be trained such that a loss function, such as mean square error (MSE), per-channel Noise-To-Signal ratio (NSR), or any other relevant loss function, between the outputs of the blocks of the trained reference neural network 700 and the individual blocks of the neural network architectures 702 is reduced, such as to within a loss function threshold; e.g., the loss function defined between the blocks of the trained reference neural network 700 and the individual blocks of the neural network architectures 702 may converge to within the loss function threshold while training the individual blocks of the neural network architectures 702; using the trained reference neural network 700, the parameters of all possible blocks B_xn in the search space may be trained; this may allow for (A) extracting quality metrics useful in building supervised regression models for accuracy and (B) initializing their weights for building unsupervised accuracy models, or for quick fine-tuning of models in the search space; training the parameters of all possible blocks B_xn in the search space may be implemented using blockwise knowledge distillation; this process may reduce MSE, per-channel Noise-To-Signal ratio (NSR), or any other relevant loss function of the output features D xn relative to the trained reference neural network 700; training the parameters of all possible blocks B_xn in the search space may be implemented using conventional knowledge distillation; a new neural network may be constructed where as few as 1 out of the N blocks of the trained reference neural network 700 is replaced by block B_xn; the new neural network is then used to train the weights W_xn of B_xn through knowledge distillation; the target parameters may be developed by fine-tuning a number of sampled neural networks 704 (e.g., rows of shaded blocks 704a, 704b, 704c, 704d, such as variably shaded and/or sized blocks) from the search space, an example of which is illustrated in FIG. 7C (e.g., in which blocks having shading, sizing, and labels may correspond to like shaded, sized, and labeled blocks in FIG. 7A); finetuning may involve training the sampled neural networks 704 using their initializations yielded from the blockwise knowledge distillation using end-to-end knowledge distillation; using the initializations yielded from the blockwise knowledge distillation may reduce the resources needed to fine-tune the sampled neural networks 704 compared to training from scratch to match from scratch training accuracy; these fine-tuned neural networks 704 may also be referred to herein as distilled neural networks; these fine-tuned neural networks 704 may then be used to build an accuracy predictor 706, an example of which is illustrated in FIG. 7F such as in which blocks having shading, sizing, and labels may correspond to like shaded, sized, and labeled blocks in FIG. 7C, and may indicate the loss function, such as MSE, for which the blocks have been trained to reduce; allow for mixing different cell-types, activation functions, quantization levels and/or attention mechanisms, while still being able to model their accuracy reliably; allow for building an accuracy model for a single search space with different attention mechanisms, activation functions, cell-types, channels, quantization settings; allow for adding quantized blocks into the search space, which may build an accuracy model for quantized networks directly; the accuracy predictor 706 may be understood as a coarse sensitivity model that may indicate which blocks require complex implementations in order to build neural networks with high accuracy; ¶¶ [0088]-[0097] with FIG. 8: building multiple neural network architectures for various use-cases; in block 802, train any number and combination of search blocks from the search space using knowledge distillation with a trained reference neural network; in block 804, extract quality features and initialize weights for each search block through knowledge distillation; in block 806, store the quality features and initialize weights for each trained search block and/or neural network architecture; in block 808, select a sub-set of neural networks from the search space as targets for building an accuracy model; in block 810, train fine-tune the sub-set of neural networks as targets for building an accuracy model using knowledge distillation; in block 812, extract targets for building an accuracy model using the fine-tuned neural networks; in block 814, generate an accuracy predictor for neural networks from the search space using the quality features and targets; ¶ [0103] with FIG. 11: any neural network 1102 from the search space (e.g., non-shaded blocks 1, 2, 3, ..., Nin neural networks 702 in FIG. 7A) may be fine-tuned using a trained neural network 1104 having blocks trained using knowledge distillation such as shaded blocks 1, 2, 3, ... , N, for which blocks having shading, sizing, and labels may correspond to like shaded, sized, and labeled blocks in FIG. 7F, neural networks 704 in FIG. 7C, or accuracy predictor 706 in FIG. 7F; the neural network 1102 from the search space and initialized with weights from the knowledge distillation of the search blocks may be fine-tuned using by using end-to-end knowledge distillation using the trained reference neural network 1104; ¶¶ [0104]-[0107] with FIG. 12: fine-tuning a neural network architecture from a search-space for distilling neural network architectures; in block 1202, select a neural network from the search space; the neural network from the search space may include any number and combination of search blocks; select the neural network based on the selection of search blocks suited to the criteria; in block 1204, initialize the neural network from the search space using weights from the training of the search blocks using knowledge distillation; in block 1206, fine-tune the neural network having search blocks initialized with the weight from the knowledge distillation by using knowledge distillation with a trained reference neural network; ¶ [0117] with FIG. 16: applying a distilled neural network architecture from a search-space for refining and compressing a trained neural network; a trained neural network 1600 may be modified to use trained search blocks (e.g., shaded blocks 1, 2, 3, 5, 7, 8, 9 in neural network 1602) from the search space to refine and compress the trained neural network 1600; refining the trained neural network may retain whole the structure of the trained neural network, replacing the blocks (e.g., non-shaded blocks 1, 2, 3, 5, 7, 8, 9 in neural network 1600) within the trained neural network with trained blocks (e.g., shaded blocks 1, 2, 3, 5, 7, 8, 9 in neural network 1602) from the search space, which may be more efficient on hardware; to accomplish this, the trained neural network 1600 may be used as the reference neural network in the search space, and as the reference neural network for training the search blocks with knowledge distillation; refining the neural network 1600 may involve designing a distilled neural network 1602 as a scenario specific version of the neural network; the parameters for the search blocks in the search space and/or the search for search blocks to generate a distilled neural network 1602 may be configured for a specific scenario, such as use-case, hardware configuration, software configuration, etc.; as such, distilling a neural network 1600 based on the reference neural network may result in a scenario specific version of the reference neural network, which may be a distilled neural network 1602 configured to perform better for the scenario; refining the neural network 1600 may involve designing a distilled neural network 1602 as a compressed version of the neural network 1600; the parameters for the search blocks in the search space may be set to be smaller than the corresponding blocks in the reference neural network; as such, distilling a neural network 1600 based on the reference neural network may result in a compressed version of the reference neural network 1600, which may be a smaller and more efficient distilled neural network 1602).
Claim 26
MOONS discloses all the elements as stated in Claim 25 and further discloses wherein the compressed machine-learned model comprises one or more uncompressed model portions of the uncompressed machine-learned model (MOONS, ¶ [0117] with FIG. 16: applying a distilled neural network architecture from a search-space for refining and compressing a trained neural network; a trained neural network 1600 may be modified to use trained search blocks (e.g., shaded blocks 1, 2, 3, 5, 7, 8, 9 in neural network 1602) from the search space to refine and compress the trained neural network 1600; refining the trained neural network may retain whole the structure of the trained neural network, replacing the blocks (e.g., non-shaded blocks 1, 2, 3, 5, 7, 8, 9 in neural network 1600) within the trained neural network with trained blocks (e.g., shaded blocks 1, 2, 3, 5, 7, 8, 9 in neural network 1602) from the search space, which may be more efficient on hardware; to accomplish this, the trained neural network 1600 may be used as the reference neural network in the search space, and as the reference neural network for training the search blocks with knowledge distillation; refining the neural network 1600 may involve designing a distilled neural network 1602 as a scenario specific version of the neural network; the parameters for the search blocks in the search space and/or the search for search blocks to generate a distilled neural network 1602 may be configured for a specific scenario, such as use-case, hardware configuration, software configuration, etc.; as such, distilling a neural network 1600 based on the reference neural network may result in a scenario specific version of the reference neural network, which may be a distilled neural network 1602 configured to perform better for the scenario; refining the neural network 1600 may involve designing a distilled neural network 1602 as a compressed version of the neural network 1600; the parameters for the search blocks in the search space may be set to be smaller than the corresponding blocks in the reference neural network; as such, distilling a neural network 1600 based on the reference neural network may result in a compressed version of the reference neural network 1600, which may be a smaller and more efficient distilled neural network 1602).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Swaminathan et al. (US 11,809,992 B1, filed on 03/31/2020) discloses in Col. 2, line 15 – Col. 3, line 3 that (1) applying compression profiles across similar neural network architectures; (2) network compression may be performed to reduce the size of a trained network, which may be applied to minimize a change in the accuracy of results provided by the neural network; (3) as very large neural networks can become cost prohibitive to implement in systems with various processing limitations (e.g., memory, bandwidth, number of nodes, latency, processor capacity, etc.), techniques to provide compressed neural networks (e.g., layer, channel, or node pruning) can expand the possible implementations for a neural network (e.g., across different systems or devices with various resource limitations to implement the neural network); e.g., compression may be implemented to lower the memory or power requirements for a neural network, or may be compressed to reduce latency be providing a faster result (e.g., a faster inference); (4) determining how to compress a neural network is also not without cost, and thus, to apply compression profiles across similar network architectures may be implemented to decrease the cost (e.g., reduce time, making the compression faster) to apply compression; e.g., channel pruning is one type of network compression that may be implemented, where the number of channels in each layer of a neural network is reduced; (5) a channel pruning algorithm may determine a number of channels to prune in each layer and which channels to prune in each layer; (6) instead of implementing iterative techniques that are time and resource intensive to determine the number and which channels to prune, fast network compression can be achieved from the application of predefined compression profiles that are specific to a network architecture (or similar network architectures) to quickly make compression decisions, such as how much to prune in each layer of a neural network; (7) since the compression profiles may be applicable to any trained network of the same architecture, using these profiles can provide a high accuracy for the corresponding compression without utilize expensive and iterative analysis and instead provide a single-pass technique to compress a neural network; (8) moreover, using compression profiles in this way can reduce time taken for compression, as analysis may not be needed on the trained neural network; (9) randomization can be used to select the features to remove, such as random pruning of channels from a neural network, as random pruning may works as well as any metric-based pruning; (10) moreover, since random pruning can be applied to the network without the analysis of the network features (as noted above) such as the weights, gradients, etc., and can be applied in a single pass without iteration, further improvements to the speed of neural network compression can be accomplished; (11) the deployment of compression techniques may also be simplified as a compression system may not need to compute complicated metrics (such as gradient, etc.) from the neural network; and (12) other techniques can be utilized to select which features to remove, such as max-metric, learned, or online techniques. Swaminathan further discloses in Col. 3, line 29 – Col. 4, line 15 with FIG. 1 that (1) compression system 110 may implement compression profile selection 120 in order to identify a compression profile to apply to trained neural network 150; e.g., compression profile mappings 122 may be maintained (e.g., as a database or other data structure or data store) to map network architectures, like architectures 124a, 124b, and 124c, to corresponding one or more compression profiles, such as compression profile(s) 126a, 126b, and 126c; (2) a compression profile may be information to determine the number and/or location of features to remove from a neural network architecture to compress that neural network architecture; e.g., a compression profile may be produced from a compression policy that is trained for compressing the same or similar neural network architectures; (3) compression profile selection 120 can rely upon the generation of new or updated compression profiles from new or provided compression policies (e.g., newly trained/generated or manually specified) in order to improve the performance of compression profile selection over time (4) in this way, compression system 110 may dynamically improve the performance of compression applied to received, trained neural networks, like trained neural network 150; (5) compression profile application 130 may apply a selected compression profile to trained neural network 150 using the profile information; e.g., a pruning profile may describe how much to prune each layer of a CNN and which features (e.g., nodes) within the layer to prune; (6) compression profiles can be generated from various compression policies and thus various compression policies that specify techniques, such as network or layer compression, heuristics, learned, or online compression, random, max-metric, learned or online feature selection, among others, may be determinative of how much and which features a selected compression profile applies; (7) the compressed version of the trained neural network can then be provided to compressed version tuning 140; (8) compression system 110 may implement compressed version tuning 140 in order to apply a tuning data set 142 to retrain the compressed version of the neural network; (9) tuning data set 142 may be provided by requesting client or other application that submitted the trained neural network 150; (10) tuning data set 142 may be provided from another source than the source of the trained neural network 150; and (11) as illustrated in FIG. 1, compressed version tuning 140 may provide a trained neural network 160 that includes compressed architecture 162 and tuned weights 164 according to the selected compression profile and tuning data set 142. Swaminathan further discloses in Col. 7, line 11 – Col. 8, line 49 with FIGS. 2-3 that (1) a request to prune a neural network artifact 350 may be received a machine learning service 210 via interface 211, wherein the request may specify an identifier or type of artifact as well as various other features of the compression to be applied; (2) the neural network artifact 360 may be obtained by model compression 213, and architecture extraction 310 may implemented to identify or otherwise obtain the architecture of the neural network (e.g., parsing the file or other encoding of the artifact); (3) external 30 neural network architecture sources may be searched or previously searched and cataloged for architecture extraction to compare for identify architecture 362; (4) architecture extraction 310 may then provide the architecture 362 to pruning profile selection 320; (5) pruning profile selection 320 may perform a lookup 374 on pruning profile index 324 ; e.g., index values may be determined from an architecture 362 (or label, indicator, type, version, or other identifier of the architecture 362) to check to see if a match is located; (6) if a match is found, then the corresponding pruning profile may be provide; (7) if a match is not found, then pruning profile generator 322 may provide a new pruning profile 372, wherein the new pruning profile 372 may also be provided to pruning profile index 324 for subsequently received requests for the same (or similar) architecture; e.g., a pruning policy may indicate, among other features for generating a pruning profile a technique to determine where and how much to prune for networks and/or layers using heuristics, learned, or online techniques, wherein a pruning policy can specify or determine how much to prune can be taken at either a network-level or at a layer-level; (8) since making layer-level decisions may implicitly make a network-level decision, often network-level decisions may be made first and propagated to layer-level; (9) several types of policies can be used to decide how much to prune by pruning profile generator 322, which can range from heuristics such as a uniform pruning policy across all layers, to learning a policy either online or offline, either using an reinforcement learning (RL) or gradient-based methods, on observing some properties of the network or layer, depending on the level such as the correlation analysis of layers; (10) for heuristic approaches, the heuristics may be designed in accordance to constraints based on the target of the compressed model; (11) pruning policies applied by pruning profile generator 322 to generate pruning profiles may also determine which features, such as which channel, to prune in the pruning profile; (12) a direct policy for which channels to prune could implicitly decide how much to prune as well; (13) the importance of channels can be weighted before deciding which to channel prune; (14) some pruning policies may provide profiles that completely freeze pruned layers or remove them altogether, while other pruning policies may allow soft pruning at prune-time allowing parameters to recover back if they are needed; typical options for which channels to prune contingent on a previously made decision for how much to prune involve random decisions such as sampling from a Bernoulli distribution or using metrics of the layer's features (e.g., tensors), such as pruning the channels with the least expected L1 of weights; (15) pruning profile selection 320 may provide the pruning profile 364 that is selected to pruner 330; (16) pruner 330 may obtain the neural network artifact 360 and apply the pruning technique using the pruning profile; e.g., pruner 330 may alter the file, structure, or other data object storing neural network artifact to remove or delete the specified features in the network, layer, channel or other portion of the neural network artifact; (17) pruner 330 may then provide the pruned artifact 366 for subsequent processing; and (18) pruner 330 may also indicate to compression job management 340 to update the compression job state to tuning 354.
Ahme et al., ("Compact CNN Structure Learning by Knowledge Distillation ", arXiv:2104.09191v1, Apr. 19, 2021) discloses in Abstract and Section I with FIG. 1 of Pages 1-2 that (1) propose a framework that leverages knowledge distillation along with customizable block-wise optimization to learn a lightweight CNN structure while preserving better control over the compression-performance tradeoff; (2) considering specific resource constraints, e.g., floating-point operations per inference (FLOPs) or model-parameters, our method results in a state of the art network compression while being capable of achieving better inference accuracy; (3) propose a powerful and adaptive method for learning an optimal network structure for the task of interest; (4) our approach advances the spirit of recently proposed method MorphNet [8] which has the advantage of being fast, scalable and adaptive to specific resource constraints (e.g., FLOPs or model-parameters);(5) most importantly, it learns network structure during training: (6) however, the current optimization technique has an intrinsically biased concentration that either pushes the optimizer to focus on high-resolution layers (towards network input) or focus more on low-resolution layers (towards network output) when optimized for FLOPs or model-parameters, respectively; (7) employ Resource-aware Optimization augmented with the Privileged Information (PI) technique in a student-teacher scheme (see Figure 1); (8) the proposed resource-aware optimization breaks down the seed network in smaller instances which curtails task complexity to learn better end-to-end network structure; (9) eventually, it enables customized optimization of each stage of the network with specific budget constraints; and (10) imposing control over model performance during optimization considering the teacher network performance as a target, which facilitates the optimizer to maintain high model performance while learning the lightweight network structure. Ahme further discloses in Section III of Pages 3-4 that (1) propose a framework that learns an optimal CNN structure to efficiently target the task of interest, considering allowed resource constraints e.g., FLOPs and model-parameters while preserving high model performance; (2) the model compression method of MorphNet [8] relies on a regularizer R, which induces sparsity in activations by putting greater cost C on neurons contributing to either FLOPs or the model-parameters; (3) the network sparsity is measured on the basis of batch normalization scaling factor associated with each neuron; i.e., if lies below than the user-defined threshold, the corresponding neuron is considered as dead and can be discarded (since its scale is negligible); (4) both the FLOPs and model-parameters are influenced by the particular layer associated with matrix multiplications - i.e. convolutions; (5) this makes sense, as the lower layers of the neural network are applied to a high-resolution image, and thus consume a large number of the total FLOPs, and whereas, the upper layers typically comprises of larger number of channels and thus contain abundant weight matrices; (6) we can define separate cost functions as shown in Eq (1) and Eq (2), where C is a function of the model-parameters θ and the hyperparameter α regulates the resource optimization intensity; (7) to overcome these challenges, our method employs Resource-aware Optimization augmented with the Privileged Information (PI) technique in a student-teacher scheme (see Figure 1); (8) the resource-aware optimization breaks the complex task of learning the entire CNN structure into comparatively simpler sub-tasks; (9) accordingly, with the reduced complexity (to find the structure of sub-network only), it enables customized optimization of each stage of the network with specific budget constraints; (10) eventually, our approach discovers a global network structure, lighter than the original end-to-end solution; (11) in addition to that, the privileged information framework , augments our method’s capability by imposing control over model performance during optimization considering the teacher network performance as a target, which facilitates the optimizer to maintain high model performance while learning the optimal network structure; (12) note that our method does not apply network expansion at all, and the student network to be compressed utilizes PI extracted almost for free from the uncompressed network itself; (13) in this way, the impact on performance is also accounted along with the existing sparsity measure that helps the optimizer to remove only the least significant neurons from the network; (14) the structure of the student network is optimized to meet the required resource budget while taking advantage of the soft predictions (from the teacher) along with the ground-truth labels. This forces the student network to keep mimicking baseline predictions during optimization; (15) the training is accomplished by minimizing the cross-entropy loss shown in Eq. (4); (16) during the inference of yi given xi, i = 1, …, N, the student network leverages privileged information zi about the sample (xi, yi), which is derived from the teacher model's prediction shown in Eq. (5); (17) thus, the student network is trained according to the optimization problem shown in Eq. (6), wherein the parameter T regulates the amount of smoothness applied to logits; (18) this not only reveals commonalities and differences between classes to be discriminated but also exploits the true potential of the soft labels; (19) Fs is the student function hypothesis space and the parameter λ [Symbol font/0xCE] [0, 1] is the imitation factor, controlling the student to mimic the teacher vs. to predict the ground-truth label; (20) therefore, the proposed approach is modeled by incorporating Eq. (6) into MorphNet’s minimization equation (in Eq. (3)), and thus we get the optimization problem shown in Eq. (7), where the standard cross-entropy loss incorporates both, the privileged information and the MorphNet optimizations; (21) propose a resource-aware optimization scheme to optimize different layers of the network using suitable optimizer which leads to Eq. (8) modification to Eq. (7); and (22) specifically, we propose a configuration in which the first half of the network is optimized for FLOPs and the second half is optimized for model-parameters.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to HWEI-MIN LU whose telephone number is (313)446-4913. The examiner can normally be reached Mon - Fri: 9:00 AM - 6:00 PM EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Mariela D. Reyes can be reached at (571) 270-1006. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/HWEI-MIN LU/Primary Examiner, Art Unit 2142