Prosecution Insights
Last updated: October 02, 2026
Application No. 17/751,089

Automated Selection of Neural Architecture Using a Smoothed Super-Net

Final Rejection §103
Filed
May 23, 2022
Examiner
RAMESH, TIRUMALE K
Art Unit
2121
Tech Center
2100 — Computer Architecture & Software
Assignee
ARM Limited
OA Round
4 (Final)
26%
Grant Probability
At Risk
5-6
OA Rounds
4m
Est. Remaining
50%
With Interview

Examiner Intelligence

Grants only 26% of cases
26%
Career Allowance Rate
13 granted / 49 resolved
-28.5% vs TC avg
Strong +24% interview lift
Without
With
+23.7%
Interview Lift
resolved cases with interview
Typical timeline
4y 9m
Avg Prosecution
19 currently pending
Career history
85
Total Applications
across all art units

Statute-Specific Performance

§101
26.8%
-13.2% vs TC avg
§103
63.9%
+23.9% vs TC avg
§102
4.2%
-35.8% vs TC avg
§112
4.6%
-35.4% vs TC avg
Black line = Tech Center average estimate • Based on career data from 49 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Amendment (5/11/2026) Applicant’s arguments with respect to claims 1, 5, 9 and 15 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument. In regard to 103 rejections - The applicant has amended the claims 1, 5, 9 and 15 as stated Page 9. The applicant states that the smoothness is a major factor in the invention and is considered as result of different architecture distributions with respect to different sub-networks with shared weights. On Page 10 , recites the amended claim 1 as supported in specification [0052-[0054]. On Page 11, the applicant argues that none of the references “Roth”, “Peng” and “Mok” teaches the second loss based on smoothness for adjusting weights in weight sharing NAS. The applicant argues that reference Peng does not teach the smoothing of updates parameters. Further, the applicant argues on Page 12 that “Peng” teaches “uncertainty” and parameter shifts” and further argues that smoothing is function of the layer itself and as a result, the applicant has specifically amended the claim to cite “ measure of smoothness of a layer solely from the weights in the layer”. Further on Page 12, the applicant has argued that “Mok” teaches the loss as a function of network and training input and does not teach the loss as function of the architectural parameters. With regard to amended claim 5, the applicant argues on Page 14 that Dong reference had to be amended to clarify the second loss over the sample of sub-networks which is not taught in the “Dong” reference. Examiner’s Response With respect to Claim 5, the examiner interprets that the second loss is not a combined loss and is the loss for the second output branch and updates all sub-networks’ parameters in back-propagation because the total loss is the sum of all individual losses for each sub-network. This approach is common in multi-output or multi-task models, where each sub-network handles a different prediction task. Without conceding the arguments from the applicant, the examiner submits that new references “Cummings”, and “Badran” strongly teaches all the amendments of the specifically weight sharing and computing the second loss in claims 1, 9 and 15. The amended dependent claim 5 is taught by reference “Dong”. In CONCLUSION, the examiner rejects the claims 1-7, 9-13, 15-19 and 21-22 under 103 and MOVES the application as FINAL REJECITION. Claim Rejections – 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-2, 9-10 and 15-16 are rejected under 35 U.S.C. 103 as being unpatentable over Daniel Cummings et.al. (hereinafter Cummings ) US 2022/0036123 A1, In view F. Badran et.al. (hereinafter Badran) Neural Network Smoothing in Correlated Time Series Context, Neural Networks, Vol. 10, No. 8, pp. 1445–1453, 1997. Claim 1: (Currently Amended) Cummings discloses: - A computer-implemented method of automated selection of neural architecture, the method comprising: [0100]: In some implementations, the processor circuitry 552 and/or acceleration circuitry 564 may include HW elements specifically tailored for machine learning functionality, such as for operating performing ANN operations, [Abstract]: The present disclosure is related to machine learning model swap (MLMS) framework for that selects and interchanges machine learning (ML) models [Abstract]: The MLMS framework includes an ML model search strategy that can flexibly adapt ML models for a wide variety of compute system and/or environmental changes [0006]: FIG. 1 depicts an overview of a machine learning model swapping (MLMS) system according to various embodiments. [0008]: The present disclosure is related to techniques for optimizing artificial intelligence (AI) and/or machine learning (ML) models to reduce resource consumption while improving AI/ML model performance. In particular, the present disclosure provides a ML architecture search (MLAS) framework that involves generalized and/or hardware (HW)-aware ML architectures. [0019:] For purposes of the present disclosure, the term “ML architecture” may refer to a particular ML model having a particular set of ML parameters and/or such an ML model configured to be operated on a particular HW platform [0017]: FIG. 1 shows the main components and interactions of an ML model scaling (MLMS) system 100. MLMS system 100 provides a holistic system that selects and interchanges ML models in an energy and communication efficient way while at the same time adapting the models to real-time changes in system constraints, conditions, and/or configurations. - accessing training data for a chosen task, the training data including a plurality of training inputs and corresponding training outputs in the training outputs in the training data for the chosen task [0022]: The supernet 130 may be an ML model configured and/or trained for a particular AI/ML task from which the MLMS system 100 is to discover or generate a smaller and/or derivative ML model (referred to herein as a “sub-network” or “subnet”). [0022]: [0022]: the supernet 130 contains one or more smaller subnets (e.g., the subnets 135 in the subnet pool 147 in FIG. 1) that, when trained in isolation, can match and/or offer trade-offs in various objectives and/or performance metrics such as accuracy and/or latency of the original supernet 130 when trained for the same number of iterations or epochs. [0015]: The MLMS system discussed herein allows prediction/inference tasks to be performed on a HW platform under dynamically varying environmental conditions (e.g., temperature, battery charge) and HW configurations (e.g. utilized cores, power modes, etc.) while maintaining the same or similar level of performance of prediction/inference results (e.g., accuracy, latency, etc.) as if the HW platform were operated in a more stable environment and/or with static HW configurations - configuring, by a supervised learning controller, a super-net to implement a sample of sub-networks of the super-net, the super-net including a plurality of nodes coupled by operational blocks, an operational block including two or more neural networks [0014]: The MLMS system implements a unique and robust ML model search and swap strategy that can flexibly adapt ML models for a wide variety of detected compute system operational changes and/or environmental changes. [0021]: in alternative implementations, the MLMS system 100 may be operated by the same compute node (e.g., client device, edge compute node, network access node, cloud service, drone, network appliance, etc.), [0100]: In some implementations, the processor circuitry 552 and/or acceleration circuitry 564 may include HW elements specifically tailored for machine learning functionality, such as for operating performing ANN operations, [0033]: In some implementations, the HW platform may be the client device 101. In some implementations, the HW platform may be some other desired compute node or device on which the user wishes to deploy an ML model or wishes to perform some AI/ML task”, [0027]: The AI/ML tasks may describe a desired problem to be solved and the AI/ML domain may describe a desired goal to be achieved. Examples of ML tasks include clustering, classification, regression, anomaly detection, data cleaning, automated ML (autoML), association rules learning, reinforcement learning, structured prediction, feature engineering, feature learning, online learning, supervised learning, semi-supervised learning (SSL), unsupervised learning, machine learned ranking (MLR), grammar induction, and/or the like. [BRI: the MLMS (Machine Learning Model Swap) system implements a unique and robust ML model search and swap strategy that can flexibly adapt ML models to a wide variety of detected compute system and/or environmental changes, treating these changes as “operational blocks” for model adaptation] - each having a plurality of network weights and associated with an architecture parameter[[s]], [0019]: model parameters are parameters derived via training, whereas hyperparameters are parameters whose values are used to control aspects of the learning process and usually have to be set before running an ML model (e.g., weights, etc.). Additionally, for purposes of the present disclosure, hyperparameters may be classified as architectural hyperparameters or training hyperparameters. Architectural hyperparameters are hyperparameters that are related to architectural aspects of an ML model - where different neural networks in the same operational block share at least some of the same network weights, [0014] : The present disclosure provides an MLMS system that selects and interchanges ML models (e.g., subnets) in an energy and communication efficient way while adapting the ML models (e.g., subnets) to real time (or near-real time) changes in system (HW platform) constraints. The MLMS system discussed herein is agnostic to particular approach used for ML model search (e.g., NAS) and for ML model (e.g., supernet) weight sharing. The MLMS system implements a unique and robust ML model search and swap strategy that can flexibly adapt ML models for a wide variety of detected compute system operational changes and/or environmental changes. - and where the architecture parameters represent the importance of different architecture choices at various locations inside the super-net [0019]: For purposes of the present disclosure, the term “ML architecture” may refer to a particular ML model having a particular set of ML parameters and/or such an ML model configured to be operated on a particular HW platform. Here, the set of ML parameters may refer to “model parameters” (also referred to simply as “parameters”) and/or “hyperparameters.” Model parameters are parameters derived via training, whereas hyperparameters are parameters whose values are used to control aspects of the learning process and usually have to be set before running an ML model (e.g., weights, etc.). Additionally, for purposes of the present disclosure, hyperparameters may be classified as architectural hyperparameters or training hyperparameters. Architectural hyperparameters are hyperparameters that are related to architectural aspects of an ML model such as, for example, the number of layers in a DNN, specific layer types in a DNN (e.g., convolutional layers, multilayer perception (MLP) layers, etc.), number of output channels, kernel size, and/or the like. [0178]: The term “attention” in the context of machine learning and/or neural networks, at least in some embodiments refers to a technique that mimics cognitive attention, which enhances important parts of a dataset where the important parts of the dataset may be determined using training data by gradient descent - generating, by the super-net, sub-network outputs responsive to the training inputs for the sample of sub-networks, where a sample sub-network includes a neural network selected from each operational block; [0014] : The present disclosure provides an MLMS system that selects and interchanges ML models (e.g., subnets) in an energy and communication efficient way while adapting the ML models (e.g., subnets) to real time (or near-real time) changes in system (HW platform) constraints [BRI: A MLMS system is an operational block] [0058]: FIG. 3 shows an example of the efficiency-aware subnet model selection after the subnet search is performed, according to various embodiments. The supernet search results 301 shows all discovered subnets 335 found during a subnet search, which includes the existing supernet 330 and a set of candidate replacement subnets 340 (individual subnets in the set of candidate replacement subnets 340 may be referred to herein as “replacement subnets 340”, “candidate subnets 340”, or the like). [BRI: candidate subnets represents sample of subents] - training, by the supervised learning controller, network weights and architecture parameters of a super-net including a plurality of sub-networks, the training including: [0027]: The AI/ML tasks may describe a desired problem to be solved and the AI/ML domain may describe a desired goal to be achieved. Examples of ML tasks include clustering, classification, regression, anomaly detection, data cleaning, automated ML (autoML), association rules learning, reinforcement learning, structured prediction, feature engineering, feature learning, online learning, supervised learning [0100]: In some implementations, the processor circuitry 552 and/or acceleration circuitry 564 may include HW elements specifically tailored for machine learning functionality, such as for operating performing ANN operations, [0178]: The term “multi-head attention” at least in some embodiments refers to an attention technique that combines several different attention mechanisms to direct the overall attention of a network or subnetwork. [0179]: the goal is to break down complicated tasks into smaller areas of attention that are processed sequentially. [0179]: The term “attention network” at least in some embodiments refers to an artificial neural networks used for attention in machine learning. [0019]: model parameters are parameters derived via training, whereas hyperparameters are parameters whose values are used to control aspects of the learning process and usually have to be set before running an ML model (e.g., weights, etc.). Additionally, for purposes of the present disclosure, hyperparameters may be classified as architectural hyperparameters or training hyperparameters. Architectural hyperparameters are hyperparameters that are related to architectural aspects of an ML model [0022: In various implementations, the ML config includes a reference ML model 130 (referred to herein as a “super network 130” or “supernet 130 [0022]: The supernet 130 may be an ML model configured and/or trained for a particular AI/ML task from which the MLMS system 100 is to discover or generate a smaller and/or derivative ML model (referred to herein as a “sub-network” or “subnet”). [0022]: This supernet 130 may include parameters and/or weights that do not significantly contribute to the prediction and/or inference determination, and these parameters and/or weights contribute to the supernet's overall computational complexity and density. Therefore, the supernet 130 contains one or more smaller subnets (e.g., the subnets 135 in the subnet pool 147 in FIG. 1) that, when trained in isolation, can match and/or offer trade-offs in various objectives and/or performance metrics such as accuracy and/or latency of the original supernet 130 when trained for the same number of iterations or epochs. [0022]: In the context of the present disclosure, the term “elastic” refers to the ability of an elastic supernet 130 to add and/or remove different ML parameters and/or subnets 135 without restarting the running framework and/or the running ML model. [BRI: Perhaps known to a POSITA that elastic supernet is a type of over-parameterized neural network that contains many candidate subnetworks within a single, larger model, enabling neural architecture search (NAS) to explore a wide range of architectures efficiently ] - accessing the training data for the chosen task. [0002]: Machine learning (ML) is the study of computer algorithms that improve automatically through experience and by the use of data. [0002]: ML algorithms build models using sample data (referred to as “training data”) and/or based on past experience in order to make predictions or decisions without being explicitly programmed to do so. [0077] : Machine learning (ML) involves programming computing systems to optimize a performance criterion using example (training) data and/or past experience. [0077]: ML involves using algorithms to perform specific task(s) without using explicit instructions to perform the specific task(s) [0013]: different solutions that need to be searched for training these separate ML models. - determining, from the sub-network outputs generated by the super-net hardware, a first loss based on accumulated differences between the sub-network outputs and corresponding training outputs [0059]: In embodiments, the subnet selector 141 implements a search function and/or NAS algorithm to identify as many possible subnet candidates 340 for the new context information (e.g., new HW configuration) that precipitated the search. [0209]: The term “loss function” or “cost function” at least in some embodiments refers to an event or values of one or more variables onto a real number that represents some “cost” associated with the event. A value calculated by a loss function may be referred to as a “loss” or “error”. Additionally or alternatively, the term “loss function” or “cost function” at least in some embodiments refers to a function used to determine the error or loss between the output of an algorithm and a target value. Additionally or alternatively, the term “loss function” or “cost function” at least in some embodiments refers to a function are used in optimization problems with the goal of minimizing a loss or error. [BRI: Perhaps known to a POSITA, that cost function is typically the average (or sum) of the loss function over all samples in a dataset. It aggregates the per-sample losses to give an overall measure of model performance. Thus, in many embodiments, especially in deep learning or multi-model setups, the “loss function” at the per-sample level and the “cost function” at the dataset level both represent accumulated differences between model outputs and target values, with the cost function being the total or averaged loss across all samples] - accumulating a second loss over the sample of sub-networks, [0059]: In embodiments, the subnet selector 141 implements a search function and/or NAS algorithm to identify as many possible subnet candidates 340 for the new context information (e.g., new HW configuration) that precipitated the search. [0209]: The term “loss function” or “cost function” at least in some embodiments refers to an event or values of one or more variables onto a real number that represents some “cost” associated with the event. A value calculated by a loss function may be referred to as a “loss” or “error”. Additionally or alternatively, the term “loss function” or “cost function” at least in some embodiments refers to a function used to determine the error or loss between the output of an algorithm and a target value. Additionally or alternatively, the term “loss function” or “cost function” at least in some embodiments refers to a function are used in optimization problems with the goal of minimizing a loss or error. [0059]: Of the possible subnet candidates 340 that are found, preference is given to a subnet 340 in the set of candidates 340 that is the most related to the existing subnet 135. The determination of the most closely related subnet 340 is based on the supernet type weight/parameter sharing scheme, such as the vectorization scheme discussed previously with respect to FIG. 2, which can be used to calculate the total transfer cost for each of the candidates 340. The subnet 340 with a minimal total transfer cost (e.g., with respect to the other candidate subnets 340) is selected as a new active subnet. [BRI: The transfer cost is the second loss] - selecting a sub-network of the plurality of sub-networks for the chosen task based on the largest adjusted architecture parameters; [0017]: The MLMS system 100 achieves energy and communication efficiency by using a similarity-based subnet selection process where a subnet is selected from a pool of subnets that has the most overlap in pre-trained parameters from the existing subnet to minimize memory write operation overhead [0019]: Architectural hyperparameters are hyperparameters that are related to architectural aspects of an ML model such as, for example, the number of layers in a DNN, specific layer types in a DNN (e.g., convolutional layers, multilayer perception (MLP) layers, etc.), number of output channels, kernel size, and/or the like. [0022]: “elasticity” refers to a component in the supernet 130 and its corresponding operator can change in several dimensions. For example, in convolutional networks this would allow for the choice of different kernel sizes (e.g., elastic kernel), depths of some group of elements (e.g., elastic depth), and number of channels of the selected components (e.g., elastic width). In the context of the present disclosure, the term “elastic” refers to the ability of an elastic supernet 130 to add and/or remove different ML parameters and/or subnets 135 without restarting the running framework and/or the running ML model. [BRI: Perhaps known to a POISTA that adjusted architecture parameters refer to the continuous or discrete variables that define a candidate network’s structure, which are optimized during the search process. These parameters can include: filter(kernel), number of layers, and number of channels. - and outputting a description of the network architecture of the selected sub-network for the chosen task. [0010]: optimizing an ML model for individual HW platforms and specific performance metrics is a very time-consuming effort, which requires highly specialized knowledge. This type of optimization is typically done manually with a great deal of in-depth understanding of the HW platform since certain characteristics of the HW platform (e.g., clock speed, number of processor cores, amount of cache memory, etc.) will affect the optimization process. The optimization process is also affected by the characteristics of the input data to the ML model (e.g., batch size, image size, number of epochs/iterations, etc.). Finally, any change to the performance metrics (e.g., going from latency to power consumption), input data characteristics (e.g., increasing the batch size), [0091]: To provide the inference, the inference engine 516 uses a model 520 that controls how the DNN inference is made on the data 514 to generate the result 518. Specifically, the model 520 includes a topology of layers of a NN. The topology includes an input layer that receives the data 514, an output layer that outputs the result 518, and one or more hidden layers between the input and output layers that provide processing between the data 14 and the result 518. The topology may be stored in a suitable information object, such as an extensible markup language (XML), JavaScript Object Notation (JSON), and/or other suitable data structure, file, and/or the like. The model 520 may also include weights and/or biases for results for any of the layers while processing the data 514 in the infence using the DNN. - and adjusting network weights and architecture parameters of the super-net to reduce a combination of the first and second losses, [0214]: The term “objective function” at least in some embodiments refers to a function to be maximized or minimized for a specific optimization problem. In some cases, an objective function is defined by its decision variables and an objective. The objective is the value, target, or goal to be optimized, such as maximizing profit or minimizing usage of a particular resource. The specific objective function chosen depends on the specific problem to be solved and the objectives to be optimized. [0209]: the term “loss function” or “cost function” at least in some embodiments refers to a function are used in optimization problems with the goal of minimizing a loss or error. [0059]: Of the possible subnet candidates 340 that are found, preference is given to a subnet 340 in the set of candidates 340 that is the most related to the existing subnet 135. The determination of the most closely related subnet 340 is based on the supernet type weight/parameter sharing scheme, such as the vectorization scheme discussed previously with respect to FIG. 2, which can be used to calculate the total transfer cost for each of the candidates 340. The subnet 340 with a minimal total transfer cost (e.g., with respect to the other candidate subnets 340) is selected as a new active subnet. [BRI: Perhaps known to a POSITA that in deep learning combing a first loss (subnetwork outputs vs. training outputs) with a second loss (e.g., sample-based subnetwork evaluation) into a single composite loss function is common in multi-objective or multi-task scenarios, such as multi-subnetwork training. This can be achieved via weighted summation. Within the context of objective function being optimized, the total transfer cost is combination of first and second loss] - where the network weights shared between different neural networks of an operational block are trained jointly, [0014] : The present disclosure provides an MLMS system that selects and interchanges ML models (e.g., subnets) in an energy and communication efficient way while adapting the ML models (e.g., subnets) to real time (or near-real time) changes in system (HW platform) constraints. The MLMS system discussed herein is agnostic to particular approach used for ML model search (e.g., NAS) and for ML model (e.g., supernet) weight sharing. The MLMS system implements a unique and robust ML model search and swap strategy that can flexibly adapt ML models for a wide variety of detected compute system operational changes and/or environmental changes. [0020]: the MLMSI 110a includes application programming interfaces (APIs) to access the other subsystems of system 100, manages ML search parameter updates (e.g., new or updated ML config), and calls any supported ML operations library (e.g., as indicated by the ML config). [0209]: the term “loss function” or “cost function” at least in some embodiments refers to a function are used in optimization problems with the goal of minimizing a loss or error. [0059]: Of the possible subnet candidates 340 that are found, preference is given to a subnet 340 in the set of candidates 340 that is the most related to the existing subnet 135. The determination of the most closely related subnet 340 is based on the supernet type weight/parameter sharing scheme, such as the vectorization scheme discussed previously with respect to FIG. 2, which can be used to calculate the total transfer cost for each of the candidates 340. The subnet 340 with a minimal total transfer cost (e.g., with respect to the other candidate subnets 340) is selected as a new active subnet. [0053]: a subnet 135 that has the most overlap in pre-trained parameters from the existing subnet 135 for a target device (e.g., compute node 550). These mechanisms help the target device adapt to a new system states and/or HW conditions/constraints, but also limit the data-transfer resource consumption when writing the new subnet 135 to memory. [BRI: Within the context of objective function being optimized, and the the total transfer cost is combination of first and second loss as a result of training all components (subnet) together is indeed a form of joint training] Cummings do not explicitly disclose: - for layers of sub-networks in a sample of sub-networks, determining a measure of a smoothness of a layer solely from network weights in the layer, the measure of smoothness related to the maximum difference between outputs from the layer relative to any possible difference between inputs to the layer, a higher maximum change in output indicating a lower smoothness; - where the second loss penalizes sub- networks with lower smoothness and improves stability of the adjustment; However, Badran discloses: - for layers of sub-networks in a sample of sub-networks, determining a measure of a smoothness of a layer solely from network weights in the layer, the measure of smoothness related to the maximum difference between outputs from the layer relative to any possible difference between inputs to the layer, [2, Page 1446]: The neural network smoother uses a basic MLP architecture (R) with an input layer supporting the temporal window, l layers of non-linear hidden neurons and one linear output neuron. Starting from this basic architecture (R), a family of multi-layer perceptrons R λ is generated. For each λ ,   the synaptic weights vector is chosen using the learning procedure described below. [2, 1446]: The input layer uses a temporal window on the time series, which is the vector of observations ( y i - d ' , … y i + d ), d’ ≥ d The two time lags d and d’ can be equal or different. They can be both positive, in which case, d’ represents the amount of information in the past used by the temporal window; d stands for the future and y i   is within the window. However, d may be negative: for example, if d= -1 and y i is not within the window, smoothing at time i uses only past information, and it can be considered as a predictor of the smoothed value at time i. [2, Page 1446]: A classical smoothing method is to choose from the family of twice continuously differentiable functions S, such that PNG media_image1.png 28 107 media_image1.png Greyscale exists, the function S which minimizes the measure of roughness PNG media_image2.png 56 433 media_image2.png Greyscale The weighting parameter λ   controls the amount of smoothing, so that increasing λ   leads to smoother estimate. [BRI: The increase λ to minimizing measure of roughness. [3, Page 1447]: The MASE can be easily expressed as PNG media_image3.png 166 511 media_image3.png Greyscale The first term on the right-hand side of (6) corresponds to a variance component and the second to the square of a bias component, both of them contributing to the MASE(   g λ ) The variance component is decreased by smoothing, but this increases the bias component. The problem is thus to choose a good compromise between these two components (which is done by minimizing (6) with respect to PNG media_image4.png 12 1 media_image4.png Greyscale [BRI: the bias component measures the average difference between the model’s predictions and the best possible model for the problem, averaged over all training data instances and increasing it underfits the data and affects smoothness as represented by equation (6)] - a higher maximum change in output indicating a lower smoothness; [3, Page 1447]: The MASE can be easily expressed as PNG media_image3.png 166 511 media_image3.png Greyscale The first term on the right-hand side of (6) corresponds to a variance component and the second to the square of a bias component, both of them contributing to the MASE(   g λ ) The variance component is decreased by smoothing, but this increases the bias component. The problem is thus to choose a good compromise between these two components (which is done by minimizing (6) with respect to PNG media_image4.png 12 1 media_image4.png Greyscale [BRI: in effect, the high difference resulted from the bias component does increases MASE and thus lower the smoothness] - the second loss based, at least in part, on a sum, over layers of a sub-network, of the measures of smoothness of the layers; [2, Page 1446]: The neural network smoother uses a basic MLP architecture (R) with an input layer supporting the temporal window, l layers of non-linear hidden neurons and one linear output neuron. Starting from this basic architecture (R), a family of multi-layer perceptrons (Rl) is generated. For each l, the synaptic weights vector is chosen using the learning procedure described below. [2, 1446]: The input layer uses a temporal window on the time series, which is the vector of observations ( y i - d ' , … y i + d ), d’ ≥ d The two time lags d and d’ can be equal or different. They can be both positive, in which case, d’ represents the amount of information in the past used by the temporal window; d stands for the future and y i   is within the window. However, d may be negative: for example, if d= -1 and y i is not within the window, smoothing at time i uses only past information, and it can be considered as a predictor of the smoothed value at time i. [2, Page 1446]: A classical smoothing method is to choose from the family of twice continuously differentiable functions S, such that PNG media_image1.png 28 107 media_image1.png Greyscale exists, the function S which minimizes the measure of roughness [BRI: function S that minimizes the measure of roughness provides the measure of smoothness] - where the second loss penalizes sub- networks with lower smoothness and improves stability of the adjustment; [Abstract, Page 1445]: The NN-smoother computes the trend in the state domain and minimizes a cost function with a regularization term. The regularization term is penalized by a parameter λ which forces the learning procedure to smooth in the time domain. We define a selecting criterion in order to select the best parameter λ . [1, Page 1445]: In most cases, smoothing techniques generate families of functions g λ .     d e pending on a parameter λ , which controls the degree of smoothing. Each of these functions is a solution to the minimization of a cost function, containing a regularization term parameterized by λ . [3, Page 1447]: The MASE can be easily expressed as PNG media_image3.png 166 511 media_image3.png Greyscale The first term on the right-hand side of (6) corresponds to a variance component and the second to the square of a bias component, both of them contributing to the MASE(   g λ ). The variance component is decreased by smoothing, but this increases the bias component. The problem is thus to choose a good compromise between these two components (which is done by minimizing (6) with respect to PNG media_image4.png 12 1 media_image4.png Greyscale [BRI: in effect, the high difference resulted from the bias component does increases MASE and thus lower the smoothness and is solved by minimizing MASE by using compilation since the regularization (penalty) is parameterized by λ ] [1, Page 1445]: Neural networks can be used for smoothing: for example, an MLP, trained through a modified gradient back propagation, can implement spline smoothing in the time domain [1, Page 1446]: The neural network smoother uses a basic MLP architecture (R) with an input layer supporting the temporal window, l layers of non-linear hidden neurons and one linear output neuron. Starting from this basic architecture (R), a family of multi-layer perceptrons (Rl) is generated. For each λ , the synaptic weights vector is chosen [BRI: in the time domain, spline smoothing with a large smoothing parameter λ does represent small changes in output with respect to input, because it penalizes curvature and enforces a smooth, slowly varying curve. In fact, this is exactly what Lipschitz constant provides (small changes to output- see spec [0022]) that minimize the loss jointly for all sub-networks] [2, Page 1447]:] We consider the measure of roughness PNG media_image5.png 85 480 media_image5.png Greyscale [2, Page 1447]: The learning phase consists in minimizing the discrete approximation given in (4) by the back propagation algorithm. [4, Page 1450]: NN performs a smoothing in the state space using a cost function C(l) (4) which acts in the time domain. So, the NN smoother combines two representations: the state space domain and the time domain. The NN network directly minimizes (4) which is the discrete form of (2) and gives the optimal smoothing directly for each discrete time unit i [BRI: In the state-space domain: the internal dynamics of the system, modeled as a set of state equations (continuous or discrete) and in the time-domain data: the actual measured inputs, outputs, and (often) states over time, used to estimate the network weights and biases. Several sub-networks are part of the same state-space model. The same set of weights (shared) is updated during training to ensure consistency across the block and the joint training leverages all available time-domain data from the block to optimize shared parameter. The NN directly minimize the cost function(loss function) given by (4)] [2, Page 1447]: The input layer has d þ d þ 3 units (for window [yi¹1¹d,…,yiþ1þd]), l hidden layers (l 0) and 3 output units. The first hidden layer has three clusters of n1 units each. The first cluster of hidden units is connected to window {i ¹ 1 ¹ d,…,i ¹ 1 þ d}, the second to the window shifted by one time unit, and the third to the window shifted by one more unit; these clusters share the same weight vector. The use of time-delay inputs convert the time series into a static mapping problem. The only constraint is that each hidden layer has three clusters of units connected only to the corresponding cluster in the previous layer. In each layer, clusters share the same weight vector. It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Cummings, and Badran. Cummings teaches a super-net including a plurality of sub-networks, first loss over a sample of sub-networks, and second loss and adjusting network weights to reduce the accumulated loss. Badran teaches relating the measure of smoothness to the maximum difference between outputs from the layer relative to any possible difference between inputs to the layer. One of ordinary skill would be motivated to combine Cummings, and Badran that choose best smoothing under different hypothesis and minimum error ( Badran [5, Page 1450]). Claim 2: (Previously Presented) Cummings discloses: - training network weights of the selected sub-network using the training data; [0022]: supernet 130 may include parameters and/or weights that do not significantly contribute to the prediction and/or inference determination, and these parameters and/or weights contribute to the supernet's overall computational complexity and density. Therefore, the supernet 130 contains one or more smaller subnets (e.g., the subnets 135 in the subnet pool 147 in FIG. 1) that, when trained in isolation, can match and/or offer trade-offs in various objectives and/or performance metrics such as accuracy and/or latency of the original supernet 130 when trained for the same number of iterations or epochs. - storing the trained network weights of the selected sub-network [0136]: one or more ML models of the set of ML models has fewer parameters than the reference ML model; and storing the set of ML models, wherein the currently deployed ML model and the replacement ML model are among the stored set of ML models. [0204]: The terms “instance-based learning” or “memory-based learning” in the context of ML at least in some embodiments refers to a family of learning algorithms that, instead of performing explicit generalization, compares new problem instances with instances seen in training, which have been stored in memory Claim 9: (Currently Amended) Cummings discloses: - A system for automated selection of neural architecture, the method comprising: a supervised learning controller [0100]: In some implementations, the processor circuitry 552 and/or acceleration circuitry 564 may include HW elements specifically tailored for machine learning functionality, such as for operating performing ANN operations, [Abstract]: The present disclosure is related to machine learning model swap (MLMS) framework for that selects and interchanges machine learning (ML) models The MLMS framework includes an ML model search strategy that can flexibly adapt ML models for a wide variety of compute system and/or environmental changes [0006]: FIG. 1 depicts an overview of a machine learning model swapping (MLMS) system according to various embodiments. [0008]: The present disclosure is related to techniques for optimizing artificial intelligence (AI) and/or machine learning (ML) models to reduce resource consumption while improving AI/ML model performance. In particular, the present disclosure provides a ML architecture search (MLAS) framework that involves generalized and/or hardware (HW)-aware ML architectures. [0019] For purposes of the present disclosure, the term “ML architecture” may refer to a particular ML model having a particular set of ML parameters and/or such an ML model configured to be operated on a particular HW platform [0017]: FIG. 1 shows the main components and interactions of an ML model scaling (MLMS) system 100. MLMS system 100 provides a holistic system that selects and interchanges ML models in an energy and communication efficient way while at the same time adapting the models to real-time changes in system constraints, conditions, and/or configurations. - a super-net operatively coupled to the supervised learning controller and including a plurality of nodes coupled by operational blocks, an operational block including two or more neural networks [0014]: The MLMS system implements a unique and robust ML model search and swap strategy that can flexibly adapt ML models for a wide variety of detected compute system operational changes and/or environmental changes. [0021] “ In alternative implementations, the MLMS system 100 may be operated by the same compute node (e.g., client device, edge compute node, network access node, cloud service, drone, network appliance, etc.)”, [0033]:”In some implementations, the HW platform may be the client device 101. In some implementations, the HW platform may be some other desired compute node or device on which the user wishes to deploy an ML model or wishes to perform some AI/ML task”, [0027]: The AI/ML tasks may describe a desired problem to be solved and the AI/ML domain may describe a desired goal to be achieved. Examples of ML tasks include clustering, classification, regression, anomaly detection, data cleaning, automated ML (autoML), association rules learning, reinforcement learning, structured prediction, feature engineering, feature learning, online learning, supervised learning, semi-supervised learning (SSL), unsupervised learning, machine learned ranking (MLR), grammar induction, and/or the like. [BRI: the MLMS (Machine Learning Model Swap) system implements a unique and robust ML model search and swap strategy that can flexibly adapt ML models to a wide variety of detected compute system and/or environmental changes, treating these changes as “operational blocks” for model adaptation] - each having a plurality of network weights [0019]: model parameters are parameters derived via training, whereas hyperparameters are parameters whose values are used to control aspects of the learning process and usually have to be set before running an ML model (e.g., weights, etc.). Additionally, for purposes of the present disclosure, hyperparameters may be classified as architectural hyperparameters or training hyperparameters. Architectural hyperparameters are hyperparameters that are related to architectural aspects of an ML model - where different neural networks in the same operational block share at least some of the same network weights, [0014] : The present disclosure provides an MLMS system that selects and interchanges ML models (e.g., subnets) in an energy and communication efficient way while adapting the ML models (e.g., subnets) to real time (or near-real time) changes in system (HW platform) constraints. The MLMS system discussed herein is agnostic to particular approach used for ML model search (e.g., NAS) and for ML model (e.g., supernet) weight sharing. The MLMS system implements a unique and robust ML model search and swap strategy that can flexibly adapt ML models for a wide variety of detected compute system operational changes and/or environmental changes. - and data loader configured to provide the training inputs to the super-net hardware to generate corresponding super-net outputs; where the supervised learning controller is configured to train network weights and architecture parameters of the super-net, [0014]: The MLMS system implements a unique and robust ML model search and swap strategy that can flexibly adapt ML models for a wide variety of detected compute system operational changes and/or environmental changes. [0015]: the MLMS system discussed herein provides a relatively low dynamic power signature (e.g., during ML model swapping) and a relatively fast memory write time and/or model loading time during model interchange operations. [0022]: The supernet 130 may be an ML model configured and/or trained for a particular AI/ML task from which the MLMS system 100 is to discover or generate a smaller and/or derivative ML model (referred to herein as a “sub-network” or “subnet”). [0022]: [0022]: the supernet 130 contains one or more smaller subnets (e.g., the subnets 135 in the subnet pool 147 in FIG. 1) that, when trained in isolation, can match and/or offer trade-offs in various objectives and/or performance metrics such as accuracy and/or latency of the original supernet 130 when trained for the same number of iterations or epochs. [0015]: The MLMS system discussed herein allows prediction/inference tasks to be performed on a HW platform under dynamically varying environmental conditions (e.g., temperature, battery charge) and HW configurations (e.g. utilized cores, power modes, etc.) while maintaining the same or similar level of performance of prediction/inference results (e.g., accuracy, latency, etc.) as if the HW platform were operated in a more stable environment and/or with static HW configurations - where the supervised learning controller is configured to train network weights and architecture parameters of the super-net, [0027]: The AI/ML tasks may describe a desired problem to be solved and the AI/ML domain may describe a desired goal to be achieved. Examples of ML tasks include clustering, classification, regression, anomaly detection, data cleaning, automated ML (autoML), association rules learning, reinforcement learning, structured prediction, feature engineering, feature learning, online learning, supervised learning [0178] The term “multi-head attention” at least in some embodiments refers to an attention technique that combines several different attention mechanisms to direct the overall attention of a network or subnetwork. [0179]: the goal is to break down complicated tasks into smaller areas of attention that are processed sequentially. [0179]: The term “attention network” at least in some embodiments refers to an artificial neural networks used for attention in machine learning. [0019]: model parameters are parameters derived via training, whereas hyperparameters are parameters whose values are used to control aspects of the learning process and usually have to be set before running an ML model (e.g., weights, etc.). Additionally, for purposes of the present disclosure, hyperparameters may be classified as architectural hyperparameters or training hyperparameters. Architectural hyperparameters are hyperparameters that are related to architectural aspects of an ML model [0022: In various implementations, the ML config includes a reference ML model 130 (referred to herein as a “super network 130” or “supernet 130 [0022]: The supernet 130 may be an ML model configured and/or trained for a particular AI/ML task from which the MLMS system 100 is to discover or generate a smaller and/or derivative ML model (referred to herein as a “sub-network” or “subnet”). [0022]: This supernet 130 may include parameters and/or weights that do not significantly contribute to the prediction and/or inference determination, and these parameters and/or weights contribute to the supernet's overall computational complexity and density. Therefore, the supernet 130 contains one or more smaller subnets (e.g., the subnets 135 in the subnet pool 147 in FIG. 1) that, when trained in isolation, can match and/or offer trade-offs in various objectives and/or performance metrics such as accuracy and/or latency of the original supernet 130 when trained for the same number of iterations or epochs. [0022]: In the context of the present disclosure, the term “elastic” refers to the ability of an elastic supernet 130 to add and/or remove different ML parameters and/or subnets 135 without restarting the running framework and/or the running ML model. [BRI: Perhaps known to a POSITA that elastic supernet is a type of over-parameterized neural network that contains many candidate subnetworks within a single, larger model, enabling neural architecture search (NAS) to explore a wide range of architectures efficiently ] - and where the architecture parameters represent the importance of different architecture choices at various locations inside the super-net [0019]: For purposes of the present disclosure, the term “ML architecture” may refer to a particular ML model having a particular set of ML parameters and/or such an ML model configured to be operated on a particular HW platform. Here, the set of ML parameters may refer to “model parameters” (also referred to simply as “parameters”) and/or “hyperparameters.” Model parameters are parameters derived via training, whereas hyperparameters are parameters whose values are used to control aspects of the learning process and usually have to be set before running an ML model (e.g., weights, etc.). Additionally, for purposes of the present disclosure, hyperparameters may be classified as architectural hyperparameters or training hyperparameters. Architectural hyperparameters are hyperparameters that are related to architectural aspects of an ML model such as, for example, the number of layers in a DNN, specific layer types in a DNN (e.g., convolutional layers, multilayer perception (MLP) layers, etc.), number of output channels, kernel size, and/or the like. [0178]: The term “attention” in the context of machine learning and/or neural networks, at least in some embodiments refers to a technique that mimics cognitive attention, which enhances important parts of a dataset where the important parts of the dataset may be determined using training data by gradient descent - for a plurality of training inputs and corresponding training outputs in training data for a chosen task [0002]: Machine learning (ML) is the study of computer algorithms that improve automatically through experience and by the use of data. [0002]: ML algorithms build models using sample data (referred to as “training data”) and/or based on past experience in order to make predictions or decisions without being explicitly programmed to do so. [0077] : Machine learning (ML) involves programming computing systems to optimize a performance criterion using example (training) data and/or past experience. [0077]: ML involves using algorithms to perform specific task(s) without using explicit instructions to perform the specific task(s) [0013]: different solutions that need to be searched for training these separate ML models. - configuring , by the supervising learning controller, the super-net to implement a sample of sub-networks [0058]: FIG. 3 shows an example of the efficiency-aware subnet model selection after the subnet search is performed, according to various embodiments. The supernet search results 301 shows all discovered subnets 335 found during a subnet search, which includes the existing supernet 330 and a set of candidate replacement subnets 340 (individual subnets in the set of candidate replacement subnets 340 may be referred to herein as “replacement subnets 340”, “candidate subnets 340”, or the like). [BRI: candidate subnets represent sample of subnets] - generating, by the super-net, sub-network outputs responsive to the training inputs for the sample of sub-networks, where a sample sub-network includes a neural network selected from each operational block; [0014] : The present disclosure provides an MLMS system that selects and interchanges ML models (e.g., subnets) in an energy and communication efficient way while adapting the ML models (e.g., subnets) to real time (or near-real time) changes in system (HW platform) constraints [BRI: A MLMS system is an operational block] [0058]: FIG. 3 shows an example of the efficiency-aware subnet model selection after the subnet search is performed, according to various embodiments. The supernet search results 301 shows all discovered subnets 335 found during a subnet search, which includes the existing supernet 330 and a set of candidate replacement subnets 340 (individual subnets in the set of candidate replacement subnets 340 may be referred to herein as “replacement subnets 340”, “candidate subnets 340”, or the like). [BRI: candidate subnets represents sample of subnets] - accumulating a first loss based on the difference between sub-network outputs and corresponding training outputs [0059]: In embodiments, the subnet selector 141 implements a search function and/or NAS algorithm to identify as many possible subnet candidates 340 for the new context information (e.g., new HW configuration) that precipitated the search. [0209]: The term “loss function” or “cost function” at least in some embodiments refers to an event or values of one or more variables onto a real number that represents some “cost” associated with the event. A value calculated by a loss function may be referred to as a “loss” or “error”. Additionally or alternatively, the term “loss function” or “cost function” at least in some embodiments refers to a function used to determine the error or loss between the output of an algorithm and a target value. Additionally or alternatively, the term “loss function” or “cost function” at least in some embodiments refers to a function are used in optimization problems with the goal of minimizing a loss or error. [BRI: Perhaps known to a POSITA, that cost function is typically the average (or sum) of the loss function over all samples in a dataset. It aggregates the per-sample losses to give an overall measure of model performance. Thus, in many embodiments, especially in deep learning or multi-model setups, the “loss function” at the per-sample level and the “cost function” at the dataset level both represent accumulated differences between model outputs and target values, with the cost function being the total or averaged loss across all samples] - accumulating a second loss over the sample of sub-networks, [0059]: In embodiments, the subnet selector 141 implements a search function and/or NAS algorithm to identify as many possible subnet candidates 340 for the new context information (e.g., new HW configuration) that precipitated the search. [0209]: The term “loss function” or “cost function” at least in some embodiments refers to an event or values of one or more variables onto a real number that represents some “cost” associated with the event. A value calculated by a loss function may be referred to as a “loss” or “error”. Additionally or alternatively, the term “loss function” or “cost function” at least in some embodiments refers to a function used to determine the error or loss between the output of an algorithm and a target value. Additionally or alternatively, the term “loss function” or “cost function” at least in some embodiments refers to a function are used in optimization problems with the goal of minimizing a loss or error. [0059]: Of the possible subnet candidates 340 that are found, preference is given to a subnet 340 in the set of candidates 340 that is the most related to the existing subnet 135. The determination of the most closely related subnet 340 is based on the supernet type weight/parameter sharing scheme, such as the vectorization scheme discussed previously with respect to FIG. 2, which can be used to calculate the total transfer cost for each of the candidates 340. The subnet 340 with a minimal total transfer cost (e.g., with respect to the other candidate subnets 340) is selected as a new active subnet. [BRI: The transfer cost is the second loss] - select a sub-network of the plurality of sub-networks for the chosen task based on the largest adjusted architecture parameters; [0017]: The MLMS system 100 achieves energy and communication efficiency by using a similarity-based subnet selection process where a subnet is selected from a pool of subnets that has the most overlap in pre-trained parameters from the existing subnet to minimize memory write operation overhead [0019]: Architectural hyperparameters are hyperparameters that are related to architectural aspects of an ML model such as, for example, the number of layers in a DNN, specific layer types in a DNN (e.g., convolutional layers, multilayer perception (MLP) layers, etc.), number of output channels, kernel size, and/or the like. [0022]: “elasticity” refers to a component in the supernet 130 and its corresponding operator can change in several dimensions. For example, in convolutional networks this would allow for the choice of different kernel sizes (e.g., elastic kernel), depths of some group of elements (e.g., elastic depth), and number of channels of the selected components (e.g., elastic width). In the context of the present disclosure, the term “elastic” refers to the ability of an elastic supernet 130 to add and/or remove different ML parameters and/or subnets 135 without restarting the running framework and/or the running ML model. [BRI: Perhaps known to a POISTA that adjusted architecture parameters refer to the continuous or discrete variables that define a candidate network’s structure, which are optimized during the search process. These parameters can include: filter(kernel), number of layers, and number of channels. - adjusting network weights and architecture parameters of the super-net to reduce a combination of the first and second losses, [0214] : The term “objective function” at least in some embodiments refers to a function to be maximized or minimized for a specific optimization problem. In some cases, an objective function is defined by its decision variables and an objective. The objective is the value, target, or goal to be optimized, such as maximizing profit or minimizing usage of a particular resource. The specific objective function chosen depends on the specific problem to be solved and the objectives to be optimized. [0209]: the term “loss function” or “cost function” at least in some embodiments refers to a function are used in optimization problems with the goal of minimizing a loss or error. [0059]: Of the possible subnet candidates 340 that are found, preference is given to a subnet 340 in the set of candidates 340 that is the most related to the existing subnet 135. The determination of the most closely related subnet 340 is based on the supernet type weight/parameter sharing scheme, such as the vectorization scheme discussed previously with respect to FIG. 2, which can be used to calculate the total transfer cost for each of the candidates 340. The subnet 340 with a minimal total transfer cost (e.g., with respect to the other candidate subnets 340) is selected as a new active subnet. [BRI: Perhaps known to a POSITA that in deep learning combing a first loss (subnetwork outputs vs. training outputs) with a second loss (e.g., sample-based subnetwork evaluation) into a single composite loss function is common in multi-objective or multi-task scenarios, such as multi-subnetwork training. This can be achieved via weighted summation. Within the context of objective function being optimized, the total transfer cost is combination of first and second loss] - where the network weights shared between different neural networks of an operational block are trained jointly, [0014] : The present disclosure provides an MLMS system that selects and interchanges ML models (e.g., subnets) in an energy and communication efficient way while adapting the ML models (e.g., subnets) to real time (or near-real time) changes in system (HW platform) constraints. The MLMS system discussed herein is agnostic to particular approach used for ML model search (e.g., NAS) and for ML model (e.g., supernet) weight sharing. The MLMS system implements a unique and robust ML model search and swap strategy that can flexibly adapt ML models for a wide variety of detected compute system operational changes and/or environmental changes. [0020]: the MLMSI 110a includes application programming interfaces (APIs) to access the other subsystems of system 100, manages ML search parameter updates (e.g., new or updated ML config), and calls any supported ML operations library (e.g., as indicated by the ML config). [0209]: the term “loss function” or “cost function” at least in some embodiments refers to a function are used in optimization problems with the goal of minimizing a loss or error. [0059]: Of the possible subnet candidates 340 that are found, preference is given to a subnet 340 in the set of candidates 340 that is the most related to the existing subnet 135. The determination of the most closely related subnet 340 is based on the supernet type weight/parameter sharing scheme, such as the vectorization scheme discussed previously with respect to FIG. 2, which can be used to calculate the total transfer cost for each of the candidates 340. The subnet 340 with a minimal total transfer cost (e.g., with respect to the other candidate subnets 340) is selected as a new active subnet. [0053]: a subnet 135 that has the most overlap in pre-trained parameters from the existing subnet 135 for a target device (e.g., compute node 550). These mechanisms help the target device adapt to a new system states and/or HW conditions/constraints, but also limit the data-transfer resource consumption when writing the new subnet 135 to memory. [BRI: Within the context of objective function being optimized, and the the total transfer cost is combination of first and second loss as a result of training all components (subnet) together is indeed a form of joint training] - and output a description of the network architecture of the selected sub-network for the chosen task. [0010]: optimizing an ML model for individual HW platforms and specific performance metrics is a very time-consuming effort, which requires highly specialized knowledge. This type of optimization is typically done manually with a great deal of in-depth understanding of the HW platform since certain characteristics of the HW platform (e.g., clock speed, number of processor cores, amount of cache memory, etc.) will affect the optimization process. The optimization process is also affected by the characteristics of the input data to the ML model (e.g., batch size, image size, number of epochs/iterations, etc.). Finally, any change to the performance metrics (e.g., going from latency to power consumption), input data characteristics (e.g., increasing the batch size), [0091]: To provide the inference, the inference engine 516 uses a model 520 that controls how the DNN inference is made on the data 514 to generate the result 518. Specifically, the model 520 includes a topology of layers of a NN. The topology includes an input layer that receives the data 514, an output layer that outputs the result 518, and one or more hidden layers between the input and output layers that provide processing between the data 14 and the result 518. The topology may be stored in a suitable information object, such as an extensible markup language (XML), JavaScript Object Notation (JSON), and/or other suitable data structure, file, and/or the like. The model 520 may also include weights and/or biases for results for any of the layers while processing the data 514 in the inference using the DNN. Cummings do not explicitly disclose: - over layers of sub-networks in a sample of sub-networks, of a measure of a smoothness based solely from network weights in the layers, where measure of smoothness related to the maximum difference between outputs from the layer relative to any possible difference between inputs to the layer, a higher maximum change in output indicating a lower smoothness; - where the second loss penalizes sub- networks with lower smoothness and improves stability of the adjustment; However, Badran discloses: - over layers of sub-networks in a sample of sub-networks, of a measure of a smoothness based solely from network weights in the layers, where measure of smoothness related to the maximum difference between outputs from the layer relative to any possible difference between inputs to the layer, [2, Page 1446]: The neural network smoother uses a basic MLP architecture (R) with an input layer supporting the temporal window, l layers of non-linear hidden neurons and one linear output neuron. Starting from this basic architecture (R), a family of multi-layer perceptrons R λ is generated. For each λ ,   the synaptic weights vector is chosen using the learning procedure described below. [2, 1446]: The input layer uses a temporal window on the time series, which is the vector of observations ( y i - d ' , … y i + d ), d’ ≥ d The two time lags d and d’ can be equal or different. They can be both positive, in which case, d’ represents the amount of information in the past used by the temporal window; d stands for the future and y i   is within the window. However, d may be negative: for example, if d= -1 and y i is not within the window, smoothing at time i uses only past information, and it can be considered as a predictor of the smoothed value at time i. [2, Page 1446]: A classical smoothing method is to choose from the family of twice continuously differentiable functions S, such that PNG media_image1.png 28 107 media_image1.png Greyscale exists, the function S which minimizes the measure of roughness PNG media_image2.png 56 433 media_image2.png Greyscale The weighting parameter λ   controls the amount of smoothing, so that increasing λ   leads to smoother estimate. [BRI: The increase λ to minimizing measure of roughness. [3, Page 1447]: The MASE can be easily expressed as PNG media_image3.png 166 511 media_image3.png Greyscale The first term on the right-hand side of (6) corresponds to a variance component and the second to the square of a bias component, both of them contributing to the MASE(   g λ ) The variance component is decreased by smoothing, but this increases the bias component. The problem is thus to choose a good compromise between these two components (which is done by minimizing (6) with respect to PNG media_image4.png 12 1 media_image4.png Greyscale [BRI: the bias component measures the average difference between the model’s predictions and the best possible model for the problem, averaged over all training data instances and increasing it underfits the data and affects smoothness as represented by equation (6)] - a higher maximum change in output indicating a lower smoothness; [3, Page 1447]: The MASE can be easily expressed as PNG media_image3.png 166 511 media_image3.png Greyscale The first term on the right-hand side of (6) corresponds to a variance component and the second to the square of a bias component, both of them contributing to the MASE(   g λ ) The variance component is decreased by smoothing, but this increases the bias component. The problem is thus to choose a good compromise between these two components (which is done by minimizing (6) with respect to PNG media_image4.png 12 1 media_image4.png Greyscale [BRI: in effect, the high difference resulted from the bias component does increases MASE and thus lower the smoothness] - the accumulated loss based, at least in part, on a sum, over layers of a sub-network, of the measures of smoothness of the layers; [2, Page 1446]: The neural network smoother uses a basic MLP architecture (R) with an input layer supporting the temporal window, l layers of non-linear hidden neurons and one linear output neuron. Starting from this basic architecture (R), a family of multi-layer perceptrons (Rl) is generated. For each l, the synaptic weights vector is chosen using the learning procedure described below. [2, 1446]: The input layer uses a temporal window on the time series, which is the vector of observations ( y i - d ' , … y i + d ), d’ ≥ d The two time lags d and d’ can be equal or different. They can be both positive, in which case, d’ represents the amount of information in the past used by the temporal window; d stands for the future and y i   is within the window. However, d may be negative: for example, if d= -1 and y i is not within the window, smoothing at time i uses only past information, and it can be considered as a predictor of the smoothed value at time i. [2, Page 1446]: A classical smoothing method is to choose from the family of twice continuously differentiable functions S, such that PNG media_image1.png 28 107 media_image1.png Greyscale exists, the function S which minimizes the measure of roughness [BRI: function S that minimizes the measure of roughness provides the measure of smoothness] - where the second loss penalizes sub- networks with lower smoothness and improves stability of the adjustment; [Abstract, Page 1445]: The NN-smoother computes the trend in the state domain and minimizes a cost function with a regularization term. The regularization term is penalized by a parameter λ which forces the learning procedure to smooth in the time domain. We define a selecting criterion in order to select the best parameter λ . [1, Page 1445]: In most cases, smoothing techniques generate families of functions g λ .     d e pending on a parameter λ , which controls the degree of smoothing. Each of these functions is a solution to the minimization of a cost function, containing a regularization term parameterized by λ . [3, Page 1447]: The MASE can be easily expressed as PNG media_image3.png 166 511 media_image3.png Greyscale The first term on the right-hand side of (6) corresponds to a variance component and the second to the square of a bias component, both of them contributing to the MASE(   g λ ). The variance component is decreased by smoothing, but this increases the bias component. The problem is thus to choose a good compromise between these two components (which is done by minimizing (6) with respect to PNG media_image4.png 12 1 media_image4.png Greyscale [BRI: in effect, the high difference resulted from the bias component does increases MASE and thus lower the smoothness and is solved by minimizing MASE by using compilation since the regularization (penalty) is parameterized by λ ] [1, Page 1445]: Neural networks can be used for smoothing: for example, an MLP, trained through a modified gradient back propagation, can implement spline smoothing in the time domain [1, Page 1446]: The neural network smoother uses a basic MLP architecture (R) with an input layer supporting the temporal window, l layers of non-linear hidden neurons and one linear output neuron. Starting from this basic architecture (R), a family of multi-layer perceptrons (Rl) is generated. For each λ , the synaptic weights vector is chosen [BRI: in the time domain, spline smoothing with a large smoothing parameter λ does represent small changes in output with respect to input, because it penalizes curvature and enforces a smooth, slowly varying curve. In fact, this is exactly what Lipschitz constant provides (small changes to output- see spec [0022]) that minimize the loss jointly for all sub-networks] [2, Page 1447]:] We consider the measure of roughness PNG media_image5.png 85 480 media_image5.png Greyscale [2, Page 1447]: The learning phase consists in minimizing the discrete approximation given in (4) by the back propagation algorithm. [4, Page 1450]: NN performs a smoothing in the state space using a cost function C(l) (4) which acts in the time domain. So, the NN smoother combines two representations: the state space domain and the time domain. The NN network directly minimizes (4) which is the discrete form of (2) and gives the optimal smoothing directly for each discrete time unit i [BRI: In the state-space domain: the internal dynamics of the system, modeled as a set of state equations (continuous or discrete) and in the time-domain data: the actual measured inputs, outputs, and (often) states over time, used to estimate the network weights and biases. Several sub-networks are part of the same state-space model. The same set of weights (shared) is updated during training to ensure consistency across the block and the joint training leverages all available time-domain data from the block to optimize shared parameter. The NN directly minimize the cost function(loss function) given by (4)] [2, Page 1447]: The input layer has d þ d þ 3 units (for window [yi¹1¹d,…,yiþ1þd]), l hidden layers (l 0) and 3 output units. The first hidden layer has three clusters of n1 units each. The first cluster of hidden units is connected to window {i ¹ 1 ¹ d,…,i ¹ 1 þ d}, the second to the window shifted by one time unit, and the third to the window shifted by one more unit; these clusters share the same weight vector. The use of time-delay inputs convert the time series into a static mapping problem. The only constraint is that each hidden layer has three clusters of units connected only to the corresponding cluster in the previous layer. In each layer, clusters share the same weight vector. It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Cummings, and Badran. Cummings teaches a super-net including a plurality of sub-networks, first loss over a sample of sub-networks, and second loss and adjusting network weights to reduce the accumulated loss. Badran teaches relating the measure of smoothness to the maximum difference between outputs from the layer relative to any possible difference between inputs to the layer. One of ordinary skill would be motivated to combine Cummings, and Badran that choose best smoothing under different hypothesis and minimum error ( Badran [5, Page 1450]). Claim 10: (Previously Presented) Cummings discloses: - train network weights of the selected sub-network using the training data; [0022]: supernet 130 may include parameters and/or weights that do not significantly contribute to the prediction and/or inference determination, and these parameters and/or weights contribute to the supernet's overall computational complexity and density. Therefore, the supernet 130 contains one or more smaller subnets (e.g., the subnets 135 in the subnet pool 147 in FIG. 1) that, when trained in isolation, can match and/or offer trade-offs in various objectives and/or performance metrics such as accuracy and/or latency of the original supernet 130 when trained for the same number of iterations or epochs. Claim 9: (Currently Amended) Cummings discloses: - A system for automated selection of neural architecture, the method comprising: [0027]: The AI/ML tasks may describe a desired problem to be solved and the AI/ML domain may describe a desired goal to be achieved. Examples of ML tasks include clustering, classification, regression, anomaly detection, data cleaning, automated ML (autoML), association rules learning, reinforcement learning, structured prediction, feature engineering, feature learning, online learning, supervised learning [0008]: The present disclosure is related to techniques for optimizing artificial intelligence (AI) and/or machine learning (ML) models to reduce resource consumption while improving AI/ML model performance. In particular, the present disclosure provides a ML architecture search (MLAS) framework that involves generalized and/or hardware (HW)-aware ML architectures. [0019] For purposes of the present disclosure, the term “ML architecture” may refer to a particular ML model having a particular set of ML parameters and/or such an ML model configured to be operated on a particular HW platform [0017]: FIG. 1 shows the main components and interactions of an ML model scaling (MLMS) system 100. MLMS system 100 provides a holistic system that selects and interchanges ML models in an energy and communication efficient way while at the same time adapting the models to real-time changes in system constraints, conditions, and/or configurations. - a super-net operatively coupled to the supervised learning controller and including a plurality of nodes coupled by operational blocks, an operational block including two or more neural networks [0019]: Model parameters are parameters derived via training, whereas hyperparameters are parameters whose values are used to control aspects of the learning process and usually have to be set before running an ML model (e.g., weights, etc.). [0091]: To provide the inference, the inference engine 516 uses a model 520 that controls how the DNN inference is made on the data 514 to generate the result 518. Specifically, the model 520 includes a topology of layers of a NN. [BRI: the supervised learning controller may regulate how the supernet is trained, including: Architecture sampling policies for evaluation of subnetworks and weight sharing strategies to maintain architectural diversity while reusing parameters] [0163]: The term “compute node” or “compute device” at least in some embodiments refers to an identifiable entity implementing an aspect of computing operations, whether part of a larger system, distributed collection of systems, or a standalone apparatus. [0100]: In some implementations, the processor circuitry 552 and/or acceleration circuitry 564 may include HW elements specifically tailored for machine learning functionality, such as for operating performing ANN operations, [0014]: The MLMS system implements a unique and robust ML model search and swap strategy that can flexibly adapt ML models for a wide variety of detected compute system operational changes and/or environmental changes. [0021] In alternative implementations, the MLMS system 100 may be operated by the same compute node (e.g., client device, edge compute node, network access node, cloud service, drone, network appliance, etc.), [0033]: In some implementations, the HW platform may be the client device 101. In some implementations, the HW platform may be some other desired compute node or device on which the user wishes to deploy an ML model or wishes to perform some AI/ML task, [0027]: The AI/ML tasks may describe a desired problem to be solved and the AI/ML domain may describe a desired goal to be achieved. Examples of ML tasks include clustering, classification, regression, anomaly detection, data cleaning, automated ML (autoML), association rules learning, reinforcement learning, structured prediction, feature engineering, feature learning, online learning, supervised learning, semi-supervised learning (SSL), unsupervised learning, machine learned ranking (MLR), grammar induction, and/or the like. [BRI: the MLMS (Machine Learning Model Swap) system implements a unique and robust ML model search and swap strategy that can flexibly adapt ML models to a wide variety of detected compute system and/or environmental changes, treating these changes as “operational blocks” for model adaptation] - each having a plurality of network weights [0019]: model parameters are parameters derived via training, whereas hyperparameters are parameters whose values are used to control aspects of the learning process and usually have to be set before running an ML model (e.g., weights, etc.). Additionally, for purposes of the present disclosure, hyperparameters may be classified as architectural hyperparameters or training hyperparameters. Architectural hyperparameters are hyperparameters that are related to architectural aspects of an ML model - where different neural networks in the same operational block share at least some of the same network weights, [0014] : The present disclosure provides an MLMS system that selects and interchanges ML models (e.g., subnets) in an energy and communication efficient way while adapting the ML models (e.g., subnets) to real time (or near-real time) changes in system (HW platform) constraints. The MLMS system discussed herein is agnostic to particular approach used for ML model search (e.g., NAS) and for ML model (e.g., supernet) weight sharing. The MLMS system implements a unique and robust ML model search and swap strategy that can flexibly adapt ML models for a wide variety of detected compute system operational changes and/or environmental changes. - and data loader configured to provide the training inputs to the super-net hardware to generate corresponding super-net outputs; where the supervised learning controller is configured to train network weights and architecture parameters of the super-net, [0022]: The supernet 130 may be an ML model configured and/or trained for a particular AI/ML task from which the MLMS system 100 is to discover or generate a smaller and/or derivative ML model (referred to herein as a “sub-network” or “subnet”). [0022]: [0022]: the supernet 130 contains one or more smaller subnets (e.g., the subnets 135 in the subnet pool 147 in FIG. 1) that, when trained in isolation, can match and/or offer trade-offs in various objectives and/or performance metrics such as accuracy and/or latency of the original supernet 130 when trained for the same number of iterations or epochs. [0015]: The MLMS system discussed herein allows prediction/inference tasks to be performed on a HW platform under dynamically varying environmental conditions (e.g., temperature, battery charge) and HW configurations (e.g. utilized cores, power modes, etc.) while maintaining the same or similar level of performance of prediction/inference results (e.g., accuracy, latency, etc.) as if the HW platform were operated in a more stable environment and/or with static HW configurations - where the supervised learning controller is configured to train network weights and architecture parameters of the super-net, [0100]: In some implementations, the processor circuitry 552 and/or acceleration circuitry 564 may include HW elements specifically tailored for machine learning functionality, such as for operating performing ANN operations, [0027]: The AI/ML tasks may describe a desired problem to be solved and the AI/ML domain may describe a desired goal to be achieved. Examples of ML tasks include clustering, classification, regression, anomaly detection, data cleaning, automated ML (autoML), association rules learning, reinforcement learning, structured prediction, feature engineering, feature learning, online learning, supervised learning [0178] The term “multi-head attention” at least in some embodiments refers to an attention technique that combines several different attention mechanisms to direct the overall attention of a network or subnetwork. [0179]: the goal is to break down complicated tasks into smaller areas of attention that are processed sequentially. [0179]: The term “attention network” at least in some embodiments refers to an artificial neural networks used for attention in machine learning. [0019]: model parameters are parameters derived via training, whereas hyperparameters are parameters whose values are used to control aspects of the learning process and usually have to be set before running an ML model (e.g., weights, etc.). Additionally, for purposes of the present disclosure, hyperparameters may be classified as architectural hyperparameters or training hyperparameters. Architectural hyperparameters are hyperparameters that are related to architectural aspects of an ML model [0022: In various implementations, the ML config includes a reference ML model 130 (referred to herein as a “super network 130” or “supernet 130 [0022]: The supernet 130 may be an ML model configured and/or trained for a particular AI/ML task from which the MLMS system 100 is to discover or generate a smaller and/or derivative ML model (referred to herein as a “sub-network” or “subnet”). [0022]: This supernet 130 may include parameters and/or weights that do not significantly contribute to the prediction and/or inference determination, and these parameters and/or weights contribute to the supernet's overall computational complexity and density. Therefore, the supernet 130 contains one or more smaller subnets (e.g., the subnets 135 in the subnet pool 147 in FIG. 1) that, when trained in isolation, can match and/or offer trade-offs in various objectives and/or performance metrics such as accuracy and/or latency of the original supernet 130 when trained for the same number of iterations or epochs. [0022]: In the context of the present disclosure, the term “elastic” refers to the ability of an elastic supernet 130 to add and/or remove different ML parameters and/or subnets 135 without restarting the running framework and/or the running ML model. [BRI: Perhaps known to a POSITA that elastic supernet is a type of over-parameterized neural network that contains many candidate subnetworks within a single, larger model, enabling neural architecture search (NAS) to explore a wide range of architectures efficiently ] - and where the architecture parameters represent the importance of different architecture choices at various locations inside the super-net [0019]: For purposes of the present disclosure, the term “ML architecture” may refer to a particular ML model having a particular set of ML parameters and/or such an ML model configured to be operated on a particular HW platform. Here, the set of ML parameters may refer to “model parameters” (also referred to simply as “parameters”) and/or “hyperparameters.” Model parameters are parameters derived via training, whereas hyperparameters are parameters whose values are used to control aspects of the learning process and usually have to be set before running an ML model (e.g., weights, etc.). Additionally, for purposes of the present disclosure, hyperparameters may be classified as architectural hyperparameters or training hyperparameters. Architectural hyperparameters are hyperparameters that are related to architectural aspects of an ML model such as, for example, the number of layers in a DNN, specific layer types in a DNN (e.g., convolutional layers, multilayer perception (MLP) layers, etc.), number of output channels, kernel size, and/or the like. [0178]: The term “attention” in the context of machine learning and/or neural networks, at least in some embodiments refers to a technique that mimics cognitive attention, which enhances important parts of a dataset where the important parts of the dataset may be determined using training data by gradient descent - for a plurality of training inputs and corresponding training outputs in training data for a chosen task [0002]: Machine learning (ML) is the study of computer algorithms that improve automatically through experience and by the use of data. [0002]: ML algorithms build models using sample data (referred to as “training data”) and/or based on past experience in order to make predictions or decisions without being explicitly programmed to do so. [0077] : Machine learning (ML) involves programming computing systems to optimize a performance criterion using example (training) data and/or past experience. [0077]: ML involves using algorithms to perform specific task(s) without using explicit instructions to perform the specific task(s) [0013]: different solutions that need to be searched for training these separate ML models. - configuring , by the supervising learning controller, the super-net to implement a sample of sub-networks [0100]: In some implementations, the processor circuitry 552 and/or acceleration circuitry 564 may include HW elements specifically tailored for machine learning functionality, such as for operating performing ANN operations, - generating, by the super-net, sub-network outputs responsive to the training inputs for the sample of sub-networks, where a sample sub-network includes a neural network selected from each operational block; [0014] : The present disclosure provides an MLMS system that selects and interchanges ML models (e.g., subnets) in an energy and communication efficient way while adapting the ML models (e.g., subnets) to real time (or near-real time) changes in system (HW platform) constraints [BRI: A MLMS system is an operational block] [0058]: FIG. 3 shows an example of the efficiency-aware subnet model selection after the subnet search is performed, according to various embodiments. The supernet search results 301 shows all discovered subnets 335 found during a subnet search, which includes the existing supernet 330 and a set of candidate replacement subnets 340 (individual subnets in the set of candidate replacement subnets 340 may be referred to herein as “replacement subnets 340”, “candidate subnets 340”, or the like). [BRI: candidate subnets represents sample of subnets] - accumulating a first loss based on the difference between sub-network outputs and corresponding training outputs [0059]: In embodiments, the subnet selector 141 implements a search function and/or NAS algorithm to identify as many possible subnet candidates 340 for the new context information (e.g., new HW configuration) that precipitated the search. [0209]: The term “loss function” or “cost function” at least in some embodiments refers to an event or values of one or more variables onto a real number that represents some “cost” associated with the event. A value calculated by a loss function may be referred to as a “loss” or “error”. Additionally or alternatively, the term “loss function” or “cost function” at least in some embodiments refers to a function used to determine the error or loss between the output of an algorithm and a target value. Additionally or alternatively, the term “loss function” or “cost function” at least in some embodiments refers to a function are used in optimization problems with the goal of minimizing a loss or error. [BRI: Perhaps known to a POSITA, that cost function is typically the average (or sum) of the loss function over all samples in a dataset. It aggregates the per-sample losses to give an overall measure of model performance. Thus, in many embodiments, especially in deep learning or multi-model setups, the “loss function” at the per-sample level and the “cost function” at the dataset level both represent accumulated differences between model outputs and target values, with the cost function being the total or averaged loss across all samples] - accumulating a second loss over the sample of sub-networks, [0059]: In embodiments, the subnet selector 141 implements a search function and/or NAS algorithm to identify as many possible subnet candidates 340 for the new context information (e.g., new HW configuration) that precipitated the search. [0209]: The term “loss function” or “cost function” at least in some embodiments refers to an event or values of one or more variables onto a real number that represents some “cost” associated with the event. A value calculated by a loss function may be referred to as a “loss” or “error”. Additionally or alternatively, the term “loss function” or “cost function” at least in some embodiments refers to a function used to determine the error or loss between the output of an algorithm and a target value. Additionally or alternatively, the term “loss function” or “cost function” at least in some embodiments refers to a function are used in optimization problems with the goal of minimizing a loss or error. [0059]: Of the possible subnet candidates 340 that are found, preference is given to a subnet 340 in the set of candidates 340 that is the most related to the existing subnet 135. The determination of the most closely related subnet 340 is based on the supernet type weight/parameter sharing scheme, such as the vectorization scheme discussed previously with respect to FIG. 2, which can be used to calculate the total transfer cost for each of the candidates 340. The subnet 340 with a minimal total transfer cost (e.g., with respect to the other candidate subnets 340) is selected as a new active subnet. [BRI: The transfer cost is the second loss] - select a sub-network of the plurality of sub-networks for the chosen task based on the largest adjusted architecture parameters; [0017]: The MLMS system 100 achieves energy and communication efficiency by using a similarity-based subnet selection process where a subnet is selected from a pool of subnets that has the most overlap in pre-trained parameters from the existing subnet to minimize memory write operation overhead [0019]: Architectural hyperparameters are hyperparameters that are related to architectural aspects of an ML model such as, for example, the number of layers in a DNN, specific layer types in a DNN (e.g., convolutional layers, multilayer perception (MLP) layers, etc.), number of output channels, kernel size, and/or the like. [0022]: “elasticity” refers to a component in the supernet 130 and its corresponding operator can change in several dimensions. For example, in convolutional networks this would allow for the choice of different kernel sizes (e.g., elastic kernel), depths of some group of elements (e.g., elastic depth), and number of channels of the selected components (e.g., elastic width). In the context of the present disclosure, the term “elastic” refers to the ability of an elastic supernet 130 to add and/or remove different ML parameters and/or subnets 135 without restarting the running framework and/or the running ML model. [BRI: Perhaps known to a POISTA that adjusted architecture parameters refer to the continuous or discrete variables that define a candidate network’s structure, which are optimized during the search process. These parameters can include: filter(kernel), number of layers, and number of channels. - adjusting network weights and architecture parameters of the super-net to reduce a combination of the first and second losses, [0214] : The term “objective function” at least in some embodiments refers to a function to be maximized or minimized for a specific optimization problem. In some cases, an objective function is defined by its decision variables and an objective. The objective is the value, target, or goal to be optimized, such as maximizing profit or minimizing usage of a particular resource. The specific objective function chosen depends on the specific problem to be solved and the objectives to be optimized. [0209]: the term “loss function” or “cost function” at least in some embodiments refers to a function are used in optimization problems with the goal of minimizing a loss or error. [0059]: Of the possible subnet candidates 340 that are found, preference is given to a subnet 340 in the set of candidates 340 that is the most related to the existing subnet 135. The determination of the most closely related subnet 340 is based on the supernet type weight/parameter sharing scheme, such as the vectorization scheme discussed previously with respect to FIG. 2, which can be used to calculate the total transfer cost for each of the candidates 340. The subnet 340 with a minimal total transfer cost (e.g., with respect to the other candidate subnets 340) is selected as a new active subnet. [BRI: Perhaps known to a POSITA that in deep learning combing a first loss (subnetwork outputs vs. training outputs) with a second loss (e.g., sample-based subnetwork evaluation) into a single composite loss function is common in multi-objective or multi-task scenarios, such as multi-subnetwork training. This can be achieved via weighted summation. Within the context of objective function being optimized, the total transfer cost is combination of first and second loss] - where the network weights shared between different neural networks of an operational block are trained jointly, [0014] : The present disclosure provides an MLMS system that selects and interchanges ML models (e.g., subnets) in an energy and communication efficient way while adapting the ML models (e.g., subnets) to real time (or near-real time) changes in system (HW platform) constraints. The MLMS system discussed herein is agnostic to particular approach used for ML model search (e.g., NAS) and for ML model (e.g., supernet) weight sharing. The MLMS system implements a unique and robust ML model search and swap strategy that can flexibly adapt ML models for a wide variety of detected compute system operational changes and/or environmental changes. [0020]: the MLMSI 110a includes application programming interfaces (APIs) to access the other subsystems of system 100, manages ML search parameter updates (e.g., new or updated ML config), and calls any supported ML operations library (e.g., as indicated by the ML config). [0209]: the term “loss function” or “cost function” at least in some embodiments refers to a function are used in optimization problems with the goal of minimizing a loss or error. [0059]: Of the possible subnet candidates 340 that are found, preference is given to a subnet 340 in the set of candidates 340 that is the most related to the existing subnet 135. The determination of the most closely related subnet 340 is based on the supernet type weight/parameter sharing scheme, such as the vectorization scheme discussed previously with respect to FIG. 2, which can be used to calculate the total transfer cost for each of the candidates 340. The subnet 340 with a minimal total transfer cost (e.g., with respect to the other candidate subnets 340) is selected as a new active subnet. [0053]: a subnet 135 that has the most overlap in pre-trained parameters from the existing subnet 135 for a target device (e.g., compute node 550). These mechanisms help the target device adapt to a new system states and/or HW conditions/constraints, but also limit the data-transfer resource consumption when writing the new subnet 135 to memory. [BRI: Within the context of objective function being optimized, and the the total transfer cost is combination of first and second loss as a result of training all components (subnet) together is indeed a form of joint training] - and output a description of the network architecture of the selected sub-network for the chosen task. [0010]” optimizing an ML model for individual HW platforms and specific performance metrics is a very time-consuming effort, which requires highly specialized knowledge. This type of optimization is typically done manually with a great deal of in-depth understanding of the HW platform since certain characteristics of the HW platform (e.g., clock speed, number of processor cores, amount of cache memory, etc.) will affect the optimization process. The optimization process is also affected by the characteristics of the input data to the ML model (e.g., batch size, image size, number of epochs/iterations, etc.). Finally, any change to the performance metrics (e.g., going from latency to power consumption), input data characteristics (e.g., increasing the batch size), [0091]: To provide the inference, the inference engine 516 uses a model 520 that controls how the DNN inference is made on the data 514 to generate the result 518. Specifically, the model 520 includes a topology of layers of a NN. The topology includes an input layer that receives the data 514, an output layer that outputs the result 518, and one or more hidden layers between the input and output layers that provide processing between the data 14 and the result 518. The topology may be stored in a suitable information object, such as an extensible markup language (XML), JavaScript Object Notation (JSON), and/or other suitable data structure, file, and/or the like. The model 520 may also include weights and/or biases for results for any of the layers while processing the data 514 in the inference using the DNN. Cummings do not explicitly disclose: - over layers of sub-networks in a sample of sub-networks, of a measure of a smoothness based solely from network weights in the layers, where measure of smoothness related to the maximum difference between outputs from the layer relative to any possible difference between inputs to the layer, a higher maximum change in output indicating a lower smoothness; - where the second loss penalizes sub- networks with lower smoothness and improves stability of the adjustment; However, Badran discloses: - over layers of sub-networks in a sample of sub-networks, of a measure of a smoothness based solely from network weights in the layers, where measure of smoothness related to the maximum difference between outputs from the layer relative to any possible difference between inputs to the layer, [2, Page 1446]: The neural network smoother uses a basic MLP architecture (R) with an input layer supporting the temporal window, l layers of non-linear hidden neurons and one linear output neuron. Starting from this basic architecture (R), a family of multi-layer perceptrons R λ is generated. For each λ ,   the synaptic weights vector is chosen using the learning procedure described below. [2, 1446]: The input layer uses a temporal window on the time series, which is the vector of observations ( y i - d ' , … y i + d ), d’ ≥ d The two time lags d and d’ can be equal or different. They can be both positive, in which case, d’ represents the amount of information in the past used by the temporal window; d stands for the future and y i   is within the window. However, d may be negative: for example, if d= -1 and y i is not within the window, smoothing at time i uses only past information, and it can be considered as a predictor of the smoothed value at time i. [2, Page 1446]: A classical smoothing method is to choose from the family of twice continuously differentiable functions S, such that PNG media_image1.png 28 107 media_image1.png Greyscale exists, the function S which minimizes the measure of roughness PNG media_image2.png 56 433 media_image2.png Greyscale The weighting parameter λ   controls the amount of smoothing, so that increasing λ   leads to smoother estimate. [BRI: The increase λ to minimizing measure of roughness. [3, Page 1447]: The MASE can be easily expressed as PNG media_image3.png 166 511 media_image3.png Greyscale The first term on the right-hand side of (6) corresponds to a variance component and the second to the square of a bias component, both of them contributing to the MASE(   g λ ) The variance component is decreased by smoothing, but this increases the bias component. The problem is thus to choose a good compromise between these two components (which is done by minimizing (6) with respect to PNG media_image4.png 12 1 media_image4.png Greyscale [BRI: the bias component measures the average difference between the model’s predictions and the best possible model for the problem, averaged over all training data instances and increasing it underfits the data and affects smoothness as represented by equation (6)] - a higher maximum change in output indicating a lower smoothness; [3, Page 1447]: The MASE can be easily expressed as PNG media_image3.png 166 511 media_image3.png Greyscale The first term on the right-hand side of (6) corresponds to a variance component and the second to the square of a bias component, both of them contributing to the MASE(   g λ ) The variance component is decreased by smoothing, but this increases the bias component. The problem is thus to choose a good compromise between these two components (which is done by minimizing (6) with respect to PNG media_image4.png 12 1 media_image4.png Greyscale [BRI: in effect, the high difference resulted from the bias component does increases MASE and thus lower the smoothness] - the accumulated loss based, at least in part, on a sum, over layers of a sub-network, of the measures of smoothness of the layers; [2, Page 1446]: The neural network smoother uses a basic MLP architecture (R) with an input layer supporting the temporal window, l layers of non-linear hidden neurons and one linear output neuron. Starting from this basic architecture (R), a family of multi-layer perceptrons (Rl) is generated. For each l, the synaptic weights vector is chosen using the learning procedure described below. [2, 1446]: The input layer uses a temporal window on the time series, which is the vector of observations ( y i - d ' , … y i + d ), d’ ≥ d The two time lags d and d’ can be equal or different. They can be both positive, in which case, d’ represents the amount of information in the past used by the temporal window; d stands for the future and y i   is within the window. However, d may be negative: for example, if d= -1 and y i is not within the window, smoothing at time i uses only past information, and it can be considered as a predictor of the smoothed value at time i. [2, Page 1446]: A classical smoothing method is to choose from the family of twice continuously differentiable functions S, such that PNG media_image1.png 28 107 media_image1.png Greyscale exists, the function S which minimizes the measure of roughness [BRI: function S that minimizes the measure of roughness provides the measure of smoothness] - where the second loss penalizes sub- networks with lower smoothness and improves stability of the adjustment; [Abstract, Page 1445]: The NN-smoother computes the trend in the state domain and minimizes a cost function with a regularization term. The regularization term is penalized by a parameter λ which forces the learning procedure to smooth in the time domain. We define a selecting criterion in order to select the best parameter λ . [1, Page 1445]: In most cases, smoothing techniques generate families of functions g λ .     d e pending on a parameter λ , which controls the degree of smoothing. Each of these functions is a solution to the minimization of a cost function, containing a regularization term parameterized by λ . [3, Page 1447]: The MASE can be easily expressed as PNG media_image3.png 166 511 media_image3.png Greyscale The first term on the right-hand side of (6) corresponds to a variance component and the second to the square of a bias component, both of them contributing to the MASE(   g λ ). The variance component is decreased by smoothing, but this increases the bias component. The problem is thus to choose a good compromise between these two components (which is done by minimizing (6) with respect to PNG media_image4.png 12 1 media_image4.png Greyscale [BRI: in effect, the high difference resulted from the bias component does increases MASE and thus lower the smoothness and is solved by minimizing MASE by using compilation since the regularization (penalty) is parameterized by λ ] [1, Page 1445]: Neural networks can be used for smoothing: for example, an MLP, trained through a modified gradient back propagation, can implement spline smoothing in the time domain [1, Page 1446]: The neural network smoother uses a basic MLP architecture (R) with an input layer supporting the temporal window, l layers of non-linear hidden neurons and one linear output neuron. Starting from this basic architecture (R), a family of multi-layer perceptrons (Rl) is generated. For each λ , the synaptic weights vector is chosen [BRI: in the time domain, spline smoothing with a large smoothing parameter λ does represent small changes in output with respect to input, because it penalizes curvature and enforces a smooth, slowly varying curve. In fact, this is exactly what Lipschitz constant provides (small changes to output- see spec [0022]) that minimize the loss jointly for all sub-networks] [2, Page 1447]:] We consider the measure of roughness PNG media_image5.png 85 480 media_image5.png Greyscale [2, Page 1447]: The learning phase consists in minimizing the discrete approximation given in (4) by the back propagation algorithm. [4, Page 1450]: NN performs a smoothing in the state space using a cost function C(l) (4) which acts in the time domain. So, the NN smoother combines two representations: the state space domain and the time domain. The NN network directly minimizes (4) which is the discrete form of (2) and gives the optimal smoothing directly for each discrete time unit i [BRI: In the state-space domain: the internal dynamics of the system, modeled as a set of state equations (continuous or discrete) and in the time-domain data: the actual measured inputs, outputs, and (often) states over time, used to estimate the network weights and biases. Several sub-networks are part of the same state-space model. The same set of weights (shared) is updated during training to ensure consistency across the block and the joint training leverages all available time-domain data from the block to optimize shared parameter. The NN directly minimize the cost function(loss function) given by (4)] [2, Page 1447]: The input layer has d þ d þ 3 units (for window [yi¹1¹d,…,yiþ1þd]), l hidden layers (l 0) and 3 output units. The first hidden layer has three clusters of n1 units each. The first cluster of hidden units is connected to window {i ¹ 1 ¹ d,…,i ¹ 1 þ d}, the second to the window shifted by one time unit, and the third to the window shifted by one more unit; these clusters share the same weight vector. The use of time-delay inputs convert the time series into a static mapping problem. The only constraint is that each hidden layer has three clusters of units connected only to the corresponding cluster in the previous layer. In each layer, clusters share the same weight vector. It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Cummings, and Badran. Cummings teaches a super-net including a plurality of sub-networks, first loss over a sample of sub-networks, and second loss and adjusting network weights to reduce the accumulated loss. Badran teaches relating the measure of smoothness to the maximum difference between outputs from the layer relative to any possible difference between inputs to the layer. One of ordinary skill would be motivated to combine Cummings and Badran that choose best smoothing under different hypothesis and minimum error ( Badran [5, Page 1450]). Claim 15: (Currently Amended) Cummings discloses: - A data processing system for automated selection of neural architecture comprising: [0004] Instead of manually designing an ML model, Neural Architecture Search (NAS) algorithms can be used to automatically discover an ideal ML model for a particular task [0014]: The MLMS system discussed herein is agnostic to particular approach used for ML model search (e.g., NAS) and for ML model (e.g., supernet) weight sharing. [Abstract]: The present disclosure is related to machine learning model swap (MLMS) framework for that selects and interchanges machine learning (ML) models The MLMS framework includes an ML model search strategy that can flexibly adapt ML models for a wide variety of compute system and/or environmental changes [0006]: FIG. 1 depicts an overview of a machine learning model swapping (MLMS) system according to various embodiments. [0099]: tasks may include AI/ML processing (e.g., including training, inferencing, and classification operations), visual data processing, network data processing, object detection, rule analysis, or the like, [0008]: The present disclosure is related to techniques for optimizing artificial intelligence (AI) and/or machine learning (ML) models to reduce resource consumption while improving AI/ML model performance. In particular, the present disclosure provides a ML architecture search (MLAS) framework that involves generalized and/or hardware (HW)-aware ML architectures. [0019] For purposes of the present disclosure, the term “ML architecture” may refer to a particular ML model having a particular set of ML parameters and/or such an ML model configured to be operated on a particular HW platform [0017]: FIG. 1 shows the main components and interactions of an ML model scaling (MLMS) system 100. MLMS system 100 provides a holistic system that selects and interchanges ML models in an energy and communication efficient way while at the same time adapting the models to real-time changes in system constraints, conditions, and/or configurations. - super-net hardware configured to implement a super-net including a plurality of selectable sub-networks, the super-net including network weights and architecture parameters, [0019]: For purposes of the present disclosure, the term “ML architecture” may refer to a particular ML model having a particular set of ML parameters and/or such an ML model configured to be operated on a particular HW platform. Here, the set of ML parameters may refer to “model parameters” (also referred to simply as “parameters”) and/or “hyperparameters.” Model parameters are parameters derived via training, whereas hyperparameters are parameters whose values are used to control aspects of the learning process and usually have to be set before running an ML model (e.g., weights, etc.). Additionally, for purposes of the present disclosure, hyperparameters may be classified as architectural hyperparameters or training hyperparameters. Architectural hyperparameters are hyperparameters that are related to architectural aspects of an ML model such as, for example, the number of layers in a DNN, specific layer types in a DNN (e.g., convolutional layers, multilayer perception (MLP) layers, etc.), number of output channels, kernel size, and/or the like. [0022]: The supernet 130 may be an ML model configured and/or trained for a particular AI/ML task from which the MLMS system 100 is to discover or generate a smaller and/or derivative ML model (referred to herein as a “sub-network” or “subnet”). [0022]: This supernet 130 may include parameters and/or weights that do not significantly contribute to the prediction and/or inference determination, and these parameters and/or weights contribute to the supernet's overall computational complexity and density. Therefore, the supernet 130 contains one or more smaller subnets (e.g., the subnets 135 in the subnet pool 147 in FIG. 1) that, when trained in isolation, can match and/or offer trade-offs in various objectives and/or performance metrics such as accuracy and/or latency of the original supernet 130 when trained for the same number of iterations or epochs. - where the architecture parameters represent the importance of different architecture choices at various locations inside super-net, where the super-net hardware includes a plurality of nodes coupled by operational blocks, an operational block including two or more neural networks each having a plurality of network weights, [0178]: The term “attention” in the context of machine learning and/or neural networks, at least in some embodiments refers to a technique that mimics cognitive attention, which enhances important parts of a dataset where the important parts of the dataset may be determined using training data by gradient descent - and where different neural networks in the same operational block share at least some of the same network weights; [0014] : The present disclosure provides an MLMS system that selects and interchanges ML models (e.g., subnets) in an energy and communication efficient way while adapting the ML models (e.g., subnets) to real time (or near-real time) changes in system (HW platform) constraints [BRI: A MLMS system is an operational block] [0014] : The MLMS system discussed herein is agnostic to particular approach used for ML model search (e.g., NAS) and for ML model (e.g., supernet) weight sharing. The MLMS system implements a unique and robust ML model search and swap strategy that can flexibly adapt ML models for a wide variety of detected compute system operational changes and/or environmental changes. [0058]: FIG. 3 shows an example of the efficiency-aware subnet model selection after the subnet search is performed, according to various embodiments. The supernet search results 301 shows all discovered subnets 335 found during a subnet search, which includes the existing supernet 330 and a set of candidate replacement subnets 340 (individual subnets in the set of candidate replacement subnets 340 may be referred to herein as “replacement subnets 340”, “candidate subnets 340”, or the like). [BRI: candidate subnets represents sample of subnets] - a data loader, coupled to the super-net hardware, configured to access training data for a designated task and provide training inputs of the training data to sample sub-networks of the super-net to generate sub-network outputs, [0014]: The MLMS system implements a unique and robust ML model search and swap strategy that can flexibly adapt ML models for a wide variety of detected compute system operational changes and/or environmental changes. [0015]: the MLMS system discussed herein provides a relatively low dynamic power signature (e.g., during ML model swapping) and a relatively fast memory write time and/or model loading time during model interchange operations. [0021]: In alternative implementations, the MLMS system 100 may be operated by the same compute node (e.g., client device, edge compute node, network access node, cloud service, drone, network appliance, etc.)”, [0033]: In some implementations, the HW platform may be the client device 101. In some implementations, the HW platform may be some other desired compute node or device on which the user wishes to deploy an ML model or wishes to perform some AI/ML task”, [0027]: The AI/ML tasks may describe a desired problem to be solved and the AI/ML domain may describe a desired goal to be achieved. Examples of ML tasks include clustering, classification, regression, anomaly detection, data cleaning, automated ML (autoML), association rules learning, reinforcement learning, structured prediction, feature engineering, feature learning, online learning, supervised learning, semi-supervised learning (SSL), unsupervised learning, machine learned ranking (MLR), grammar induction, and/or the like. - where a sample sub-network includes a neural network selected from each operational block; [0017]: FIG. 1 shows the main components and interactions of an ML model scaling (MLMS) system 100. MLMS system 100 provides a holistic system that selects and interchanges ML models in an energy and communication efficient way while at the same time adapting the models to real-time changes in system constraints, conditions, and/or configurations - a supervised learning controller, coupled to the super-net hardware, configured to train network weights and architecture parameters of the super-net, including: [0027]: The AI/ML tasks may describe a desired problem to be solved and the AI/ML domain may describe a desired goal to be achieved. Examples of ML tasks include clustering, classification, regression, anomaly detection, data cleaning, automated ML (autoML), association rules learning, reinforcement learning, structured prediction, feature engineering, feature learning, online learning, supervised learning, semi-supervised learning (SSL), unsupervised learning, machine learned ranking (MLR), grammar induction, and/or the like. [0100]: an AI engine chip that can run many different kinds of AI instruction sets once loaded with the appropriate weightings and training code - accumulate a loss over the sample of sub-networks, the accumulated loss based, at least in part, on a sum, over layers of sub-networks in the sample, [0059]: In embodiments, the subnet selector 141 implements a search function and/or NAS algorithm to identify as many possible subnet candidates 340 for the new context information (e.g., new HW configuration) that precipitated the search. [0209]: The term “loss function” or “cost function” at least in some embodiments refers to an event or values of one or more variables onto a real number that represents some “cost” associated with the event. A value calculated by a loss function may be referred to as a “loss” or “error”. Additionally or alternatively, the term “loss function” or “cost function” at least in some embodiments refers to a function used to determine the error or loss between the output of an algorithm and a target value. Additionally or alternatively, the term “loss function” or “cost function” at least in some embodiments refers to a function are used in optimization problems with the goal of minimizing a loss or error. [0059]: Of the possible subnet candidates 340 that are found, preference is given to a subnet 340 in the set of candidates 340 that is the most related to the existing subnet 135. The determination of the most closely related subnet 340 is based on the supernet type weight/parameter sharing scheme, such as the vectorization scheme discussed previously with respect to FIG. 2, which can be used to calculate the total transfer cost for each of the candidates 340. The subnet 340 with a minimal total transfer cost (e.g., with respect to the other candidate subnets 340) is selected as a new active subnet. [BRI: The transfer cost is accumulated loss] - adjust the architecture parameters and network weights of the super-network to reduce the accumulated loss, [0214]: The term “objective function” at least in some embodiments refers to a function to be maximized or minimized for a specific optimization problem. In some cases, an objective function is defined by its decision variables and an objective. The objective is the value, target, or goal to be optimized, such as maximizing profit or minimizing usage of a particular resource. The specific objective function chosen depends on the specific problem to be solved and the objectives to be optimized. [0209]: the term “loss function” or “cost function” at least in some embodiments refers to a function are used in optimization problems with the goal of minimizing a loss or error. [0059]: Of the possible subnet candidates 340 that are found, preference is given to a subnet 340 in the set of candidates 340 that is the most related to the existing subnet 135. The determination of the most closely related subnet 340 is based on the supernet type weight/parameter sharing scheme, such as the vectorization scheme discussed previously with respect to FIG. 2, which can be used to calculate the total transfer cost for each of the candidates 340. The subnet 340 with a minimal total transfer cost (e.g., with respect to the other candidate subnets 340) is selected as a new active subnet. - where the network weights shared between different neural networks of an operational block are trained jointly, [0014]: The present disclosure provides an MLMS system that selects and interchanges ML models (e.g., subnets) in an energy and communication efficient way while adapting the ML models (e.g., subnets) to real time (or near-real time) changes in system (HW platform) constraints. The MLMS system discussed herein is agnostic to particular approach used for ML model search (e.g., NAS) and for ML model (e.g., supernet) weight sharing. The MLMS system implements a unique and robust ML model search and swap strategy that can flexibly adapt ML models for a wide variety of detected compute system operational changes and/or environmental changes. [0020]: the MLMSI 110a includes application programming interfaces (APIs) to access the other subsystems of system 100, manages ML search parameter updates (e.g., new or updated ML config), and calls any supported ML operations library (e.g., as indicated by the ML config). [0209]: the term “loss function” or “cost function” at least in some embodiments refers to a function are used in optimization problems with the goal of minimizing a loss or error. [0059]: Of the possible subnet candidates 340 that are found, preference is given to a subnet 340 in the set of candidates 340 that is the most related to the existing subnet 135. The determination of the most closely related subnet 340 is based on the supernet type weight/parameter sharing scheme, such as the vectorization scheme discussed previously with respect to FIG. 2, which can be used to calculate the total transfer cost for each of the candidates 340. The subnet 340 with a minimal total transfer cost (e.g., with respect to the other candidate subnets 340) is selected as a new active subnet. [0053]: a subnet 135 that has the most overlap in pre-trained parameters from the existing subnet 135 for a target device (e.g., compute node 550). These mechanisms help the target device adapt to a new system states and/or HW conditions/constraints, but also limit the data-transfer resource consumption when writing the new subnet 135 to memory. [BRI: Within the context of objective function being optimized, and the the total transfer cost is combination of first and second loss as a result of training all components (subnet) together is indeed a form of joint training] - where the supervised learning controller is further configured to: select a sub-network of the plurality of sub-networks based on the largest adjusted architecture parameters; [0017]: The MLMS system 100 achieves energy and communication efficiency by using a similarity-based subnet selection process where a subnet is selected from a pool of subnets that has the most overlap in pre-trained parameters from the existing subnet to minimize memory write operation overhead [0019]: Architectural hyperparameters are hyperparameters that are related to architectural aspects of an ML model such as, for example, the number of layers in a DNN, specific layer types in a DNN (e.g., convolutional layers, multilayer perception (MLP) layers, etc.), number of output channels, kernel size, and/or the like. [0022]: “elasticity” refers to a component in the supernet 130 and its corresponding operator can change in several dimensions. For example, in convolutional networks this would allow for the choice of different kernel sizes (e.g., elastic kernel), depths of some group of elements (e.g., elastic depth), and number of channels of the selected components (e.g., elastic width). In the context of the present disclosure, the term “elastic” refers to the ability of an elastic supernet 130 to add and/or remove different ML parameters and/or subnets 135 without restarting the running framework and/or the running ML model. [BRI: Perhaps known to a POISTA that adjusted architecture parameters refer to the continuous or discrete variables that define a candidate network’s structure, which are optimized during the search process. These parameters can include: filter(kernel), number of layers, and number of channels. - and output a description of the network architecture of the selected sub-network for the chosen task. [0010]” optimizing an ML model for individual HW platforms and specific performance metrics is a very time-consuming effort, which requires highly specialized knowledge. This type of optimization is typically done manually with a great deal of in-depth understanding of the HW platform since certain characteristics of the HW platform (e.g., clock speed, number of processor cores, amount of cache memory, etc.) will affect the optimization process. The optimization process is also affected by the characteristics of the input data to the ML model (e.g., batch size, image size, number of epochs/iterations, etc.). Finally, any change to the performance metrics (e.g., going from latency to power consumption), input data characteristics (e.g., increasing the batch size), [0091]: To provide the inference, the inference engine 516 uses a model 520 that controls how the DNN inference is made on the data 514 to generate the result 518. Specifically, the model 520 includes a topology of layers of a NN. The topology includes an input layer that receives the data 514, an output layer that outputs the result 518, and one or more hidden layers between the input and output layers that provide processing between the data 14 and the result 518. The topology may be stored in a suitable information object, such as an extensible markup language (XML), JavaScript Object Notation (JSON), and/or other suitable data structure, file, and/or the like. The model 520 may also include weights and/or biases for results for any of the layers while processing the data 514 in the inference using the DNN. Cummings do not explicitly disclose: - accumulate a loss over the sample of sub-networks, the accumulated loss based, at least in part on a sum, over layers of sub-networks in the sample of measures of smoothness based on network weights in the layers, where a measure of smoothness relates to the maximum change in output from a layer relative to any possible change in input to the layer, - a higher maximum change in output indicating a lower smoothness; - where the accumulated loss penalizes sub- networks with lower smoothness and improves stability of the adjustment; However, Badran discloses: - accumulate a loss over the sample of sub-networks, the accumulated loss based, at least in part on a sum, over layers of sub-networks in the sample of measures of smoothness based on network weights in the layers, where a measure of smoothness relates to the maximum change in output from a layer relative to any possible change in input to the layer [2, Page 1446]: The neural network smoother uses a basic MLP architecture (R) with an input layer supporting the temporal window, l layers of non-linear hidden neurons and one linear output neuron. Starting from this basic architecture (R), a family of multi-layer perceptrons R λ is generated. For each λ ,   the synaptic weights vector is chosen using the learning procedure described below. [2, 1446]: The input layer uses a temporal window on the time series, which is the vector of observations ( y i - d ' , … y i + d ), d’ ≥ d The two time lags d and d’ can be equal or different. They can be both positive, in which case, d’ represents the amount of information in the past used by the temporal window; d stands for the future and y i   is within the window. However, d may be negative: for example, if d= -1 and y i is not within the window, smoothing at time i uses only past information, and it can be considered as a predictor of the smoothed value at time i. [2, Page 1446]: A classical smoothing method is to choose from the family of twice continuously differentiable functions S, such that PNG media_image1.png 28 107 media_image1.png Greyscale exists, the function S which minimizes the measure of roughness PNG media_image2.png 56 433 media_image2.png Greyscale The weighting parameter λ   controls the amount of smoothing, so that increasing λ   leads to smoother estimate. [BRI: The increase λ to minimizing measure of roughness]. [3, Page 1447]: The MASE can be easily expressed as PNG media_image3.png 166 511 media_image3.png Greyscale The first term on the right-hand side of (6) corresponds to a variance component and the second to the square of a bias component, both of them contributing to the MASE(   g λ ) The variance component is decreased by smoothing, but this increases the bias component. The problem is thus to choose a good compromise between these two components (which is done by minimizing (6) with respect to PNG media_image4.png 12 1 media_image4.png Greyscale [BRI: the bias component measures the average difference between the model’s predictions and the best possible model for the problem, averaged over all training data instances and increasing it underfits the data and affects smoothness as represented by equation (6)] - a higher maximum change in output indicating a lower smoothness; [3, Page 1447]: The MASE can be easily expressed as PNG media_image3.png 166 511 media_image3.png Greyscale The first term on the right-hand side of (6) corresponds to a variance component and the second to the square of a bias component, both of them contributing to the MASE(   g λ ) The variance component is decreased by smoothing, but this increases the bias component. The problem is thus to choose a good compromise between these two components (which is done by minimizing (6) with respect to PNG media_image4.png 12 1 media_image4.png Greyscale [BRI: in effect, the high difference resulted from the bias component does increases MASE and thus lower the smoothness] - the accumulated loss based, at least in part, on a sum, over layers of a sub-network, of the measures of smoothness of the layers; [2, Page 1446]: The neural network smoother uses a basic MLP architecture (R) with an input layer supporting the temporal window, l layers of non-linear hidden neurons and one linear output neuron. Starting from this basic architecture (R), a family of multi-layer perceptrons (Rl) is generated. For each l, the synaptic weights vector is chosen using the learning procedure described below. [2, 1446]: The input layer uses a temporal window on the time series, which is the vector of observations ( y i - d ' , … y i + d ), d’ ≥ d The two time lags d and d’ can be equal or different. They can be both positive, in which case, d’ represents the amount of information in the past used by the temporal window; d stands for the future and y i   is within the window. However, d may be negative: for example, if d= -1 and y i is not within the window, smoothing at time i uses only past information, and it can be considered as a predictor of the smoothed value at time i. [2, Page 1446]: A classical smoothing method is to choose from the family of twice continuously differentiable functions S, such that PNG media_image1.png 28 107 media_image1.png Greyscale exists, the function S which minimizes the measure of roughness [BRI: function S that minimizes the measure of roughness provides the measure of smoothness] - where the second loss penalizes sub- networks with lower smoothness and improves stability of the adjustment; [Abstract, Page 1445]: The NN-smoother computes the trend in the state domain and minimizes a cost function with a regularization term. The regularization term is penalized by a parameter λ which forces the learning procedure to smooth in the time domain. We define a selecting criterion in order to select the best parameter λ . [1, Page 1445]: In most cases, smoothing techniques generate families of functions g λ .     d e pending on a parameter λ , which controls the degree of smoothing. Each of these functions is a solution to the minimization of a cost function, containing a regularization term parameterized by λ . [3, Page 1447]: The MASE can be easily expressed as PNG media_image3.png 166 511 media_image3.png Greyscale The first term on the right-hand side of (6) corresponds to a variance component and the second to the square of a bias component, both of them contributing to the MASE(   g λ ). The variance component is decreased by smoothing, but this increases the bias component. The problem is thus to choose a good compromise between these two components (which is done by minimizing (6) with respect to PNG media_image4.png 12 1 media_image4.png Greyscale [BRI: in effect, the high difference resulted from the bias component does increases MASE and thus lower the smoothness and is solved by minimizing MASE by using compilation since the regularization (penalty) is parameterized by λ ] [1, Page 1445]: Neural networks can be used for smoothing: for example, an MLP, trained through a modified gradient back propagation, can implement spline smoothing in the time domain [1, Page 1446]: The neural network smoother uses a basic MLP architecture (R) with an input layer supporting the temporal window, l layers of non-linear hidden neurons and one linear output neuron. Starting from this basic architecture (R), a family of multi-layer perceptrons (Rl) is generated. For each λ , the synaptic weights vector is chosen [BRI: in the time domain, spline smoothing with a large smoothing parameter λ does represent small changes in output with respect to input, because it penalizes curvature and enforces a smooth, slowly varying curve. In fact, this is exactly what Lipschitz constant provides (small changes to output- see spec [0022]) that minimize the loss jointly for all sub-networks] [2, Page 1447]:] We consider the measure of roughness PNG media_image5.png 85 480 media_image5.png Greyscale [2, Page 1447]: The learning phase consists in minimizing the discrete approximation given in (4) by the back propagation algorithm. [4, Page 1450]: NN performs a smoothing in the state space using a cost function C(l) (4) which acts in the time domain. So, the NN smoother combines two representations: the state space domain and the time domain. The NN network directly minimizes (4) which is the discrete form of (2) and gives the optimal smoothing directly for each discrete time unit i [BRI: In the state-space domain: the internal dynamics of the system, modeled as a set of state equations (continuous or discrete) and in the time-domain data: the actual measured inputs, outputs, and (often) states over time, used to estimate the network weights and biases. Several sub-networks are part of the same state-space model. The same set of weights (shared) is updated during training to ensure consistency across the block and the joint training leverages all available time-domain data from the block to optimize shared parameter. The NN directly minimize the cost function(loss function) given by (4)] [2, Page 1447]: The input layer has d þ d þ 3 units (for window [yi¹1¹d,…,yiþ1þd]), l hidden layers (l 0) and 3 output units. The first hidden layer has three clusters of n1 units each. The first cluster of hidden units is connected to window {i ¹ 1 ¹ d,…,i ¹ 1 þ d}, the second to the window shifted by one time unit, and the third to the window shifted by one more unit; these clusters share the same weight vector. The use of time-delay inputs convert the time series into a static mapping problem. The only constraint is that each hidden layer has three clusters of units connected only to the corresponding cluster in the previous layer. In each layer, clusters share the same weight vector. It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Cummings, and Badran. Cummings teaches a super-net including a plurality of sub-networks, first loss over a sample of sub-networks, and second loss and adjusting network weights to reduce the accumulated loss. Badran teaches relating the measure of smoothness to the maximum difference between outputs from the layer relative to any possible difference between inputs to the layer. One of ordinary skill would be motivated to combine Cummings, and Badran that choose Claim 16: (Previously Presented) Cummings discloses: - train network weights of the selected sub-network using the training data; [0022]: supernet 130 may include parameters and/or weights that do not significantly contribute to the prediction and/or inference determination, and these parameters and/or weights contribute to the supernet's overall computational complexity and density. Therefore, the supernet 130 contains one or more smaller subnets (e.g., the subnets 135 in the subnet pool 147 in FIG. 1) that, when trained in isolation, can match and/or offer trade-offs in various objectives and/or performance metrics such as accuracy and/or latency of the original supernet 130 when trained for the same number of iterations or epochs. Claims 3-4, 7, 11-13, 17-19 and 21 are rejected under 35 U.S.C. 103 as being unpatentable over Daniel Cummings et.al. (hereinafter Cummings ) US 2022/0036123 A1, In view F.Badran et.al. (hereinafter Badran) Neural Network Smoothing in Correlated Time Series Context, Neural Networks, Vol. 10, No. 8, pp. 1445–1453, 1997. further in view of Jisoo Mok et.al. (hereinafter Mok) AdvRush: Searching for adversarially robust neural architectures, Proceedings of the IEEE/CVF international conference on computer vision. 2021. Claim 3: (Previously Presented) Cummings, and Badran do not explicitly disclose: - and the second loss is a product of a smoothness scale factor and the sum, over layers of the sample sub-network, of measure of smoothness of the layers However, Mok discloses: - and the second loss is a product of a smoothness scale factor and the sum, over layers of the sample sub-network, of measures of smoothness of the layers [2.1, Page 12323]: Our work is closely related to the defense approaches that utilize a regularization term derived from the curvature information of the neural network’s loss landscape [6.1. Page 12327]: Effect of Regularization Strength The regularization strength γ is empirically set to be 0.01 to match the scale of L v a l and L λ . [BRI: The regularization strength is a smoothness scale factor as a result of its matching the scale] [4.2 , Page 12326]: H s t d   F can be expressed in terms of l2 norm: PNG media_image6.png 6 156 media_image6.png Greyscale where the expectation is taken over z ∼ N(0, Id). Because the direct computation of H s t d is expensive, we linearly approximate it through the finite difference approximation of the Hessian: PNG media_image7.png 77 422 media_image7.png Greyscale where h controls the scale of the loss landscape on which we induce smoothness. However, computing multiple H s t d z in directions drawn from z ∼ N(0, I d ) and taking its average would be computationally inefficient because each computation of Hstdz requires calculation of the gradient. Therefore, we minimize the input loss landscape along the the high curvature direction, PNG media_image8.png 25 178 media_image8.png Greyscale to maximize the effect of L λ , . With the approximated L λ , the bi-level optimization problem of AdvRush can be expressed as: PNG media_image9.png 121 472 media_image9.png Greyscale x v a l   is the clean input data from Dval. The value of h in the denominator of Eq. (9) is absorbed by the regularization strength γ. It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Cummings, Badran and Mok. Cummings teaches a super-net including a plurality of sub-networks, first loss over a sample of sub-networks, and second loss and adjusting network weights to reduce the accumulated Badran teaches relating the measure of smoothness to the maximum difference between outputs from the layer relative to any possible difference between inputs to the layer Mok teaches relating the measure of smoothness to the maximum difference between outputs from the layer relative to any possible difference between inputs to the layer. One of ordinary skill would be motivated to combine Cummings, Badran and Mok that can place a proper choice of an architectures to improve the performance of the neural network on a target task (Mok [ 1, Page 12323]). Claim 4: (Original) Cummings, and Badran do not explicitly disclose: - the second loss for sub-network k is: PNG media_image10.png 47 318 media_image10.png Greyscale λ   is a smoothness scale factor, W k ,   j   is a matrix of network weights in layer j of sub-network k and σ m a x   ( W k ,   j     ) is the maximum singular of the matrix W k ,   j   or an estimate thereof. However, Mok discloses: - the second loss for sub-network k is: PNG media_image10.png 47 318 media_image10.png Greyscale λ   is a smoothness scale factor, W k ,   j   is a matrix of network weights in layer j of sub-network k and σ m a x   ( W k ,   j     ) is the maximum singular of the matrix W k ,   j   or an estimate thereof. [4.2, Page 12326]: Approximation of L λ H s t d   F can be expressed in terms of l2 norm: PNG media_image6.png 6 156 media_image6.png Greyscale where the expectation is taken over z ∼ N(0, I d ). Because the direct computation of H s t d is expensive, we linearly approximate it through the finite difference approximation of the Hessian: PNG media_image7.png 77 422 media_image7.png Greyscale where h controls the scale of the loss landscape on which we induce smoothness. However, computing multiple H s t d z in directions drawn from z ∼ N(0, I d ) and taking its average would be computationally inefficient because each computation of H s t d z requires calculation of the gradient. Therefore, we minimize the input loss landscape along the high curvature direction, PNG media_image11.png 27 182 media_image11.png Greyscale to maximize the effect of [4.2, Page 12326]: Approximation of L λ H s t d   F can be expressed in terms of l2 norm: PNG media_image6.png 6 156 media_image6.png Greyscale where the expectation is taken over z ∼ N(0, I d ). Because the direct computation of H s t d is expensive, we linearly approximate it through the finite difference approximation of the Hessian: PNG media_image7.png 77 422 media_image7.png Greyscale where h controls the scale of the loss landscape on which we induce smoothness. However, computing multiple H s t d z in directions drawn from z ∼ N(0, I d ) and taking its average would be computationally inefficient because each computation of H s t d z requires calculation of the gradient. Therefore, we minimize the input loss landscape along the high curvature direction. Claim 7: (Previously Presented) Cummings, and Badran do not explicitly disclose: - where said adjust the architecture parameters and the network weights includes: determining gradients of the accumulated first and second loss functions; However, Mok discloses: - where said adjust the architecture parameters and the network weights includes: determining gradients of the accumulated first and second loss functions; [4.2, Page 12326]: Approximation of L λ H s t d   F can be expressed in terms of l2 norm: PNG media_image6.png 6 156 media_image6.png Greyscale where the expectation is taken over z ∼ N(0, I d ). Because the direct computation of H s t d is expensive, we linearly approximate it through the finite difference approximation of the Hessian: PNG media_image7.png 77 422 media_image7.png Greyscale where h controls the scale of the loss landscape on which we induce smoothness. However, computing multiple H s t d z in directions drawn from z ∼ N(0, I d ) and taking its average would be computationally inefficient because each computation of Hstdz requires calculation of the gradient. Therefore, we minimize the input loss landscape along the the high curvature direction, [4.2, Page 12326]: each computation of H s t d z requires calculation of the gradient. Therefore, we minimize the input loss landscape along the high curvature direction, PNG media_image12.png 23 176 media_image12.png Greyscale - and update the architecture parameters and the network weights of the super-net based on the gradients. [4.1, Page 12325]: To make the search space continuous for gradient-based optimization, categorical choice of a particular operation is continuously relaxed by applying a softmax function over all the possible operations: PNG media_image13.png 63 442 media_image13.png Greyscale where α   ( i , j ) is a set of operation mixing weights (i.e., architecture parameters). O is the pre-defined set of operations that are used to construct the supernet. By definition, the size of α   ( i , j )   must be equal to |O|. Through continuous relaxation, both the architecture parameters α and the weight parameters ω in the supernet can be updated via gradient descent. It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Cummings, Badran and Mok. Cummings teaches a super-net including a plurality of sub-networks, first loss over a sample of sub-networks, and second loss and adjusting network weights to reduce the accumulated Badran teaches relating the measure of smoothness to the maximum difference between outputs from the layer relative to any possible difference between inputs to the layer Mok teaches relating the measure of smoothness to the maximum difference between outputs from the layer relative to any possible difference between inputs to the layer. One of ordinary skill would be motivated to combine Cummings, Badran and Mok that can place a proper choice of an architectures to improve the performance of the neural network on a target task (Mok [ 1, Page 12323]). Claim 11: (Previously Presented) Cummings, and Badran do not explicitly disclose: - and the second loss is a product of a smoothness scale factor and the sum, over layers of the sample sub-network, of measure of smoothness of the layers However, Mok discloses: - and the second loss is a product of a smoothness scale factor and the sum, over layers of the sample sub-network, of measures of smoothness of the layers [2.1, Page 12323]: Our work is closely related to the defense approaches that utilize a regularization term derived from the curvature information of the neural network’s loss landscape [6.1. Page 12327]: Effect of Regularization Strength The regularization strength γ is empirically set to be 0.01 to match the scale of L v a l and L λ . (BRI: The regularization strength is a smoothness scale factor as a result of its matching the scale) [4.2 , Page 12326]: H s t d   F can be expressed in terms of l2 norm: PNG media_image6.png 6 156 media_image6.png Greyscale where the expectation is taken over z ∼ N(0, Id). Because the direct computation of H s t d is expensive, we linearly approximate it through the finite difference approximation of the Hessian: PNG media_image7.png 77 422 media_image7.png Greyscale where h controls the scale of the loss landscape on which we induce smoothness. However, computing multiple H s t d z in directions drawn from z ∼ N(0, I d ) and taking its average would be computationally inefficient because each computation of Hstdz requires calculation of the gradient. Therefore, we minimize the input loss landscape along the the high curvature direction, PNG media_image8.png 25 178 media_image8.png Greyscale to maximize the effect of L λ , . With the approximated L λ , the bi-level optimization problem of AdvRush can be expressed as: PNG media_image9.png 121 472 media_image9.png Greyscale x v a l   is the clean input data from Dval. The value of h in the denominator of Eq. (9) is absorbed by the regularization strength γ. It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Cummings, Badran and Mok. Cummings teaches a super-net including a plurality of sub-networks, first loss over a sample of sub-networks, and second loss and adjusting network weights to reduce the accumulated Badran teaches relating the measure of smoothness to the maximum difference between outputs from the layer relative to any possible difference between inputs to the layer Mok teaches relating the measure of smoothness to the maximum difference between outputs from the layer relative to any possible difference between inputs to the layer. One of ordinary skill would be motivated to combine Cummings, Badran and Mok that can place a proper choice of an architectures to improve the performance of the neural network on a target task (Mok [ 1, Page 12323]). Claim 12: (Original) Cummings, and Badran do not explicitly disclose: PNG media_image10.png 47 318 media_image10.png Greyscale λ   is a smoothness scale factor, W k ,   j   is a matrix of network weights in layer j of sub-network k and σ m a x   ( W k ,   j     ) is the maximum singular of the matrix W k ,   j   or an estimate thereof. However, Mok discloses: PNG media_image10.png 47 318 media_image10.png Greyscale λ   is a smoothness scale factor, W k ,   j   is a matrix of network weights in layer j of sub-network k and σ m a x   ( W k ,   j     ) is the maximum singular of the matrix W k ,   j   or an estimate thereof. [4.2, Page 12326]: Approximation of L λ H s t d   F can be expressed in terms of l2 norm: PNG media_image6.png 6 156 media_image6.png Greyscale where the expectation is taken over z ∼ N(0, I d ). Because the direct computation of H s t d is expensive, we linearly approximate it through the finite difference approximation of the Hessian: PNG media_image7.png 77 422 media_image7.png Greyscale where h controls the scale of the loss landscape on which we induce smoothness. However, computing multiple H s t d z in directions drawn from z ∼ N(0, I d ) and taking its average would be computationally inefficient because each computation of H s t d z requires calculation of the gradient. Therefore, we minimize the input loss landscape along the the high curvature direction, PNG media_image11.png 27 182 media_image11.png Greyscale to maximize the effect of [4.2, Page 12326]: Approximation of L λ H s t d   F can be expressed in terms of l2 norm: PNG media_image6.png 6 156 media_image6.png Greyscale where the expectation is taken over z ∼ N(0, I d ). Because the direct computation of H s t d is expensive, we linearly approximate it through the finite difference approximation of the Hessian: PNG media_image7.png 77 422 media_image7.png Greyscale where h controls the scale of the loss landscape on which we induce smoothness. However, computing multiple H s t d z in directions drawn from z ∼ N(0, I d ) and taking its average would be computationally inefficient because each computation of H s t d z requires calculation of the gradient. Therefore, we minimize the input loss landscape along the high curvature direction. Claim 13: (Previously Presented) Cummings, and Badran do not explicitly disclose: - where said adjust the architecture parameters and the network weights includes: determining gradients of the accumulated first and second loss functions; However, Mok discloses: - where said adjust the architecture parameters and the network weights includes: determining gradients of the accumulated first and second loss functions; [4.2, Page 12326]: Approximation of L λ H s t d   F can be expressed in terms of l2 norm: PNG media_image6.png 6 156 media_image6.png Greyscale where the expectation is taken over z ∼ N(0, I d ). Because the direct computation of H s t d is expensive, we linearly approximate it through the finite difference approximation of the Hessian: PNG media_image7.png 77 422 media_image7.png Greyscale where h controls the scale of the loss landscape on which we induce smoothness. However, computing multiple H s t d z in directions drawn from z ∼ N(0, I d ) and taking its average would be computationally inefficient because each computation of H s t d z requires calculation of the gradient. Therefore, we minimize the input loss landscape along the the high curvature direction, [4.2, Page 12326]: each computation of H s t d z requires calculation of the gradient. Therefore, we minimize the input loss landscape along the high curvature direction, PNG media_image12.png 23 176 media_image12.png Greyscale - and update the architecture parameters and the network weights of the super-net based on the gradients. [4.1, Page 12325]: To make the search space continuous for gradient-based optimization, categorical choice of a particular operation is continuously relaxed by applying a softmax function over all the possible operations: PNG media_image13.png 63 442 media_image13.png Greyscale where α   ( i , j ) is a set of operation mixing weights (i.e., architecture parameters). O is the pre-defined set of operations that are used to construct the supernet. By definition, the size of α   ( i , j )   must be equal to |O|. Through continuous relaxation, both the architecture parameters α and the weight parameters ω in the supernet can be updated via gradient descent. [0052]: FIG. 39 is a system diagram for an example system for training, adapting, instantiating and deploying machine learning models in an advanced computing pipeline, in accordance with at least one embodiment; [0005]: FIG. 3 illustrates a diagram of an overall framework on how a neural network is selected for an input at a federated learning (FL) client site, according to at least one embodiment; [0005]: FIG. 3 illustrates a diagram of an overall framework on how a neural network is selected for an input at a federated learning (FL) client site, according to at least one embodiment; - store the trained network weights. [0064]: in at least one embodiment, results 114, 120 include updated model weights (or their gradients) from trained portion of supernet for client A 112 and trained portion of supernet for client B 118, and updated model weights are sent to client server 102 for aggregation. In at least one embodiment, after aggregation, new weights are redistributed to client A 106 and client B 108 and a next round of local training is executed. Claim 17: (Original) Cummings, and Badran do not explicitly disclose: - and the second loss is a product of a smoothness scale factor and the sum, over layers of the sample sub-network, of measure of smoothness of the layers However, Mok discloses: - and the second loss is a product of a smoothness scale factor and the sum, over layers of the sample sub-network, of measures of smoothness of the layers [2.1, Page 12323]: Our work is closely related to the defense approaches that utilize a regularization term derived from the curvature information of the neural network’s loss landscape [6.1. Page 12327]: Effect of Regularization Strength The regularization strength γ is empirically set to be 0.01 to match the scale of L v a l and L λ . (BRI: The regularization strength is a smoothness scale factor as a result of its matching the scale) [4.2 , Page 12326]: H s t d   F can be expressed in terms of l2 norm: PNG media_image6.png 6 156 media_image6.png Greyscale where the expectation is taken over z ∼ N(0, Id). Because the direct computation of H s t d is expensive, we linearly approximate it through the finite difference approximation of the Hessian: PNG media_image7.png 77 422 media_image7.png Greyscale where h controls the scale of the loss landscape on which we induce smoothness. However, computing multiple H s t d z in directions drawn from z ∼ N(0, I d ) and taking its average would be computationally inefficient because each computation of Hstdz requires calculation of the gradient. Therefore, we minimize the input loss landscape along the the high curvature direction, PNG media_image8.png 25 178 media_image8.png Greyscale to maximize the effect of L λ , . With the approximated L λ , the bi-level optimization problem of AdvRush can be expressed as: PNG media_image9.png 121 472 media_image9.png Greyscale x v a l   is the clean input data from Dval. The value of h in the denominator of Eq. (9) is absorbed by the regularization strength γ. It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Cummings, Badran and Mok. Cummings teaches a super-net including a plurality of sub-networks, first loss over a sample of sub-networks, and second loss and adjusting network weights to reduce the accumulated Badran teaches relating the measure of smoothness to the maximum difference between outputs from the layer relative to any possible difference between inputs to the layer Mok teaches relating the measure of smoothness to the maximum difference between outputs from the layer relative to any possible difference between inputs to the layer. One of ordinary skill would be motivated to combine Cummings, Badran and Mok that can place a proper choice of an architectures to improve the performance of the neural network on a target task (Mok [ 1, Page 12323]). Claim 18: (Original) Cummings, and Badran do not explicitly disclose: PNG media_image10.png 47 318 media_image10.png Greyscale λ   is a smoothness scale factor, W k ,   j   is a matrix of network weights in layer j of sub-network k and σ m a x   ( W k ,   j     ) is the maximum singular of the matrix W k ,   j   or an estimate thereof. However, Mok discloses: - PNG media_image10.png 47 318 media_image10.png Greyscale   λ   is a smoothness scale factor, W k ,   j   is a matrix of network weights in layer j of sub-network k and σ m a x   ( W k ,   j     ) is the maximum singular of the matrix W k ,   j   or an estimate thereof. [4.2, Page 12326]: Approximation of L λ H s t d   F can be expressed in terms of l2 norm: PNG media_image6.png 6 156 media_image6.png Greyscale where the expectation is taken over z ∼ N(0, I d ). Because the direct computation of H s t d is expensive, we linearly approximate it through the finite difference approximation of the Hessian: PNG media_image7.png 77 422 media_image7.png Greyscale where h controls the scale of the loss landscape on which we induce smoothness. However, computing multiple H s t d z in directions drawn from z ∼ N(0, I d ) and taking its average would be computationally inefficient because each computation of H s t d z requires calculation of the gradient. Therefore, we minimize the input loss landscape along the the high curvature direction, PNG media_image11.png 27 182 media_image11.png Greyscale to maximize the effect of [4.2, Page 12326]: Approximation of L λ H s t d   F can be expressed in terms of l2 norm: PNG media_image6.png 6 156 media_image6.png Greyscale where the expectation is taken over z ∼ N(0, I d ). Because the direct computation of H s t d is expensive, we linearly approximate it through the finite difference approximation of the Hessian: PNG media_image7.png 77 422 media_image7.png Greyscale where h controls the scale of the loss landscape on which we induce smoothness. However, computing multiple H s t d z in directions drawn from z ∼ N(0, I d ) and taking its average would be computationally inefficient because each computation of H s t d z requires calculation of the gradient. Therefore, we minimize the input loss landscape along the high curvature direction. Claim 19:(Previously Presented) Cummings, and Badran do not explicitly disclose: - where said adjust the architecture parameters and the network weights includes: determining gradients of the accumulated first and second loss functions; However, Mok discloses: - where said adjust the architecture parameters and the network weights includes: determining gradients of the accumulated first and second loss functions; [4.2, Page 12326]: Approximation of L λ H s t d   F can be expressed in terms of l2 norm: PNG media_image6.png 6 156 media_image6.png Greyscale where the expectation is taken over z ∼ N(0, I d ). Because the direct computation of H s t d is expensive, we linearly approximate it through the finite difference approximation of the Hessian: PNG media_image7.png 77 422 media_image7.png Greyscale where h controls the scale of the loss landscape on which we induce smoothness. However, computing multiple H s t d z in directions drawn from z ∼ N(0, I d ) and taking its average would be computationally inefficient because each computation of Hstdz requires calculation of the gradient. Therefore, we minimize the input loss landscape along the the high curvature direction, [4.2, Page 12326]: each computation of H s t d z requires calculation of the gradient. Therefore, we minimize the input loss landscape along the high curvature direction, PNG media_image12.png 23 176 media_image12.png Greyscale - and update the architecture parameters and the network weights of the super-net based on the gradients. [4.1, Page 12325]: To make the search space continuous for gradient-based optimization, categorical choice of a particular operation is continuously relaxed by applying a softmax function over all the possible operations: PNG media_image13.png 63 442 media_image13.png Greyscale where α   ( i , j ) is a set of operation mixing weights (i.e., architecture parameters). O is the pre-defined set of operations that are used to construct the supernet. By definition, the size of α   ( i , j )   must be equal to |O|. Through continuous relaxation, both the architecture parameters α and the weight parameters ω in the supernet can be updated via gradient descent. It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Cummings, Badran and Mok. Cummings teaches a super-net including a plurality of sub-networks, first loss over a sample of sub-networks, and second loss and adjusting network weights to reduce the accumulated Badran teaches relating the measure of smoothness to the maximum difference between outputs from the layer relative to any possible difference between inputs to the layer Mok teaches relating the measure of smoothness to the maximum difference between outputs from the layer relative to any possible difference between inputs to the layer. One of ordinary skill would be motivated to combine Cummings, Badran and Mok that can place a proper choice of an architectures to improve the performance of the neural network on a target task (Mok [ 1, Page 12323]). Claim 21: (Previously Presented) Cummings, and Badran do not explicitly disclose: - where the super-net has at least 100,000 parameters including the plurality of network weights and the plurality of architecture parameters. However, Mok discloses: - where the super-net has at least 100,000 parameters including the plurality of network weights and the plurality of architecture parameters. [5.2, Page 12326]: White-box Attacks We evaluate the adversarial robustness of architectures standard and adversarially trained on CIFAR-10 using various white-box attacks. [5.2, Page 12326]: White-box Attacks White-box attack evaluation results are presented in Table 1. [5.2. Page 12327]: PNG media_image14.png 380 995 media_image14.png Greyscale (BRI: The added new claim has used CIFAR-10 data set (see [0068]). See AdvRush uses 4.2 M parameters] It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Cummings, Badran and Mok. Cummings teaches a super-net including a plurality of sub-networks, first loss over a sample of sub-networks, and second loss and adjusting network weights to reduce the accumulated Badran teaches relating the measure of smoothness to the maximum difference between outputs from the layer relative to any possible difference between inputs to the layer Mok teaches relating the measure of smoothness to the maximum difference between outputs from the layer relative to any possible difference between inputs to the layer. One of ordinary skill would be motivated to combine Cummings, Badran and Mok that can place a proper choice of an architectures to improve the performance of the neural network on a target task (Mok [ 1, Page 12323]). Claims 5-6 and 22 are rejected under 35 U.S.C. 103 as being unpatentable over Daniel Cummings et.al. (hereinafter Cummings ) US 2022/0036123 A1, In view F.Badran et.al. (hereinafter Badran) Neural Network Smoothing in Correlated Time Series Context, Neural Networks, Vol. 10, No. 8, pp. 1445–1453, 1997. further in view of Minjing Dong et.al. (hereinafter Dong) arXiv:2009.00902v1 Claim 5: (Currently Amended) Cummings and Badran do not explicitly disclose: - the second loss is determined by: for each layer of [[the]]of each sub-network of the sample of sub-networks, - computationally estimating a maximum singular value of a matrix of network weights of the layers of a sub-network and over sub-networks of the sample of sub-networks - and accumulating the estimated maximum singular values over layers of a sub- network and over sub-networks of the sample of sub-networks. However, Dong discloses: - the second loss is determined by: for each layer of [[the]]of each sub-network of the sample of sub-networks, [Abstract, Page 1]: For NAS framework, all the architecture parameters are equally treated when the discrete architecture is sampled from supernet. However, the importance of architecture parameters could vary from operation to operation or connection to connection, which is not explored and might reduce the confidence of robust architecture sampling. Thus, we propose to sample architecture parameters from trainable multivariate log-normal distributions, with which the Lipschitz constant of entire network can be approximated using a univariate log-normal distribution with mean and variance related to architecture parameters [2.2, Page 3]: We now explore the relationship between architecture parameters α, β and Lipschitz constant of network. Since the entire neural network is constructed by stacking cells in series as [ I 1 , I 2 , … I N ], Eq. 2 can be further decomposed as as PNG media_image15.png 71 823 media_image15.png Greyscale where λ l , λ C and λ( I N ) denote the Lipschitz constants of loss function, classifier and cell I N   respectively. [4, Page 9]: feeding adversarial examples into the training stage to form a min-max game where the inner maximum generates adversarial samples to maximize the classification loss and outer minimum optimizes model parameters to minimize the loss. [BRI: the second loss can be determined using the Lipschitz constants of loss function] - computationally estimating a maximum singular value of a matrix of network weights of the layers of a sub-network and over sub-networks of the sample of sub-networks. [2.2, Page 2]: Lipschitz Constraints in Neural Architecture The discrete architecture A is determined by both connections and operations, which creates a huge search space. Differentiable Architecture Search algorithms provide an efficient solution through [2.2, Page 3]: the continuous relaxation of the architecture representation [18, 20, 21]. Within the differentiable NAS framework, we decompose the entire neural network into cells. Each cell I is a directed acyclic graph (DAG) consisting of an ordered sequence of n nodes, where each node denotes a latent representation which is transformed from two previous latent representations and each edge (i ,j) denotes an operation o from a pre-defined search space O which transforms I ( i ) . Following [19], the architecture parameters α which weighs operations and β which weighs input flows are introduced to form an operation mixture with weighted inputs. The intermediate node is computed as PNG media_image16.png 83 732 media_image16.png Greyscale where I ( 0 ) and I ( 1 )   are fixed as inputs nodes during the searching phase and the last node is formed by channel-wise concatenating of previous intermediate nodes     ∪ ∑ i = 2 n - 1 I ( i )   as the output of cell. [2.2, Page 3]: after the searching phase, the operation o with the maximum β(i, j) α(i,j) for each edge (i,j) is selected and the connection of each node j to its two precedents i < j with maximum o β(i, j) α(i,j) is selected so that a discrete superior architecture can be sampled from the supernet [BRI: selecting the “maximum operation” in the supernet’s weight space is a valid way to sample a discrete, superior architecture by choosing the most active or optimalselect of weights from the supernet’s parameter matrix, which defines a subnet] [2.2, Page 4]: we focus on the L2 bounded perturbations and according to the definition of spectral norm, the Lipschitz constant of these operations is the spectral norm of its weight matrix where   λ 2 0 =   W 2 0 , which also is the maximum singular value of W ,marked as Λ 1 . However, directly computing Λ 1   is not practical through gradient descent. The power iteration method can be applied for an efficient approximation of Λ 1 [22]. Note that although the perturbation is L2 bounded, the robustness against L ∞   can be also achieved, as stated by [23]. - and accumulating the estimated maximum singular values over layers of a sub- network and over sub-networks of the sample of sub-networks. [2.1, Page 2]: regularizing the weight matrix of each layer to form a Lipschitz constrained network has been proven to be beneficial for the adversarial robustness, [2.2, Page 4]: the Lipschitz constant of operations without convolutional layers can be summarized as follows, (1). average pooling: S−0.5 where S denotes the stride of pooling layer, (2). max pooling: 1, (3). identity connection: 1, (4) [BRI: a pooling layer is not a convolutional layer. It the layer after the convolutional layer in a CNN. Average pooling computes the mean of elements in a pooling window and reduces spatial dimensions deterministically. Max pooling selects the maximum value in each pooling window. Unlike average pooling, max pooling is non-linear and non-smooth, so its exact Lipschitz constant depends on the input configuration. Accumulating these estimates across consecutive layers or sub-networks to approximate the overall network Lipschitz constant It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Cummings, Badran and Dong. Cummings teaches a super-net including a plurality of sub-networks, first loss over a sample of sub-networks, and adjusting network weights to reduce the accumulated loss. Badran teaches relating the measure of smoothness to the maximum difference between outputs from the layer relative to any possible difference between inputs to the layer. Dong teaches maximum singular value and training set. One of ordinary skill would be motivated to combine Cummings, Badran, and Dong to constraint Lipschitz constant to improve the robustness (Dong [2.2, Page 2]). Claim 6: (Previously Presented) Cummings, and Badran do not explicitly disclose: - where the measure of smoothness of a layer of a sub-network is based on an estimate of the maximum ratio between variations in the output from the layer and variations in the input to the layer a higher maximum ratio indicating a lower smoothness. However, Dong discloses: - where the measure of smoothness of a layer of a sub-network is based on an estimate of the maximum ratio between variations in the output from the layer and variations in the input to the layer a higher maximum ratio indicating a lower smoothness. [2.3, Page 3]: One concern of NAS for adversarial robustness is the computational cost since both adverarial training and supernet optimization can be time-consuming. [3.2, Page 4]: the relationship between architecture parameters α, β and Lipschitz constant of the network. Since the entire neural network is constructed by stacking cells in series as [I1, I2, ..., IN ], Eq. 2 can be further decomposed as a PNG media_image17.png 88 513 media_image17.png Greyscale where λ l , λ C , λ( I N ) denote the Lipschitz constants of the loss function, classifier, and cell IN respectively. [3.2, Page 4]: Eq. 6 can be unfolded recursively and rewritten as PNG media_image18.png 52 487 media_image18.png Greyscale It is obvious that the adversarial robustness can be bounded by the Lipschitz constants of cells. Eq. 7 also suggests that the impact of perturbation grows exponentially with the number of cells, which further highlights the influence of cell designing. [2.2, Page 3]: the Lipschitz constant of operations without convolutional layers can be summarized as follows, (1). average pooling: S −0.5 where S denotes the stride of pooling layer, (2). max pooling: 1, (3). identity connection: 1, (4). Zeroize: 0. For the rest operations including depthwise separate conv and dilated depth-wise separate conv, we focus on the L 2 bounded perturbations and according to the definition of spectral norm, the Lipschitz constant of these operations is the spectral norm of its weight matrix where PNG media_image19.png 25 116 media_image19.png Greyscale which is also is the maximum singular value of the W. It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Cummings, Badran and Dong. Cummings teaches a super-net including a plurality of sub-networks, first loss over a sample of sub-networks, and adjusting network weights to reduce the accumulated loss. Badran teaches relating the measure of smoothness to the maximum difference between outputs from the layer relative to any possible difference between inputs to the layer. Dong teaches maximum singular value and training set. One of ordinary skill would be motivated to combine Cummings, Badran, and Dong to constraint Lipschitz constant to improve the robustness (Dong [2.2, Page 2]). Claim 22: (Previously Presented) Cummings and Badran do not explicitly disclose: - where the training data includes at least 50,000 training inputs and corresponding training outputs. However, Dong discloses: - where the training data includes at least 50,000 training inputs and corresponding training outputs. [3.1, Page 6]: we search the robust neural architectures on CIFAR-10 dataset which contains 50K training images and 10K validation images over 10 classes It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Cummings, Badran and Dong. Cummings teaches a super-net including a plurality of sub-networks, first loss over a sample of sub-networks, and adjusting network weights to reduce the accumulated loss. Badran teaches relating the measure of smoothness to the maximum difference between outputs from the layer relative to any possible difference between inputs to the layer. Dong teaches maximum singular value and training set. One of ordinary skill would be motivated to combine Cummings, Badran, and Dong to constraint Lipschitz constant to improve the robustness (Dong [2.2, Page 2]). Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to TIRUMALE KRISHNASWAMY RAMESH whose telephone number is (571)272-4605. The examiner can normally be reached by phone. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Li B Zhen can be reached on phone (571-272-3768). The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /TIRUMALE K RAMESH/Examiner, Art Unit 2121 /Li B. Zhen/Supervisory Patent Examiner, Art Unit 2121
Read full office action

Prosecution Timeline

Show 6 earlier events
Nov 24, 2025
Interview Requested
Dec 02, 2025
Applicant Interview (Telephonic)
Dec 11, 2025
Examiner Interview Summary
Dec 12, 2025
Request for Continued Examination
Dec 20, 2025
Response after Non-Final Action
Feb 12, 2026
Non-Final Rejection mailed — §103
May 11, 2026
Response Filed
Aug 25, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12699873
NEURAL NETWORK PROCESSING USING MIXED-PRECISION DATA REPRESENTATION
6y 6m to grant Granted Aug 04, 2026
Patent 12688395
Neural Network Processor with On-Chip Convolution Kernel Storage
8y 5m to grant Granted Jul 21, 2026
Patent 12518153
TRAINING MACHINE LEARNING SYSTEMS
5y 12m to grant Granted Jan 06, 2026
Patent 12293284
META COOPERATIVE TRAINING PARADIGMS
4y 4m to grant Granted May 06, 2025
Patent 12229651
BLOCK-BASED INFERENCE METHOD FOR MEMORY-EFFICIENT CONVOLUTIONAL NEURAL NETWORK IMPLEMENTATION AND SYSTEM THEREOF
4y 4m to grant Granted Feb 18, 2025
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

5-6
Expected OA Rounds
26%
Grant Probability
50%
With Interview (+23.7%)
4y 9m (~4m remaining)
Median Time to Grant
High
PTA Risk
Based on 49 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month