Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 05/21/2024 was filed before the mailing date of the first office action. The submission is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Objections
Claim 3 is objected to because of the following informalities: the claim recites “generating a random sample of Deep Neural Networks (DNN) (DNNs) within the defined initial search space”, which should be corrected to “generating a random sample of Deep Neural Networks (DNNs) within the defined initial search space”.
Claim 5 is objected to because of the following informalities: the claim recites “types of blocks and/or topological interconnections between the types of learning units”, which should be corrected to recite “and” or “or”.
Claim 7 is objected to because of the following informalities: the preamble recites “the method as in clam 1”, which should be corrected to “the method as in claim 1”. Appropriate correction is required.
Claim 8 is objected to because of the following informalities: the claim recites: “a system monitor comprising a Predictor Monitor and a Performance Monitor, which cooperate to detect concept drifts and/or data drifts efficiently and react accordingly by triggering Automatic Refactor”, which should be corrected to recite “and” or “or”.
Claim 9 is objected to because of the following informalities: the claim recites “and/or a framework obsolescence signal, which is handled by the design principles search module”, which should be corrected to recite “and” or “or”.
Appropriate correction is required.
Claim Interpretation
Examiner notes that the following claims recite limitations that do not actively claim the described steps:
“a machine learning application server enabled to receive user requests from client devices” in claim 8.
“a system monitor comprising a Predictor Monitor and a Performance Monitor, which cooperate to detect concept drifts and/or data drifts efficiently” in claim 8.
“wherein the Neural Architecture Search module is enabled to predict performance” in claim 11.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 8-11 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Claim 8 recites the limitation “detect concept drifts and/or data drifts efficiently”. The term “efficiently” in claim 8 is a relative term which renders the claim indefinite. The term “efficiently” is not defined by the claim, the specification does not provide a standard for ascertaining the requisite degree, and one of ordinary skill in the art would not be reasonably apprised of the scope of the invention. Examiner notes that this limitation is being interpreted such that concept drifts or data drifts can be detected by a system monitor.
Dependent claims 9-11 are also rejected because they fail to correct the deficiencies of
independent claim 8 on which they depend.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-7 are rejected under 35 U.S.C. 101. Claims 1-7 are directed to a method; therefore, claims 1-7 fall within one of the four statutory categories (i.e., process, machine, manufacture, or composition of matter). However, claims 1-7 fall within the judicial exception of an abstract idea, specifically the abstract ideas of “Mental Processes” (including observation, evaluation, and opinion) and “Mathematical Concepts (including mathematical calculations and relationships)”.
Claim 1:
Claim 1 is directed to a method; therefore, the claim does fall within one of the four statutory categories (i.e., process, machine, manufacture, or composition of matter).
Claim 1 recites the following abstract ideas:
Step 2A Prong 1:
checking whether there is new data (mental step directed to observation, evaluation – a person could check, or observe, whether there is new data in their mind);
detecting obsolescence criteria based on: framework obsolescence; architectural obsolescence; and obsolescence of knowledge (mental step directed to observation, evaluation – a person could detect obsolescence criteria based on observed or mentally determined framework, architectural, and knowledge obsolescence in their mind),
automatically refactoring a model database including: based on the framework obsolescence being detected, searching for new design principles (mental step directed to observation, evaluation – a person could automatically refactor a model database in their mind, potentially assisted by pen and paper, see MPEP 2106.04(a)(2)(III)), by searching for new design principles based on framework obsolescence being mentally observed, or detected);
based on the architectural obsolescence being detected, searching for new network architectures (mental step directed to observation, evaluation – a person could search for new network architectures in their mind based on architectural obsolescence being mentally observed, or detected);
based on a number of validated deep neural networks being greater than a threshold, calculating performance predictors for each validated neural network (mental step directed to observation, evaluation – a person could calculate performance predictors for each observed validated neural network in their mind based on an observed or mentally determined number of observed validated deep neural networks being greater than a threshold).
Claim 1 recites the following additional elements:
receiving a user request; and based on the obsolescence of knowledge being detected, training and validating new deep neural networks.
Step 2A Prong 2:
Receiving a user request is interpreted as insignificant extra-solution activity directed to mere data gathering. As the claim does not recite any particular deep neural network architecture nor technical steps for training and validation, training and validating new deep neural networks is interpreted as generic computer activity in the technological environment in which the claimed abstract ideas are performed. Wherein the networks are trained and validated based on the obsolescence of knowledge being detected is interpreted as a description of the state of the technological environment under which this step is performed. These additional elements, when considered as a whole with the aforementioned abstract ideas, do not integrate those abstract ideas into a practical application (see MPEP 2106.05(g) and MPEP 2106.05(h)).
Step 2B:
Receiving a user request is interpreted as well-understood, routine, conventional activity directed to receiving data over a network. As the claim does not recite any particular deep neural network architecture nor technical steps for training and validation, training and validating new deep neural networks is interpreted as generic computer activity. Wherein the networks are trained and validated based on the obsolescence of knowledge being detected is interpreted as a description of the state of the technological environment under which this step is performed. These additional elements, when considered as a whole with the aforementioned abstract ideas, do not amount to significantly more than those abstract ideas (see MPEP 2106.05(d)).
Claim 2 recites wherein the user request is a model search signal sent by a user.
This limitation is interpreted as a further description of the kind of data merely gathered and received over a network in the analysis of claim 1, and does not integrate the claimed abstract ideas into a practical application or amount to significantly more than the claimed abstract ideas.
Claim 3 recites wherein the searching for the new design principles comprises: defining an initial search space; generating a random sample of Deep Neural Networks (DNN) (DNNs) within the defined initial search space; analyzing hyperparameters of best DNNs using empirical bootstrap, wherein the empirical bootstrap resamples random DNNs with replacement and selects a DNN with a highest score predicted by zero-cost performance predictors.
Searching for new design principles by defining an initial search space is interpreted as a mental step directed to observation, evaluation – a person could define an initial search space in their mind, potentially assisted by pen and paper. Generating a random sample of DNNs within a defined initial search space is interpreted as a mental step directed to observation, evaluation – a person could generate a random sample of observed DNNs within a defined initial search space in their mind. Analyzing hyperparameters of best DNNs using an empirical bootstrap method to resample random DNNs is interpreted as a mental step directed to observation, evaluation – a person could analyze hyperparameters of DNNs by using an empirical bootstrap method to resample random DNNs in their mind. Selecting a DNN with a highest score predicted by zero-cost performance predictors is interpreted as a mental step directed to observation, evaluation – a person could select a DNN with an observed highest score predicted by zero-cost performance predictors in their mind.
Claim 4 recites wherein a resampling size is an operational parameter that defines a number of DNNs that will be resampled in each empirical bootstrap attempt.
This limitation is interpreted as a further description of the resampling parameter used in the abstract ideas recited in claim 3, and does not integrate those abstract ideas into a practical application or amount to significantly more than those abstract ideas.
Claim 5 recites wherein the neural network design principles are divided between: design principles relating to quantitative hyperparameters, including a number of neurons, a number of layers, a number of filters; and design principles relating to non-quantitative hyperparameters including types of learning units, types of blocks and/or topological interconnections between the types of learning units, types of activation functions.
This limitation is interpreted as a further description of the kinds of neural network design principals that were searched for in the abstract idea analysis of claim 1. This further description, when considered as a whole with the aforementioned abstract ideas, does not integrate those abstract ideas into a practical application or amount to significantly more than those abstract ideas.
Claim 6 recites wherein an empirical distribution function (EDF) is computed as a function of a scaled_score, as follows: EDF(x) =
|
s
∈
O
:
s
≤
x
|
|
s
∈
O
|
where “x” is a scaled score threshold, “s∈O” is a scaled score of a network from an observation set, s∈O:s≤x” is a subset of all scaled scores that are less than or equal a given scaled score threshold “x”, and “|⋅|” is a number of elements in the subset.
Computing an empirical distribution function as a function of a scaled score according to the claimed equation is interpreted as an abstract idea directed to a mathematical calculation.
Claim 7 recites wherein an area under a curve (AUC) EDF indicator is defined as follows:
A
U
C
E
D
F
=
∫
0
1
E
D
F
x
d
x
, wherein the AUC EDF indicator is computed numerically with a trapezoidal rule.
Computing an AUC EDF indicator with a trapezoidal rule according to the claimed equation is interpreted as an abstract idea directed to a mathematical calculation.
Viewed as a whole, these additional claim elements do not provide meaningful limitations to transform the abstract idea into a patent eligible application of the abstract idea such that the claims amount to significantly more than the abstract idea itself. Therefore, the claims are rejected under 35 U.S.C. 101.
Examiner notes that claims 8-11 do not actively recite any claim limitations directed to a judicial exception and therefore are eligible under 35 U.S.C. 101 (see the claim interpretation section for analysis of claims 8-9 and 11).
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-2, 4-6, and 8-11 are rejected under 35 U.S.C. 103 as being unpatentable over Ramamonjison et al (US 20230281974 A1, herein Ramamonjison) in view of Kartoun et al (US 20240428127 A1, herein Kartoun), in further view of Radosavovic et al (“Designing Network Design Spaces”, herein Radosavovic).
Regarding claim 1, Ramamonjison teaches a method of automatically optimizing a variety of neural network design principles in production environments, the method comprising: receiving a user request (para. [0020] recites “the method can include receiving a request through a network from computing device to adapt the machine learning model that is configured by a learned set of configuration parameters; and returning the final set of adapted configuration parameters for the machine learning model to the requesting computer device (i.e., a method of optimizing a neural network design in response to a request from a device such as the edge, or user device described in at least para. [0080]));
checking whether there is new data; detecting obsolescence criteria based on: [framework obsolescence]; architectural obsolescence; and obsolescence of knowledge (para. [0036] recites “When a trained machine learning model is used to perform objection detection (referred to hereinafter as a trained object detection model) and the trained object detection model is deployed to a computing system to perform inference using a new set of images which are different than the set of labeled training images that was used to train the object detection model, domain shift can occur” (i.e., receiving new data and detecting that the previous data, or knowledge, model architecture, and model framework may be outdated, or obsolete)),
based on the architectural obsolescence being detected, searching for new network architectures (para. [0065] recites “it may be desirable to use an object detection model in a shifted domain that has a different model architecture than that of the source model 100. For example, a simplified model that can run on a computationally constrained edge device may be desired. According to example embodiments, the method 200 can be applied with additional steps to apply domain shift adaptation to a further model that has a different model architecture than source model 100” (i.e., searching for new network architectures when a previous architecture is no longer desirable, or obsolete));
based on the obsolescence of knowledge being detected, training and validating new deep neural networks (para. [0006] recites “One solution that has been proposed to address the domain shift problem for a trained object detection model involves further training a trained object detection model with a set of new images that are from the target domain” (i.e., training newly configured models with new data when domain shift, or knowledge obsolescence, is determined to have occurred)); based on a number of validated deep neural networks being greater than a threshold, calculating performance predictors for each validated neural network (para. [0012] recites “The method further includes performing a plurality of model adaptation epochs, where for an initial model adaptation epoch the learned set of configuration parameters is used as a current set of configuration parameters for the machine learning model. Each model adaptation epoch may include: predicting for each of the plurality of targets image samples, using the machine learning model configured by the current set of configuration parameters. When performing the plurality of model adaptation epochs is completed, the current set of configuration parameters is output as a final set of adapted configuration parameters for the machine learning model”. Para. [0071] recites “The goal is to obtain an adapted object detection model that can achieve good performance on both the set of clean test images and the unlabeled set of corrupted images from the target domain” (i.e., predicting the performance of models that have been retrained and determined to have improved over the original models before domain adaptation)).
However, Ramamonjison does not explicitly teach detecting obsolescence criteria based on: framework obsolescence or automatically refactoring a model database including: based on the framework obsolescence being detected, searching for new design principles.
Radosavovic teaches detecting obsolescence criteria based on: framework obsolescence and automatically refactoring a model database including: based on the framework obsolescence being detected, searching for new design principles (Radosavovic section 3 para. 3 recites “we propose to design progressively simplified versions of an initial, unconstrained design space. We refer to this process as design space design. Design space design is akin to sequential manual network design, but elevated to the population level. Specifically, in each step of our design process the input is an initial design space and the output is a refined design space, where the aim of each design step is to discover design principles that yield populations of simpler or better performing models” (i.e., determining, or detecting, a model that could be improved, or a model with framework obsolescence, and refactoring the model with new design principles)).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine these teachings by utilizing the design space design method from Radosavovic to supplement the domain shift model adaptation system from Ramamonjison. Radosavovic and Ramamonjison are both directed to methods of analyzing and improving neural networks. One of ordinary skill in the art would be motivated to combine these model adaptation methods to improve the ability to modify neural network models, as at least section 2 of Radosavovic teaches “better design spaces can improve the efficiency of [architecture] search algorithms and also lead to existence of better models by enriching the design space”.
Regarding claim 2, the combination of Ramamonjison and Radosavovic teaches the method of claim 1 as mentioned above, wherein the user request is a model search signal sent by a user (Ramamonjison para. [0020] recites “the method can include receiving a request through a network from computing device to adapt the machine learning model that is configured by a learned set of configuration parameters; and returning the final set of adapted configuration parameters for the machine learning model to the requesting computer device (i.e., searching for an optimized model in response to a request from a device such as the edge, or user device described in at least para. [0080])).
Regarding claim 4, the combination of Ramamonjison and Radosavovic teaches the method of claim 3 as mentioned above, wherein a resampling size is an operational parameter that defines a number of DNNs that will be resampled in each empirical bootstrap attempt (Radosavovic section 3.1 para. 4-5 recite “we employ an empirical bootstrap to estimate the likely range in which the best models fall. To summarize: (1) we generate distributions of models obtained by sampling and training n models from a design space, (2) we compute and plot error EDFs to summarize design space quality, (3) we visualize various properties of a design space and use an empirical bootstrap to gain insight, and (4) we use these insights to refine the design space. Given n pairs (xi, ei) of model statistic xi (e.g. depth) and corresponding error ei, we compute the empirical bootstrap by: (1) sampling with replacement 25% of the pairs, (2) selecting the pair with min error in the sample, (3) repeating this 104 times, and finally (4) computing the 95% CI for the min x value. The median gives the most likely best value” (i.e., resampling determines the number of DNNs in each empirical bootstrap run)).
Regarding claim 5, the combination of Ramamonjison and Radosavovic teaches the method of claim 1 as mentioned above, wherein the neural network design principles are divided between: design principles relating to quantitative hyperparameters, including a number of neurons, a number of layers, a number of filters; and design principles relating to non-quantitative hyperparameters including types of learning units, types of blocks and/or topological interconnections between the types of learning units, types of activation functions (Radosavovic section 3.2 para. 1 recites “Our focus is on exploring the structure of neural networks assuming standard, fixed network blocks (e.g., residual bottleneck blocks). In our terminology the structure of the network includes elements such as the number of blocks (i.e. network depth), block widths (i.e. number of channels), and other block parameters such as bottleneck ratios or group widths”. Ramamonjison para. [0065] recites “In some scenarios, it may be desirable to use an object detection model in a shifted domain that has a different model architecture than that of the source model 100. Student model 700S may be a smaller, faster model in that it has fewer computation blocks 104 (and thus fewer layers 106, 110) than source model 100” (i.e., the design principles can include quantitative parameters such as a number of layers from Ramamonjison and non-quantitative parameters such as types of blocks)).
Regarding claim 6, the combination of Ramamonjison and Radosavovic teaches the method of claim 1 as mentioned above, wherein an empirical distribution function (EDF) is computed as a function of a scaled_score, as follows: EDF(x) =
|
s
∈
O
:
s
≤
x
|
|
s
∈
O
|
where “x” is a scaled score threshold, “s∈O” is a scaled score of a network from an observation set, s∈O:s≤x” is a subset of all scaled scores that are less than or equal a given scaled score threshold “x”, and “|⋅|” is a number of elements in the subset (Radosavovic section 3.1 para. 3-5 recite “our primary tool for analyzing design space quality is the error empirical distribution function (EDF). The error EDF of n models with errors ei is given by: (EQ1) F(e) gives the fraction of models with error less than e. Given a population of trained models, we can plot and analyze various network properties versus network error, see Figure 2. For these plots, we employ an empirical bootstrap4 to estimate the likely range in which the best models fall. Given n pairs (xi, ei) of model statistic xi (e.g. depth) and corresponding error ei, we compute the empirical bootstrap by: (1) sampling with replacement 25% of the pairs, (2) selecting the pair with min error in the sample, (3) repeating this 104 times, and finally (4) computing the 95% CI for the min x value. The median gives the most likely best value” (i.e., calculating an empirical distribution function as a function of the networks in the observation set)).
Regarding claim 8, Ramamonjison teaches a system of automatically optimizing a variety of neural network design principles in production environments, the system comprising: client devices; and a machine learning application server enabled to receive user requests from client devices (para. [0012] recites “A system of one or more computers can be configured to perform particular operations or actions by virtue of having software, firmware, hardware, or a combination”. Para. [0020] recites “the method can include receiving a request through a network from computing device to adapt the machine learning model that is configured by a learned set of configuration parameters; and returning the final set of adapted configuration parameters for the machine learning model to the requesting computer device (i.e., a method of optimizing a neural network design in response to a request from a device such as the edge, or user device connected to the server as described in at least para. [0080])), wherein the machine learning application server comprises:
a system monitor comprising a Predictor Monitor and a Performance Monitor, which cooperate to detect concept drifts and/or data drifts efficiently and react accordingly by triggering Automatic Refactor; and an Automatic Refactor, comprising [a design principles search module], a neural architecture search module and a train and validation module (para. [0065] recites “it may be desirable to use an object detection model in a shifted domain that has a different model architecture than that of the source model 100. For example, a simplified model that can run on a computationally constrained edge device may be desired. According to example embodiments, the method 200 can be applied with additional steps to apply domain shift adaptation to a further model that has a different model architecture than source model 100” (i.e., prediction and performance monitoring capable of searching for and retraining, or refactoring, new network architectures when a previous architecture is no longer desirable, or obsolete, due to domain shift));
a models database (para. [0080] recites “In a further example, the adapted object detection models 100A, 100B are stored by cloud computing service provider 1201 in a cloud network 1204 and access to such models is offered as an inference service” (i.e., a models database));
and a user request processor (para. [0020] recites “the method can include receiving a request through a network from computing device to adapt the machine learning model that is configured by a learned set of configuration parameters; and returning the final set of adapted configuration parameters for the machine learning model to the requesting computer device (i.e., a method of optimizing a neural network design in response to a request from a device such as the edge, or user device connected to the server as described in at least para. [0080])).
However, Ramamonjison does not explicitly teach a design principles search module.
Radosavovic teaches a design principles search module (section 3 para. 3 recites “we propose to design progressively simplified versions of an initial, unconstrained design space. We refer to this process as design space design. Design space design is akin to sequential manual network design, but elevated to the population level. Specifically, in each step of our design process the input is an initial design space and the output is a refined design space, where the aim of each design step is to discover design principles that yield populations of simpler or better performing models” (i.e., a module for determining, or detecting, a model that could be improved, or a model with framework obsolescence, and refactoring the model with new design principles)).
See claim 1 for motivation to combine.
Regarding claim 9, the combination of Ramamonjison and Radosavovic teaches the system of claim 8 as mentioned above, wherein depending on a performance degradation detected, the Performance Monitor reports a knowledge obsolescence signal, which is handled by the train and validation module, an architecture obsolescence signal, which is handled by the neural architecture search module (Ramamonjison para. [0065] recites “it may be desirable to use an object detection model in a shifted domain that has a different model architecture than that of the source model 100. For example, a simplified model that can run on a computationally constrained edge device may be desired. According to example embodiments, the method 200 can be applied with additional steps to apply domain shift adaptation to a further model that has a different model architecture than source model 100” (i.e., a performance monitor capable of searching for new network architectures when a previous architecture is no longer desirable, or obsolete, when it has been determined that degradation, or domain shift, has occurred)),
and/or a framework obsolescence signal, which is handled by the design principles search module (Radosavovic section 3 para. 3 recites “we propose to design progressively simplified versions of an initial, unconstrained design space. We refer to this process as design space design. Design space design is akin to sequential manual network design, but elevated to the population level. Specifically, in each step of our design process the input is an initial design space and the output is a refined design space, where the aim of each design step is to discover design principles that yield populations of simpler or better performing models” (i.e., a module determining, or detecting, signals that a model that could be improved, or a model with framework obsolescence, and refactoring the model with new design principles))).
Regarding claim 10, the combination of Ramamonjison and Radosavovic teaches the system of claim 8 as mentioned above, system as in claim 8, wherein the Neural Architecture Search module is implemented as Genetic Algorithm, Hill Climbing, Evolution Strategy, Particle Swarm Optimization, or other algorithms (Ramamonjison teaches a neural network architecture search and adaptation algorithm in at least fig. 7).
Regarding claim 11, the combination of Ramamonjison and Radosavovic teaches the system of claim 8 as mentioned above, system as in claim 8, wherein the Neural Architecture Search module is enabled to predict performance of DNNs with other performance predictors including training with reduced epochs, training with reduced dataset, training surrogate models (Radosavovic section 1 para. 8 recites “We compare top REGNET models to existing networks in various settings. First, REGNET models are surprisingly effective in the mobile regime. We hope that these simple models can serve as strong baselines for future work. Next, REGNET models lead to considerable improvements over standard RESNE(X)T models in all metrics” (i.e., determining performance of DNNs”). Radosavovic section 3.2 para. 4 recites “We repeat the sampling until we obtain n = 500 models in our target complexity regime (360MF to 400MF), and train each model for 10 epochs” (i.e., training with reduced epochs)).
Claim 3 is rejected under 35 U.S.C. 103 as being unpatentable over Ramamonjison et al (US 20230281974 A1, herein Ramamonjison) in view of Kartoun et al (US 20240428127 A1, herein Kartoun), in further view of Radosavovic et al (“Designing Network Design Spaces”, herein Radosavovic), in further view of Abdelfattah et al (“Zero-cost Proxies for Lightweight NAS”, herein Abdelfattah).
Regarding claim 3, the combination of Ramamonjison and Radosavovic teaches the method of claim 1 as mentioned above, wherein the searching for the new design principles comprises: defining an initial search space; generating a random sample of Deep Neural Networks (DNN) (DNNs) within the defined initial search space (Radosavovic section 3 para. 3 recites “we propose to design progressively simplified versions of an initial, unconstrained design space. Specifically, in each step of our design process the input is an initial design space and the output is a refined design space, where the aim of each design step is to discover design principles that yield populations of simpler or better performing models”. Radosavovic section 3.1 para. 2 and para. 5 recite “To obtain a distribution of models, we sample and train models from a design space n. we generate distributions of models obtained by sampling and training n models from a design space, (2) we compute and plot error EDFs to summarize design space quality, (3) we visualize various properties of a design space and use an empirical bootstrap to gain insight, and (4) we use these insights to refine the design space” (i.e., defining an initial search space with a sample of DNNs));
and analyzing hyperparameters of best DNNs using empirical bootstrap, wherein the empirical bootstrap resamples random DNNs with replacement (Radosavovic section 3.1 para. 4-5 recite “we employ an empirical bootstrap to estimate the likely range in which the best models fall. To summarize: (1) we generate distributions of models obtained by sampling and training n models from a design space, (2) we compute and plot error EDFs to summarize design space quality, (3) we visualize various properties of a design space and use an empirical bootstrap to gain insight, and (4) we use these insights to refine the design space. Given n pairs (xi, ei) of model statistic xi (e.g. depth) and corresponding error ei, we compute the empirical bootstrap by: (1) sampling with replacement 25% of the pairs, (2) selecting the pair with min error in the sample, (3) repeating this 104 times, and finally (4) computing the 95% CI for the min x value. The median gives the most likely best value” (i.e., analyzing and resampling DNNs with replacement using empirical bootstrap))
However, the combination of Ramamonjison and Radosavovic does not explicitly teach selecting a DNN with a highest score predicted by zero-cost performance predictors.
Abdelfattah teaches selecting a DNN with a highest score predicted by zero-cost performance predictors (Abdelfattah section 5 recites “Mellor proposed using jacob cov to score a set of randomly sampled models and to greedily choose the model with the highest score. This “NAS without training” methodology is very attractive thanks to its simplicity and low computational cost. In this section, we evaluate our metrics in this setting that we simply call “random search” (RAND). We extend this methodology slightly: instead of just training the top model, we keep training models (from best to worst as ranked by the zero-cost metric) until the desired accuracy is achieved” (i.e., using zero-cost performance predictors to select a highest predicted score)).
it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine these teachings by applying the zero-cost performance predictors from Abdelfattah to evaluate the models from Ramamonjison (as modified by Radosavovic). Abdelfattah teaches in at least section 5 that is zero-cost metrics can be integrated with other neural network optimization algorithms, stating “we also investigate how to integrate zero-cost metrics within existing NAS algorithms such as reinforcement learning (RL), aging evolution (AE) search, and predictor-based search. More specifically, we investigate enhancing these search algorithms through either (a) zero-cost warmup phase or (b) zero-cost move proposal”. One of ordinary skill in the art would be motivated to enhance the model optimization method from Ramamonjison with these zero-cost performance predictors from Abdelfattah.
Claim 7 is rejected under 35 U.S.C. 103 as being unpatentable over Ramamonjison et al (US 20230281974 A1, herein Ramamonjison) in view of Kartoun et al (US 20240428127 A1, herein Kartoun), in further view of Radosavovic et al (“Designing Network Design Spaces”, herein Radosavovic), in further view of Kartoun et al (US 20240428127 A1, herein Kartoun).
Regarding claim 7, the combination of Ramamonjison and Radosavovic teaches the method of claim 1 as mentioned above, wherein an area under a curve (AUC) EDF indicator is defined as follows:
A
U
C
E
D
F
=
∫
0
1
E
D
F
x
d
x
, (Radosavovic section 3.1 para. 3 recites “our primary tool for analyzing design space quality is the error empirical distribution function” (i.e., an empirical distribution function shown in at least the AUC graphs from fig. 5, 7, and 9-10).
However, the combination of Ramamonjison and Radosavovic does not explicitly teach wherein the AUC EDF indicator is computed numerically with a trapezoidal rule; however, Kartoun explicitly teaches wherein the AUC EDF indicator can be computed numerically with a trapezoidal rule (para. [0049] recites “any mathematical technique for approximating definite integrals can be applied to calculate AUC values. For example, trapezoidal rule approximation or Riemann sum approximation can be used to calculate AUC values” (i.e., calculating the area under the curve with a trapezoidal rule)). One of ordinary skill in the art would recognize that the method of approximating integrals from Kartoun could be applied to use the trapezoidal rule for the area under the curve calculations from Radosavovic.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
“On Network Design Spaces for Visual Recognition” (Radosavovic et al) teaches a method for comparison of distribution estimates, in which network design spaces are compared by applying statistical techniques to populations of sampled models, while controlling for confounding factors like network complexity.
“Refactoring Neural Networks for Verification” (Shriver et al) teaches an automated framework for DNN refactoring of the architecture and distillation of learning relationships within a transformed network.
US 20190180186 A1 (Liang et al) teaches a method for evolving and implementing different architectures for deep neural networks using a genetic algorithm.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to LEAH M FEITL whose telephone number is (571) 272-8350. The examiner can normally be reached on M-F 0900-1700 EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Viker Lamardo can be reached on (571) 270-5871. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/L.M.F./ Examiner, Art Unit 2147
/VIKER A LAMARDO/Supervisory Patent Examiner, Art Unit 2147