Prosecution Insights
Last updated: October 01, 2026
Application No. 18/839,379

Neural Architecture Search with Improved Computational Efficiency

Non-Final OA §101§103
Filed
Aug 16, 2024
Priority
Feb 16, 2022 — provisional 63/310,837 +1 more
Examiner
ASEGDEW, NATNAEL AREGA
Art Unit
Tech Center
Assignee
Google LLC
OA Round
1 (Non-Final)
Grant Probability
Favorable
1-2
OA Rounds

Examiner Intelligence

Grants only 0% of cases
0%
Career Allowance Rate
0 granted / 0 resolved
-60.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
Avg Prosecution
13 currently pending
Career history
9
Total Applications
across all art units
This examiner has no resolved cases yet (career too new); statute-level performance unavailable. The Grant Probability card shows Tech Center averages instead.

Office Action

§101 §103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . This action is in response to the instant application filed on 08/16/2024. Information Disclosure Statement The information disclosure statement (IDS) submitted on 02/13/2025 is being considered by the examiner. Claim Objections Claims 7 and 17 objected to because of the following informalities: typo in this limitation: the one or more constraints comprise a runtime latency constraint that requires that ta runtime… Appropriate correction is required. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Regarding claim 1: Step 1: The claim recites a method which falls into the statutory category of process. Step 2A Prong 1: The claim recites multiple abstract ideas: defining, by a computing system comprising one or more computing devices, a plurality of searchable parameters that control an architecture of a neural network, a mental process given a human being can define search parameters in their mind; determining, by the computing system using a controller model, a new set of values for the plurality of searchable parameters to generate a new architecture for the neural network, a mental process given a human being can pick new values for search parameters in their mind; determining, by the computing system, whether the neural network with the new architecture satisfies one or more constraints, a mental process given human being can determine whether or not a neural network satisfies constraints in their mind; when the neural network with the new architecture does not satisfy the one or more constraints: discarding, by the computing system, the new architecture, a mental process given a human being can remove an architecture from consideration in their mind; and when the neural network with the new architecture satisfies the one or more constraints: determining, by the computing system, one or more performance metrics for the neural network with the new architecture relative to production of inferences for a set of validation data, a mental process given a human being can choose performance metrics in their mind; evaluating, by the computing system, a value function that provides a value based at least in part on the one or more performance metrics and a conditional probability of the new architecture for the neural network given that the new architecture for the neural network satisfies the one or more constraints, a mental process given a human being can evaluate the value of a neural network in their mind; and updating, by the computing system, one or more values of one or more parameters of the controller model based on the value function, a mental process given a human being can change parameters in their mind or with the aid of a generic computer. Step 2A Prong 2: The claim does not recite additional elements that integrate the abstract idea into a practical application given: a computing system comprising one or more computing device, amounts to mere instructions to apply the judicial exception wherein the neural network is configured to process input data to produce inferences, amounts to generally linking the judicial exception to a technological environment a controller model, amounts to mere instructions to apply the judicial exception Step 2B: The claim does not include additional elements, when taken alone or in combination, that are sufficient to amount to significantly more than the judicial exception. As discussed above, the computing system and controller model are considered mere instructions to apply the judicial exception through a generic computer (see MPEP 2106.05(f)). Additionally the neural network configured to produce inferences is considered generally linking the judicial exception to a technological environment given it only recites a general neural network that produces an output (see MPEP 2106.05(h)). Claim 1 is not patent eligible. Regarding claim 2, the rejection of claim 1 is incorporated, further the claim recites: wherein the conditional probability comprises an exact conditional probability. This limitation amounts to more specifics of the abstract idea of evaluating a value function. The claim does not include any additional elements that amount to an integration of the judicial exceptions into a practical application, nor to significantly more than the judicial exceptions. Claim 2 is not patent eligible. Regarding claim 3, the rejection of claim 1 is incorporated, further the claim recites: wherein the conditional probability comprises an estimated conditional probability. This limitation amounts to more specifics of the abstract idea of evaluating a value function. The claim does not include any additional elements that amount to an integration of the judicial exceptions into a practical application, nor to significantly more than the judicial exceptions. Claim 3 is not patent eligible. Regarding claim 4, the rejection of claim 3 is incorporated, further the claim recites: wherein evaluating, by the computing system, the value function comprises performing, by the computing system, a Monte-Carlo sampling technique to determine the estimated conditional probability. This limitation amounts to a mathematical calculation. The claim does not include any additional elements that amount to an integration of the judicial exceptions into a practical application, nor to significantly more than the judicial exceptions. Claim 4 is not patent eligible. Regarding claim 5, the rejection of claim 1 is incorporated, further the claim recites: wherein: the one or more constraints comprise a size constraint that requires that a number of parameters included in the new network architecture does not exceed a threshold number of parameters. This limitation amounts to more specifics of the abstract idea of determining whether the neural network satisfies constraints. The claim does not include any additional elements that amount to an integration of the judicial exceptions into a practical application, nor to significantly more than the judicial exceptions. Claim 5 is not patent eligible. Regarding claim 6, the rejection of claim 1 is incorporated, further the claim recites: wherein: the one or more constraints comprise a training latency constraint that requires that training of neural network with the new architecture does not exceed a threshold training time. This limitation amounts to more specifics of the abstract idea of determining whether the neural network satisfies constraints. The claim does not include any additional elements that amount to an integration of the judicial exceptions into a practical application, nor to significantly more than the judicial exceptions. Claim 6 is not patent eligible. Regarding claim 7, the rejection of claim 1 is incorporated, further the claim recites: wherein: the one or more constraints comprise a runtime latency constraint that requires that ta runtime latency of neural network with the new architecture does not exceed a threshold runtime. This limitation amounts to more specifics of the abstract idea of determining whether the neural network satisfies constraints. The claim does not include any additional elements that amount to an integration of the judicial exceptions into a practical application, nor to significantly more than the judicial exceptions. Claim 7 is not patent eligible. Regarding claim 8, the rejection of claim 1 is incorporated, further the claim recites: wherein the validation data comprises tabular data. This limitation amounts to generally linking the judicial exception to a field of use (tabular data) see MPEP 2106.05(h). The claim does not include any additional elements that amount to an integration of the judicial exceptions into a practical application, nor to significantly more than the judicial exceptions. Claim 8 is not patent eligible. Regarding claim 9, the rejection of claim 1 is incorporated, further the claim recites: wherein discarding, by the computing system, the new architecture comprises discarding, by the computing system, the new architecture prior to completion of training of a neural network having the new architecture. This limitation amounts to more specifics of the abstract idea of discarding architecture since it describes making the choice to not train the neural network having the new architecture. The claim does not include any additional elements that amount to an integration of the judicial exceptions into a practical application, nor to significantly more than the judicial exceptions. Claim 9 is not patent eligible. Regarding claim 10, the rejection of claim 1 is incorporated, further the claim recites: herein the controller model comprises a plurality of sets of logits that are respectively associated with the plurality of searchable parameters, and wherein each of the plurality of sets of logits generates its respective prediction independent of the other sets of logits. This limitation amounts to mere instructions to apply the abstract idea given it merely describes how the generic controller model generates predictions. The claim does not include any additional elements that amount to an integration of the judicial exceptions into a practical application, nor to significantly more than the judicial exceptions. Claim 10 is not patent eligible. Regarding claims 11-20, the claims teach essentially the same limitations as the claims above (or some combination thereof), therefore the claims are rejected for at least the reasons listed above. Claims 11-20 are not patent eligible. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claim(s) 1-3, 5-7, 9, 11-13, 15-17, 19 are rejected under 35 U.S.C. 103 as being unpatentable over Bender (Can weight sharing outperform random architecture search? An investigation with TuNAS) (As found in IDS dated 02/13/2025) in view of Wang (AttentiveNAS: Improving Neural Architecture Search via Attentive Sampling). Regarding claim 1, Bender teaches defining, by a computing system comprising one or more computing devices, a plurality of searchable parameters that control an architecture of a neural network, wherein the neural network is configured to process input data to produce inferences (Section 3, Our goal is to develop a NAS method that can reliably find high quality models at a specific inference cost across multiple search spaces. We next present three progressively larger search spaces and show that they are non-trivial); and for one or more iterations: determining, by the computing system using a controller model (Section 5, For architecture search experiments, we always repeat the entire search process 5 times as suggested by Lindauer and Hutter, Section 4.2, models sampled by the RL controller), a new set of values for the plurality of searchable parameters to generate a new architecture for the neural network (Section 4, We alternate between learning the shared weights W using gradient descent and learning the policy π using REINFORCE [38]. At each step, we first sample a network architecture α ∼ π.); evaluating, by the computing system, a value function that provides a value based at least in part on the one or more performance metrics (Section 4, During a search, we learn a policy π, a probability distribution from which we can sample high quality architectures…. The accuracy Q(α) and inference time T(α) jointly determine the reward r(α) which is used to update the policy π, the value function is the reward which is based on the quality and inference time and sampling from the policy (a probability distribution)); and updating, by the computing system, one or more values of one or more parameters of the controller model based on the value function (Section 4, We alternate between learning the shared weights W using gradient descent and learning the policy π using REINFORCE [40].). Bender fails to teach determining, by the computing system, whether the neural network with the new architecture satisfies one or more constraints; when the neural network with the new architecture does not satisfy the one or more constraints: discarding, by the computing system, the new architecture; and when the neural network with the new architecture satisfies the one or more constraints: determining, by the computing system, one or more performance metrics for the neural network with the new architecture relative to production of inferences for a set of validation data and Wang teaches determining, by the computing system, whether the neural network with the new architecture satisfies one or more constraints; when the neural network with the new architecture does not satisfy the one or more constraints: discarding, by the computing system, the new architecture (Section 4.2, To draw an architecture sample given a FLOPs constraint, a straightforward strategy is to leverage rejection sampling, i.e., draw samples uniformly from the entire search space and reject samples if the targeted FLOPs constraint is not satisfied); and when the neural network with the new architecture satisfies the one or more constraints: determining, by the computing system, one or more performance metrics for the neural network with the new architecture relative to production of inferences for a set of validation data (Section 4.3, Accuracy predictor as performance estimator: train an accuracy predictor on a validation set; then for each architecture, use the predicted accuracy given by the accuracy predictor as its performance estimation) and (Section 4.2, At each sampling step, one needs to first draw a sample of target FLOPs τ0 according to the prior distribution π(τ); and then sample k architectures {a1,··· ,ak} from π(α | τ0), the value function can be calculated based on whether the architecture is chosen which is based on a conditional probability). Bender and Wang are analogous to the claimed invention because they are in the field of neural architecture search. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have used the sampling method in Wang along with Bender to sample models that are constrained (on inference time in Bender and on FLOPs in Wang) to speed up the constrained sampling process (Wang Section 4.1, This naive sampling strategy, however, is inefficient especially when the search space is large. To speedup the FLOPs-constrained sampling process we propose to approximate π(α | τ) empirically). Regarding claim 2, Wang teaches, which Bender is silent on, wherein the conditional probability comprises an exact conditional probability (Section 4.2, At each sampling step, one needs to first draw a sample of target FLOPs τ0 according to the prior distribution π(τ); and then sample k architectures {a1,··· ,ak} from π(α | τ0)). Regarding claim 3, Wang teaches, which Bender fails to teach, wherein the conditional probability comprises an estimated conditional probability (Section 4.2, Let ˆπ(α | τ) denote an empirical approximation of π(α | τ)). Regarding claim 4, Wang teaches, which Bender fails to teach, wherein evaluating, by the computing system, the value function comprises performing, by the computing system, a Monte-Carlo sampling technique to determine the estimated conditional probability (Section 4.1, To draw an architecture sample given a FLOPs constraint, a straightforward strategy is to leverage rejection sampling, i.e., draw samples uniformly from the entire search space and reject samples if the targeted FLOPs constraint is not satisfied…. To speedup the FLOPs-constrained sampling process, we propose to approximate π(α | τ) empirically, Section 3.1, To solve this optimization, in practice, we can approximate the expectation over π(τ) with n Monte Carlo samples of FLOPs {τo}. Then, for each targeted FLOPs τo, we can approximate the summation over π(α | τo)). Regarding claim 5, Bender teaches wherein: the one or more constraints comprise a size constraint that requires that a number of parameters included in the new network architecture does not exceed a threshold number of parameters (Section 2, Recently, Neural Architecture Search has been used intensively to find architectures that have better tradeoff between accuracy and latency [36, 39, 8, 5, 35], FLOPS [37], power consumption [14], and memory usage [9], memory usage and FLOPS implies parameter limit). Regarding claim 6, Bender teaches wherein: the one or more constraints comprise a training latency constraint that requires that training of neural network with the new architecture does not exceed a threshold training time (Section 2, In our experiments we typically needed to run 7 searches to achieve the desired latency…... Our focus is on larger and more challenging searches which incorporate latency constraints.) Regarding claim 7, Bender teaches wherein: the one or more constraints comprise a runtime latency constraint that requires that ta runtime latency of neural network with the new architecture does not exceed a threshold runtime (Table 2, We use a target inference time of 84ms for the first two (to compare against ProxylessNAS and MnasNet) and 57ms for the third search space (to compare against MobileNetV3)). Regarding claim 9, Wang teaches wherein discarding, by the computing system, the new architecture comprises discarding, by the computing system, the new architecture prior to completion of training of a neural network having the new architecture (Section 4.2, To draw an architecture sample given a FLOPs constraint, a straightforward strategy is to leverage rejection sampling, i.e., draw samples uniformly from the entire search space and reject samples if the targeted FLOPs constraint is not satisfied). Regarding claims 11-17, 19, the claims cover essentially same limitation as the claims above and are rejected for the same reasons. Claim(s) 8, 18, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Bender in view of Wang as applied to claim 1 above, and further in view of Égelé (AgEBO-Tabular: Joint Neural Architecture and Hyperparameter Search with Autotuned Data-Parallel Training for Tabular Data) (As found in IDS dated 02/13/2025). Regarding claim 8, Bender in view of Wang teaches the method of claim 1 but fails to teach wherein the validation data comprises tabular data. Égelé teaches wherein the validation data comprises tabular data (Abs, To that end, we develop AgEBO-Tabular, which combines Aging Evolution (AE) to search over neural architectures and asynchronous Bayesian optimization (BO) to search over hyperparameters to adapt data-parallel training. We evaluate the efficacy of our approach on two large predictive modeling tabular data sets from the Exascale Computing Project CANcer Distributed Learning Environment (ECP-CANDLE).) Égelé and Bender are analogous to the claimed invention because they are in the field of neural architecture search. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have combined the method of Bender with Égelé to extend neural architecture search to tabular data while avoiding large computation time for large data sets (Abs, A key issue in NAS, particularly for large data sets, is the large computation time required to evaluate each generated architecture). Regarding claims 18 and 20, the claims teach essentially the same limitation and as such are rejected for the reasons above. Claim 10 is rejected under 35 U.S.C. 103 as being unpatentable over Bender in view of Wang as applied to claim 1 above, and further in view of Zoph (NEURAL ARCHITECTURE SEARCH WITH REINFORCEMENT LEARNING) (as cited in IDS dated 02/13/2025). Regarding claim 10, Bender in view of Wang teaches the method of claim 1 but fails to teach wherein the controller model comprises a plurality of sets of logits that are respectively associated with the plurality of searchable parameters, and wherein each of the plurality of sets of logits generates its respective prediction independent of the other sets of logits. Zoph teaches wherein the controller model comprises a plurality of sets of logits that are respectively associated with the plurality of searchable parameters, and wherein each of the plurality of sets of logits generates its respective prediction independent of the other sets of logits (Section 3.2 Training with Reinforce, The list of tokens that the controller predicts can be viewed as a list of actions a1:T to design an architecture for a child network, In this work, we use the REINFORCE rule from Williams (1992): PNG media_image1.png 222 904 media_image1.png Greyscale Where m is the number of different architectures that the controller samples in one batch and T is the number of hyperparameters our controller has to predict to design a neural network architecture, different calculation for each architecture and hyperparameter involving logits as shown In the equation.) Zoph is analogous to the claimed invention because it is in the field of network architecture search using REINFORCE. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have used the empirical approximation of the REINFORCE update rule used in Zoph along with Bender and Wang because it is a simple substitution of the regular REINFORCE method used in Bender with a predictable outcome. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to NATNAEL A ASEGDEW whose telephone number is (571)270-0407. The examiner can normally be reached 7:30-5. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kakali Chaki can be reached at (571) 272-3719. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /NATNAEL A ASEGDEW/ Examiner, Art Unit 2122 /MICHAEL H HOANG/ PRIMARY EXAMINER, Art Unit 2122
Read full office action

Prosecution Timeline

Aug 16, 2024
Application Filed
Sep 23, 2026
Non-Final Rejection mailed — §101, §103 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
Grant Probability
Low
PTA Risk
Based on 0 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month