Prosecution Insights
Last updated: October 02, 2026
Application No. 18/173,347

Proxy Task Design Tools for Neural Architecture Search

Final Rejection §103§112
Filed
Feb 23, 2023
Examiner
AGRAWAL, SHISHIR
Art Unit
2123
Tech Center
2100 — Computer Architecture & Software
Assignee
Google LLC
OA Round
2 (Final)
8%
Grant Probability
At Risk
3-4
OA Rounds
4m
Est. Remaining
24%
With Interview

Examiner Intelligence

Grants only 8% of cases
8%
Career Allowance Rate
2 granted / 24 resolved
-46.7% vs TC avg
Strong +15% interview lift
Without
With
+15.4%
Interview Lift
resolved cases with interview
Typical timeline
4y 0m
Avg Prosecution
12 currently pending
Career history
49
Total Applications
across all art units

Statute-Specific Performance

§101
23.9%
-16.1% vs TC avg
§103
40.0%
+0.0% vs TC avg
§102
6.8%
-33.2% vs TC avg
§112
29.4%
-10.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 24 resolved cases

Office Action

§103 §112
DETAILED ACTION Status of Claims This Office action is responsive to communications filed on 2026-07-07. Claim(s) 1-20 is/are pending and are examined herein. Claim(s) 1-20 is/are objected to. Claim(s) 1-20 is/are rejected under 35 USC 112(b). Claim(s) 1-20 is/are rejected under 35 USC 103. Notice of Pre-AIA or AIA Status The present application, filed on or after 2013-03-16, is being examined under the first inventor to file provisions of the AIA . Response to Arguments Regarding the objections for informalities, the applicant’s amendments resolve the issues raised in the previous Office action but introduce new issues as described below. Regarding the rejections under 35 USC 112(b), the applicant’s amendments resolve some of the issues raised in the previous Office action. However, they do not adequately address the substance of at least one of the issues raised previously (e.g., the relationship of the “reduced period of time” to other claim elements remains unclear). They also introduce new issues (e.g., regarding the ambiguity regarding “full-training” as discussed during the interview of 2026-06-25). Issues in the pending claim are listed below. Regarding the rejections under 35 USC 103, the applicant’s arguments have been fully considered but they are not persuasive: The applicant asserts that “the cited references are silent on determining a plurality of correlation candidate models to evaluate a plurality of proxy task choices” [remarks, pages 10-11]. The examiner respectfully disagrees: Sinapov discloses a set T_source of source “tasks for which the agent has learned a policy” (with the goal being to “select a task T_i in T_{source} such that T_i serves as an effective source for learning T_j”, where T_j is a given target task) [Sinapov, section 4.1]. The source tasks in T_{source} map to the “plurality of proxy task choices” of the claim, and the policies that are learned for these tasks map to the “plurality of correlation candidate models” of the claim. The applicant asserts that “the cited references are silent on… generating correlation scores for the plurality of proxy task choices” [remarks, pages 10-11]. The examiner respectfully disagrees. Sinapov discloses computing hat{B}(T_i, T_j) which is “the expected benefit of transferring from T_i to T_j” [Sinapov, section 4.2]. The score hat{B}(T_i, T_j) maps to the “correlation score” for the task T_i, as required by the claim. The applicant quotes broad portions of the amended claim and then broadly asserts that “none of the references cited disclose, teach, or suggest” these features [remarks, page 11] without indicating any specific claim elements introduced in their amendment that are not disclosed by the prior art made of record. These broad and unsubstantiated assertions do not comply with 37 CFR 1.111(c) because they do not clearly point out the patentable novelty which the applicant thinks the claims present in view of the state of the art disclosed by the references cited. Further, they do not show how the amendments avoid such references. Moreover, the examiner respectfully disagrees with the applicant’s broad assertions. The substance of the amendments to the independent claims amounts to a rolling up of certain limitations from certain dependent claims, but these limitations are disclosed at least in Wang (as described in the previous Office action and again below). The complete prior art rejection, updated in view of the applicant’s amendments, is given below. Claim Objections Claim(s) 1-20 is/are objected to because of the following informalities: Claims 1, 11, and 17 recite with respect to correlation between proxy task training and full-training [emphasis added] but this should be “with respect to a correlation between a proxy task training and a full training” for grammaticality and proper punctuation. Similarly, the subsequent recitation of the correlation between the proxy task training and the full-training [emphasis added] should be “the correlation between the proxy task training and the full training” for proper punctuation and proper antecedent basis. The applicant is also invited to consult a related 112(b) rejection. Dependent claims 2-10, 12-16, and 18-20 inherit the objections. Claims 1, 11, and 17 recite until meeting a minimum number of correlation candidate models iteratively: [emphasis added] but there should be a comma after “models” for grammaticality and proper punctuation (i.e., “until meeting a minimum number of correlation candidate models, iteratively:”). Dependent claims 2-10, 12-16, and 18-20 inherit the objection. Appropriate correction is required. Claim Rejections - 35 USC 112(b) The following is a quotation of 35 USC 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 USC 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claim(s) 1-20 is/are rejected under 35 USC 112(b) or 35 USC 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 USC 112, the applicant), regards as the invention. Claims 1, 11, and 17 recite a correlation between proxy task training and full-training [emphasis added]. This is subjective language: a model can always be trained for longer to achieve better performance, so deciding that the training of a model has been fully completed involves a subjective judgment. MPEP 2173.05(b)(IV) indicates, in the presence of subjective claim language, an “objective standard must be provided [in the specification] in order to allow the public to determine the scope of the claim”. However, in the present instance, the specification does not provide any objective criterion for determining what training is a “full training”. For the purpose of compact prosecution, the claim is interpreted broadly as encompassing any training of a model performing the target task. Dependent claims 2-10, 12-16, and 18-20 inherit the rejection. Claims 9, 16, and 20 recite a reduced period of time of a total training time [emphasis added] but the meaning of this phrase is not clear for at least two reasons. First, it is not clear whether or not the recitation of “total training time” is intended to be bound in scope by the entity of the same name that is introduced in the parent claim. If it is, the claim should be “a reduced period of time of the total training time” for proper antecedent basis. However, even if this was the case, it would still not be clear what it means for the “reduced period of time” to be “of the total training time” since this language is questionably grammatical and does not adequately clarify whether the “reduced period of time” is intended to refer to any period of time that is shorter than the “total training time” or to something different. MPEP 2173.05(b) indicates that a “claim may be rendered indefinite when a limitation of the claim is defined by reference to an object and the relationship between the limitation and the object is not sufficiently defined” and, in the present instance, the relationship between the “reduced period of time” and the “total training time” is not sufficiently defined. Consequently, the claim is indefinite. For the purpose of compact prosecution, the claim is interpreted broadly as encompassing at least the situation where the “reduced period of time” is any period of time that is shorter than the “total training time”. Alternative language clarifying relationships between claim elements is advised. Dependent claim 10 inherits the rejection. Claim Rejections - 35 USC 103 The following is a quotation of 35 USC 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 USC 102(b)(2)(C) for any potential 35 USC 102(a)(2) prior art against the later invention. Claim(s) 1-6, 11-13, and 17-18 is/are rejected under 35 USC 103 as being unpatentable over Jivko SINAPOV et al. (Learning Inter-Task Transferability in the Absence of Target Task Samples, published 2015-05-04; hereafter, “Sinapov”) in view of Ruochen WANG et al. (RANK-NOSH: Efficient Predictor-Based Architecture Search via Non-Uniform Successive Halving, published 2021-08-18; hereafter, “Wang”). Claim 1 Sinapov discloses: A method for automatically determining a proxy task ([Sinapov, section 4.1]: Sinapov discloses a system whose goal, given a target task T_j in T_{target}, is to “select a task T_i in T_{source} such that T_i serves as an effective source for learning T_j” [Sinapov, section 4.1 second paragraph]. Selecting T_i in T_{source} maps to “determining a proxy task” as recited by the claim.) comprising: determining, by one or more processors, ([Sinapov, section 5.2]: Sinapov further discloses implementing the methods on a Condor Cluster system [Sinapov, section 5.2 paragraph beginning “All told”]. The processors in the cluster map to the “one or more processors” of the claim.) a plurality of correlation candidate models to evaluate respective proxy task choices of plurality of proxy task choices with respect to correlation between proxy task training and full-training ([Sinapov, section 4.1]: Sinapov discloses an agent learning a policy for each of the source tasks in T_{source} [Sinapov, section 4.1 first paragraph], and a function B where B(T_i, T_j) is “the value of transferring the policy learned in T_i to the task T_j” [Sinapov, section 4.1 second paragraph]. The set T_{source} of source tasks maps to the “plurality of proxy task choices” of the claim, and the policies learned for these tasks map to the “plurality of correlation candidate models” of the claim. Learning any of the policies for the tasks in T_{source} maps to the “proxy task training” of the claim, learning for the target task T_j maps to the “full-training” of the claim, and the function B maps to the “correlation” of the claim.) generating, by the one or more processors, a full-training score for each of the plurality of correlation candidate models; ([Sinapov, figure 2]: Sinapov discloses computing rewards after each episode of training [Sinapov, figure 2; see also, section 5.2 paragraph beginning “Varying”]. The reward for each source task after the last training episode for that task maps to the “full-training score for each of the plurality of correlation candidate models” of the claim.) generating, by the one or more processors, a correlation score for each of the plurality of proxy task choices using the plurality of correlation candidate models and the full-training scores, the correlation score indicating the correlation between the proxy task training and the full-training; ([Sinapov, sections 4.2 and 5.2]: Sinapov discloses that the system computes, for each T_i in T_{source}, the value “hat{B}(T_i, T_j), i.e., the expected benefit of transferring T_i to T_j” [Sinapov, section 4.2 first paragraph]. To estimate this benefit of transfer, the agent is trained on task T_j starting with the policy learned on task T_i [Sinapov, section 5.2 paragraph beginning “One the baseline curves”]. The quantity hat{B}(T_i, T_j) for each source task T_i maps to the “correlation score for each of the plurality of task choices” of the claim since it “indicates the correlation” as recited by the claim (with B mapping to the “correlation” of the claim, as noted above). Since estimating these quantities uses the policies for the source tasks, it “us[es] the plurality of correlation candidate models and the full-training scores” as recited by the claim, with the “plurality of correlation candidate models” and the “full-training scores” being mapped as described above.) ranking, by the one or more processors, the plurality of proxy task choices based on the correlation scores and training time; ([Sinapov, sections 4.3 and 5.2]: Sinapov discloses creating a ranked list R_j = [T_{{1}}, T_{{2}}, …, T_{{P}}] of source tasks according to the expected benefits, i.e., hat{B}(T_{{k}}, T_j) geq hat{B}(T_{{k+1}}, T_j} for all k [Sinapov, section 4.3.1 second paragraph]. In other words, the ranked list R_j is “based on the correlation scores” as mapped above. It is also based on “training time” since the estimates hat{B}(T_i, T_j) are computed after the baseline models for each task have been trained [Sinapov, section 5.2 paragraph beginning “Once the baseline curves”].) selecting, by the one or more processors, a proxy task choice of the plurality of proxy task choices based on the ranking; and outputting, by the one or more processors, instructions associated with the selected proxy task choice. ([Sinapov, section 4.3]: Sinapov indicates that the best possible source task is defined as T^* = argmax_{T_i in T_{source}} B(T_i, T_j). In other words, T^* = T_{{1}} maps to the “selected proxy task choice” of the claim. The use of T^* for transfer learning maps to the “instructions associated with the selected proxy task choice” of the claim.) While Sinapov discusses source task selection in the general context of transfer learning, it does not describe a specific application to neural architecture search. In other words, Sinapov does not distinctly disclose: for a neural architecture search… for the neural architecture search, wherein determining the plurality of correlation candidate models comprises, until meeting a minimum number of correlation candidate models iteratively: training a plurality of models for a fraction of a total training time for the neural architecture search; and rejecting a portion of the plurality of models which do not add to a score distribution for one or more metrics associated with the neural architecture search among the plurality of models; Wang is in the field of machine learning. Moreover, Sinapov in view of Wang discloses: for a neural architecture search… for the neural architecture search; ([Wang, abstract and algorithm 1]: Wang discloses a method of neural architecture search called “NOn-uniform Successive Halving (NOSH)” [Wang, abstract] which includes training steps [Wang, algorithm 1; see also, section 3.2 paragraph beginning “Initialization”]. In the combination, the source task selection procedure of Sinapov is used to determine the training tasks used for the neural architecture search of Wang. The applicant is also invited to consult YLi and Zoph as cited in the conclusion of a previous Office action.) wherein determining the plurality of correlation candidate models comprises, until meeting a minimum number of correlation candidate models iteratively: training a plurality of models for a fraction of a total training time for the neural architecture search; ([Wang, sections 3.2 and 3.4, algorithm 1, and figure 2]: Wang discloses “initializ[ing] the pool by randomly sampling K_{init} architectures from the search space” [Wang, section 3.4 first paragraph]. This pool forms part of the input to the algorithm [Wang, algorithm 1] alongside the “schedule E = {e^{(l)}}_{l = 1}^N [which] represents the training epoch for every architecture at each level, where e^{(i)} < e^{(i+1)}, i = 1 ~ (N – 1), and e^{(N)} is the maximum number of epochs (fully trained)” [Wang, section 3.2 paragraph beginning “We introduce”; see also, algorithm 1]. The initial pool of K_{init} architectures is trained for e^{(1)} epochs, and then “architectures with the validation accuracy in the bottom K_{init}(1-r) will be terminated (kept in level 1), while the top K_{init}r architectures will be trained further to e^{(2)} epochs and upgrade to level 2” [Wang, section 3.2 paragraph beginning “Initialization”; see also, algorithm 1 line 7 and figure 2 leftmost pyramid]. This “process repeats until the maximum training epoch e^{(N)} is reached” [Wang, section 3.2 paragraph beginning “Initialization”; see also, algorithm 1 and figure 2 leftmost pyramid]. In other words, the initial pool of K_{init} architectures maps to the “plurality of models” of the claim, the maximum number e^{(N)} in the schedule maps to the “total training time” of the claim, the first number e^{(1)} in the schedule maps to the “fraction” of the total training time of the claim, and the K_{init}r^{N-1} candidates that reach level-N map to the “minimum number of correlation candidate models” of the claim.) and rejecting a portion of the plurality of models which do not add to a score distribution for one or more metrics associated with the neural architecture search among the plurality of models; ([Wang, section 3.2 and algorithm 1]: As noted above, Wang discloses that “architectures with the validation accuracy in the bottom K_{init}(1-r) will be terminated (kept in level 1), while the top K_{init}r architectures will be trained further to e^{(2)} epochs and upgrade to level 2” [Wang, section 3.2 paragraph beginning “Initialization”]. In other words, the bottom K_{init}(1-r) architectures of the first iteration map to the “portion of the plurality of models” of the claims (i.e., the ones that are “reject[ed]” as recited by the claim). The validation accuracies map to the “one or metrics” of the claim.) Before the effective filing date of the invention, it would have been obvious to a person of ordinary skill in the art to combine the source/proxy task selection method of Sinapov with the neural architecture search method of Wang because the latter “reduces the search budget by ~5x while achieving competitive or even better than previous state-of-the-art predictor based methods” [Wang, abstract], thereby resulting in an efficient system overall. Claim 2 Sinapov in view of Wang discloses the elements of the parent claim(s). It also discloses: [The method of claim 1, further comprising] receiving, by the one or more processors, the plurality of proxy task choices for the neural architecture search. ([Sinapov, section 4.1]: As noted above, the set T_{source} of source tasks maps to the “plurality of proxy task choices” of the claim.) The same motivation to combine applies. Claim 3 Sinapov in view of Wang discloses the elements of the parent claim(s). It also discloses: [The method of claim 1, wherein determining the plurality of correlation candidate models further comprises:] randomly sampling a first plurality of models from a search space for the neural architecture search; ([Wang, section 3.4]: As noted under the parent claim, Wang discloses “initializ[ing] the pool by randomly sampling K_{init} architectures from the search space” [Wang, section 3.4 first paragraph]. The pool of K_{init} architectures from the search space map to the “first plurality” of the claim.) training each of the first plurality of models for a first fraction of the total training time for the neural architecture search; ([Wang, section 3.2 and algorithm 1]: As noted under the parent claim, each architecture in K_{init} is trained for e^{(1)} epochs, with some of the architectures eventually reaching e^{(N)} epochs [Wang, section 3.2 and algorithm 1], and this maximum number e^{(N)} of epochs maps to the “total training time” of the claim. The first number e^{(1)} in the schedule maps to the “first fraction” of the claim.) and rejecting a first portion of the first plurality of models which do not add to a score distribution for the one or more metrics, wherein a second plurality of models corresponds to the first portion subtracted from the first plurality of models. ([Wang, section 3.2 and algorithm 1]: As noted under the parent claim, Wang discloses that “architectures in with the validation accuracy in the bottom K_{init}(1-r) will be terminated (kept in level 1), while the top K_{init}r architectures will be trained further to e^{(2)} epochs and upgrade to level 2” [Wang, section 3.2 paragraph beginning “Initialization”]. The bottom K_{init}(1-r) map to the “first portion” of the claim, and the top K_{init}r map to the “second plurality” of the claim.) The same motivation to combine applies. Claim 4 Sinapov in view of Wang discloses the elements of the parent claim(s). It also discloses: [The method of claim 3, wherein determining the plurality of correlation candidate models further comprises:] training each of the second plurality of models for a second fraction of the total training time for the neural architecture search, ([Wang, section 3.2 and algorithm 1]: As noted above, Wang discloses that “the top K_{init}r architectures will be trained further to e^{(2)} epochs” [Wang, section 3.2 paragraph beginning “Initialization”]. In other words, e^{(2)} maps to the “second fraction” of the full training time of the claim.) the second fraction being greater than the first fraction; ([Wang, section 3.2]: As noted above, Wang discloses that e^{(i)} < e^{(i+1)} for all i [Wang, section 3.2 paragraph beginning “We introduce”]. Since e^{(1)} maps to the “first fraction” and e^{(2)} to the “second fraction” of the claim, it is indeed the case that the “second fraction [is] greater than the first fraction” as required by the claim.) and rejecting a second portion of the second plurality of models which do not add to a score distribution for the one or more metrics. ([Wang, section 3.2 and figure 2]: Wang indicates that the “process repeats until the maximum training epoch e^{(N)} is reached” [Wang, section 3.2 paragraph beginning “Initialization”; see also, figure 2 leftmost pyramid]. In other words, the K_{init}r(1-r) architectures which terminate at level 2 map to the “second portion” of the claim.) The same motivation to combine applies. Claim 5 Sinapov in view of Wang discloses the elements of the parent claim(s). It also discloses: [The method of claim 4, wherein] the first fraction and the second fraction respectively correspond to a number of models in the first plurality of models and the second plurality of models. ([Wang, section 3.2]: As noted above, e^{(1)} maps to the “first fraction” of the claim, and it “corresponds to a number of models in the first plurality of models” since it represents the minimum number of epochs to which all models in the “first plurality” are trained. Similarly, e^{(2)} maps to the “second fraction” of the claim and it “corresponds to a number of models in the… second plurality of models” since it represents the minimum number of epochs to which all models in the “second plurality” are trained.) The same motivation to combine applies. Claim 6 Sinapov in view of Wang discloses the elements of the parent claim(s). [The method of claim 1, wherein generating the correlation score for each of the plurality of proxy task choices further comprises:] training each of the plurality of correlation candidate models; ([Sinapov, section 4.1]: As noted under the parent claim, the agent learns a policy for each source task in T_{source} [Sinapov, section 4.1 first paragraph], with the policies mapping to the “plurality of correlation candidate models” of the claim. The learning of these policies maps to the “training” of the claim.) and during the training: monitoring one or more metrics and training time; ([Sinapov, figure 1 and section 5.2]: Sinapov discloses 2500 training episodes for each task [Sinapov, section 5.2 paragraph beginning “Varying”] and it tracks rewards as the number of training episodes changes [Sinapov, figure 1]. In other words, keeping track of present reward maps to “monitoring one or more metrics” as recited by the claim, and keeping track of the training episode counter maps to “monitoring… training time” as recited by the claim. Both of these are performed “during the training” as required by the claim.) and continuously computing the correlation score based on the full-training scores. ([Sinapov, sections 4.2 and 5.2, and figure 3]: As noted under the parent claim, an expected benefit hat{B}(T_i, T_j) maps to a “correlation-score” of the claim, and these are “based on the full-training scores” because their computation uses the models trained on the source tasks. Sinapov also shows computing the expected benefit of transfer hat{B}(T_i, T_j) “continuously” as training progresses [Sinapov, figure 3].) The same motivation to combine applies. Claim 11 Sinapov discloses: A system comprising: one or more processors; and one or more storage devices coupled to the one or more processors and storing instructions that, when executed by the one or more processors, cause the one or more processors to perform operations for automatically ([Sinapov, section 5.2]: Sinapov discloses that the methods disclosed therein are implemented on a Condor Cluster system [Sinapov, section 5.2 paragraph beginning “All told”]. The cluster maps to the “system” of the claim, the processors in the cluster map to the “one or more processors” of the claim, and memory or hard drives in the cluster to the “one or more storage devices” of the claim.) determining a proxy task ([Sinapov, section 4.1]: Sinapov discloses a system whose goal, given a target task T_j in T_{target}, is to “select a task T_i in T_{source} such that T_i serves as an effective source for learning T_j” [Sinapov, section 4.1 second paragraph]. Selecting T_i in T_{source} maps to “determining a proxy task” as recited by the claim.) the operations comprising: determining a plurality of correlation candidate models to evaluate respective proxy task choices of plurality of proxy task choices with respect to correlation between proxy task training and full-training ([Sinapov, section 4.1]: Sinapov discloses an agent learning a policy for each of the source tasks in T_{source} [Sinapov, section 4.1 first paragraph], and a function B where B(T_i, T_j) is “the value of transferring the policy learned in T_i to the task T_j” [Sinapov, section 4.1 second paragraph]. The set T_{source} of source tasks maps to the “plurality of proxy task choices” of the claim, and the policies learned for these tasks map to the “plurality of correlation candidate models” of the claim. Learning any of the policies for the tasks in T_{source} maps to the “proxy task training” of the claim, learning for the target task T_j maps to the “full-training” of the claim, and the function B maps to the “correlation” of the claim.) generating a full-training score for each of the plurality of correlation candidate models; ([Sinapov, figure 2]: Sinapov discloses computing rewards after each episode of training [Sinapov, figure 2; see also, section 5.2 paragraph beginning “Varying”]. The reward for each source task after the last training episode for that task maps to the “full-training score for each of the plurality of correlation candidate models” of the claim.) generating a correlation score for each of the plurality of proxy task choices using the plurality of correlation candidate models and the full-training scores, the correlation score indicating the correlation between the proxy task training and the full-training; ([Sinapov, sections 4.2 and 5.2]: Sinapov discloses that the system computes, for each T_i in T_{source}, the value “hat{B}(T_i, T_j), i.e., the expected benefit of transferring T_i to T_j” [Sinapov, section 4.2 first paragraph]. To estimate this benefit of transfer, the agent is trained on task T_j starting with the policy learned on task T_i [Sinapov, section 5.2 paragraph beginning “One the baseline curves”]. The quantity hat{B}(T_i, T_j) for each source task T_i maps to the “correlation score for each of the plurality of task choices” of the claim since it “indicates the correlation” as recited by the claim (with B mapping to the “correlation” of the claim, as noted above). Since estimating these quantities uses the policies for the source tasks, it “us[es] the plurality of correlation candidate models and the full-training scores” as recited by the claim, with the “plurality of correlation candidate models” and the “full-training scores” being mapped as described above.) ranking the plurality of proxy task choices based on the correlation scores and training time; ([Sinapov, sections 4.3 and 5.2]: Sinapov discloses creating a ranked list R_j = [T_{{1}}, T_{{2}}, …, T_{{P}}] of source tasks according to the expected benefits, i.e., hat{B}(T_{{k}}, T_j) geq hat{B}(T_{{k+1}}, T_j} for all k [Sinapov, section 4.3.1 second paragraph]. In other words, the ranked list R_j is “based on the correlation scores” as mapped above. It is also based on “training time” since the estimates hat{B}(T_i, T_j) are computed after the baseline models for each task have been trained [Sinapov, section 5.2 paragraph beginning “Once the baseline curves”].) selecting a proxy task choice of the plurality of proxy task choices based on the ranking; and outputting instructions associated with the selected proxy task choice. ([Sinapov, section 4.3]: Sinapov indicates that the best possible source task is defined as T^* = argmax_{T_i in T_{source}} B(T_i, T_j). In other words, T^* = T_{{1}} maps to the “selected proxy task choice” of the claim. The use of T^* for transfer learning maps to the “instructions associated with the selected proxy task choice” of the claim.) While Sinapov discusses source task selection in the general context of transfer learning, it does not describe a specific application to neural architecture search. In other words, Sinapov does not distinctly disclose: for a neural architecture search… for the neural architecture search, wherein determining the plurality of correlation candidate models comprises, until meeting a minimum number of correlation candidate models iteratively: training a plurality of models for a fraction of a total training time for the neural architecture search; and rejecting a portion of the plurality of models which do not add to a score distribution for one or more metrics associated with the neural architecture search among the plurality of models; Wang is in the field of machine learning. Moreover, Sinapov in view of Wang discloses: for a neural architecture search… for the neural architecture search; ([Wang, abstract and algorithm 1]: Wang discloses a method of neural architecture search called “NOn-uniform Successive Halving (NOSH)” [Wang, abstract] which includes training steps [Wang, algorithm 1; see also, section 3.2 paragraph beginning “Initialization”]. In the combination, the source task selection procedure of Sinapov is used to determine the training tasks used for the neural architecture search of Wang. The applicant is also invited to consult YLi and Zoph as cited in the conclusion of a previous Office action.) wherein determining the plurality of correlation candidate models comprises, until meeting a minimum number of correlation candidate models iteratively: training a plurality of models for a fraction of a total training time for the neural architecture search; ([Wang, sections 3.2 and 3.4, algorithm 1, and figure 2]: Wang discloses “initializ[ing] the pool by randomly sampling K_{init} architectures from the search space” [Wang, section 3.4 first paragraph]. This pool forms part of the input to the algorithm [Wang, algorithm 1] alongside the “schedule E = {e^{(l)}}_{l = 1}^N [which] represents the training epoch for every architecture at each level, where e^{(i)} < e^{(i+1)}, i = 1 ~ (N – 1), and e^{(N)} is the maximum number of epochs (fully trained)” [Wang, section 3.2 paragraph beginning “We introduce”; see also, algorithm 1]. The initial pool of K_{init} architectures is trained for e^{(1)} epochs, and then “architectures with the validation accuracy in the bottom K_{init}(1-r) will be terminated (kept in level 1), while the top K_{init}r architectures will be trained further to e^{(2)} epochs and upgrade to level 2” [Wang, section 3.2 paragraph beginning “Initialization”; see also, algorithm 1 line 7 and figure 2 leftmost pyramid]. This “process repeats until the maximum training epoch e^{(N)} is reached” [Wang, section 3.2 paragraph beginning “Initialization”; see also, algorithm 1 and figure 2 leftmost pyramid]. In other words, the initial pool of K_{init} architectures maps to the “plurality of models” of the claim, the maximum number e^{(N)} in the schedule maps to the “total training time” of the claim, the first number e^{(1)} in the schedule maps to the “fraction” of the total training time of the claim, and the K_{init}r^{N-1} candidates that reach level-N map to the “minimum number of correlation candidate models” of the claim.) and rejecting a portion of the plurality of models which do not add to a score distribution for one or more metrics associated with the neural architecture search among the plurality of models; ([Wang, section 3.2 and algorithm 1]: As noted above, Wang discloses that “architectures with the validation accuracy in the bottom K_{init}(1-r) will be terminated (kept in level 1), while the top K_{init}r architectures will be trained further to e^{(2)} epochs and upgrade to level 2” [Wang, section 3.2 paragraph beginning “Initialization”]. In other words, the bottom K_{init}(1-r) architectures of the first iteration map to the “portion of the plurality of models” of the claims (i.e., the ones that are “reject[ed]” as recited by the claim). The validation accuracies map to the “one or metrics” of the claim.) Before the effective filing date of the invention, it would have been obvious to a person of ordinary skill in the art to combine the source/proxy task selection method of Sinapov with the neural architecture search method of Wang because the latter “reduces the search budget by ~5x while achieving competitive or even better than previous state-of-the-art predictor based methods” [Wang, abstract], thereby resulting in an efficient system overall. Claims 12-13 inherit limitations from claim 11 and recite additional limitations which are substantially similar to those recited by claims 3 and 5, respectively (with claim 13 requiring only one of the two limitations required in claim 5), so they are rejected by the same rationale. Claim 17 Sinapov discloses: A non-transitory computer readable medium for storing instructions that, when executed by one or more processors, cause the one or more processors to perform operations for automatically ([Sinapov, section 5.2]: Sinapov discloses that the methods disclosed therein are implemented on a Condor Cluster system [Sinapov, section 5.2 paragraph beginning “All told”]. The processors in the cluster map to the “one or more processors” of the claim, and any hard drive in the cluster to the “non-transitory computer readable storage medium” of the claim.) determining a proxy task ([Sinapov, section 4.1]: Sinapov discloses a system whose goal, given a target task T_j in T_{target}, is to “select a task T_i in T_{source} such that T_i serves as an effective source for learning T_j” [Sinapov, section 4.1 second paragraph]. Selecting T_i in T_{source} maps to “determining a proxy task” as recited by the claim.) the operations comprising: determining a plurality of correlation candidate models to evaluate respective proxy task choices of plurality of proxy task choices with respect to correlation between proxy task training and full-training ([Sinapov, section 4.1]: Sinapov discloses an agent learning a policy for each of the source tasks in T_{source} [Sinapov, section 4.1 first paragraph], and a function B where B(T_i, T_j) is “the value of transferring the policy learned in T_i to the task T_j” [Sinapov, section 4.1 second paragraph]. The set T_{source} of source tasks maps to the “plurality of proxy task choices” of the claim, and the policies learned for these tasks map to the “plurality of correlation candidate models” of the claim. Learning any of the policies for the tasks in T_{source} maps to the “proxy task training” of the claim, learning for the target task T_j maps to the “full-training” of the claim, and the function B maps to the “correlation” of the claim.) generating a full-training score for each of the plurality of correlation candidate models; ([Sinapov, figure 2]: Sinapov discloses computing rewards after each episode of training [Sinapov, figure 2; see also, section 5.2 paragraph beginning “Varying”]. The reward for each source task after the last training episode for that task maps to the “full-training score for each of the plurality of correlation candidate models” of the claim.) generating a correlation score for each of the plurality of proxy task choices using the plurality of correlation candidate models and the full-training scores, the correlation score indicating the correlation between the proxy task training and the full-training; ([Sinapov, sections 4.2 and 5.2]: Sinapov discloses that the system computes, for each T_i in T_{source}, the value “hat{B}(T_i, T_j), i.e., the expected benefit of transferring T_i to T_j” [Sinapov, section 4.2 first paragraph]. To estimate this benefit of transfer, the agent is trained on task T_j starting with the policy learned on task T_i [Sinapov, section 5.2 paragraph beginning “One the baseline curves”]. The quantity hat{B}(T_i, T_j) for each source task T_i maps to the “correlation score for each of the plurality of task choices” of the claim since it “indicates the correlation” as recited by the claim (with B mapping to the “correlation” of the claim, as noted above). Since estimating these quantities uses the policies for the source tasks, it “us[es] the plurality of correlation candidate models and the full-training scores” as recited by the claim, with the “plurality of correlation candidate models” and the “full-training scores” being mapped as described above.) ranking the plurality of proxy task choices based on the correlation scores and training time; ([Sinapov, sections 4.3 and 5.2]: Sinapov discloses creating a ranked list R_j = [T_{{1}}, T_{{2}}, …, T_{{P}}] of source tasks according to the expected benefits, i.e., hat{B}(T_{{k}}, T_j) geq hat{B}(T_{{k+1}}, T_j} for all k [Sinapov, section 4.3.1 second paragraph]. In other words, the ranked list R_j is “based on the correlation scores” as mapped above. It is also based on “training time” since the estimates hat{B}(T_i, T_j) are computed after the baseline models for each task have been trained [Sinapov, section 5.2 paragraph beginning “Once the baseline curves”].) selecting a proxy task choice of the plurality of proxy task choices based on the ranking; and outputting instructions associated with the selected proxy task choice. ([Sinapov, section 4.3]: Sinapov indicates that the best possible source task is defined as T^* = argmax_{T_i in T_{source}} B(T_i, T_j). In other words, T^* = T_{{1}} maps to the “selected proxy task choice” of the claim. The use of T^* for transfer learning maps to the “instructions associated with the selected proxy task choice” of the claim.) While Sinapov discusses source task selection in the general context of transfer learning, it does not describe a specific application to neural architecture search. In other words, Sinapov does not distinctly disclose: for a neural architecture search… for the neural architecture search, wherein determining the plurality of correlation candidate models comprises, until meeting a minimum number of correlation candidate models iteratively: training a plurality of models for a fraction of a total training time for the neural architecture search; and rejecting a portion of the plurality of models which do not add to a score distribution for one or more metrics associated with the neural architecture search among the plurality of models; Wang is in the field of machine learning. Moreover, Sinapov in view of Wang discloses: for a neural architecture search… for the neural architecture search; ([Wang, abstract and algorithm 1]: Wang discloses a method of neural architecture search called “NOn-uniform Successive Halving (NOSH)” [Wang, abstract] which includes training steps [Wang, algorithm 1; see also, section 3.2 paragraph beginning “Initialization”]. In the combination, the source task selection procedure of Sinapov is used to determine the training tasks used for the neural architecture search of Wang. The applicant is also invited to consult YLi and Zoph as cited in the conclusion of a previous Office action.) wherein determining the plurality of correlation candidate models comprises, until meeting a minimum number of correlation candidate models iteratively: training a plurality of models for a fraction of a total training time for the neural architecture search; ([Wang, sections 3.2 and 3.4, algorithm 1, and figure 2]: Wang discloses “initializ[ing] the pool by randomly sampling K_{init} architectures from the search space” [Wang, section 3.4 first paragraph]. This pool forms part of the input to the algorithm [Wang, algorithm 1] alongside the “schedule E = {e^{(l)}}_{l = 1}^N [which] represents the training epoch for every architecture at each level, where e^{(i)} < e^{(i+1)}, i = 1 ~ (N – 1), and e^{(N)} is the maximum number of epochs (fully trained)” [Wang, section 3.2 paragraph beginning “We introduce”; see also, algorithm 1]. The initial pool of K_{init} architectures is trained for e^{(1)} epochs, and then “architectures with the validation accuracy in the bottom K_{init}(1-r) will be terminated (kept in level 1), while the top K_{init}r architectures will be trained further to e^{(2)} epochs and upgrade to level 2” [Wang, section 3.2 paragraph beginning “Initialization”; see also, algorithm 1 line 7 and figure 2 leftmost pyramid]. This “process repeats until the maximum training epoch e^{(N)} is reached” [Wang, section 3.2 paragraph beginning “Initialization”; see also, algorithm 1 and figure 2 leftmost pyramid]. In other words, the initial pool of K_{init} architectures maps to the “plurality of models” of the claim, the maximum number e^{(N)} in the schedule maps to the “total training time” of the claim, the first number e^{(1)} in the schedule maps to the “fraction” of the total training time of the claim, and the K_{init}r^{N-1} candidates that reach level-N map to the “minimum number of correlation candidate models” of the claim.) and rejecting a portion of the plurality of models which do not add to a score distribution for one or more metrics associated with the neural architecture search among the plurality of models; ([Wang, section 3.2 and algorithm 1]: As noted above, Wang discloses that “architectures with the validation accuracy in the bottom K_{init}(1-r) will be terminated (kept in level 1), while the top K_{init}r architectures will be trained further to e^{(2)} epochs and upgrade to level 2” [Wang, section 3.2 paragraph beginning “Initialization”]. In other words, the bottom K_{init}(1-r) architectures of the first iteration map to the “portion of the plurality of models” of the claims (i.e., the ones that are “reject[ed]” as recited by the claim). The validation accuracies map to the “one or metrics” of the claim.) Before the effective filing date of the invention, it would have been obvious to a person of ordinary skill in the art to combine the source/proxy task selection method of Sinapov with the neural architecture search method of Wang because the latter “reduces the search budget by ~5x while achieving competitive or even better than previous state-of-the-art predictor based methods” [Wang, abstract], thereby resulting in an efficient system overall. Claim 18 inherits limitations from claim 11 and recites additional limitations which are substantially similar to those recited by claim 3, so it is rejected by the same rationale. Claim(s) 7-8 and 14-15, and 19 is/are rejected under 35 USC 103 as being unpatentable over Sinapov in view of Wang, further in view of Benjamin WILSON et al. (US20230244982A1, effectively filed 2022-01-28; hereafter, “Wilson”). Claim 7 Sinapov in view of Wang discloses the elements of the parent claim(s). It might not distinctly disclose: [The method of claim 6, wherein generating the correlation scores for each of the plurality of proxy task choices further comprises] stopping the training based on obtaining a threshold correlation score. Wilson is in the field of machine learning. Moreover, Sinapov in view of Wang and Wilson discloses: [The method of claim 6, wherein generating the correlation scores for each of the plurality of proxy task choices further comprises] stopping the training based on obtaining a threshold correlation score. ([Wilson, 0076]: Wilson discloses making a selection of a model based on predetermined criteria, such as the criterion of “a speed by which a model provides… a prediction that satisfies a minimum threshold” [Wilson, 0076]. In the combination, the models of Wilson are the models trained on source tasks on which transfer learning is performed, as disclosed in Sinapov, and the threshold of Wilson is taken to be the “threshold correlation score” of the claim.) Before the effective filing date of the invention, it would have been obvious to a person of ordinary skill in the art to combine the method of source task selection for neural architecture search as disclosed by Sinapov in view of Wang with the selection of source tasks which quickly achieve threshold performance as described in Wilson because this would ensure that the system runs quickly. Claim 8 Sinapov in view of Wang and Wilson discloses the elements of the parent claim(s). It also discloses: [The method of claim 7, wherein selecting the proxy task choice of the plurality of proxy tasks further comprises] selecting a proxy task choice that obtained the threshold correlation score within a shortest amount of time. ([Wilson, 0076]: As noted above, Wang discloses making a selection of a model based on predetermined criteria, such as the criterion of “a speed by which a model provides… a prediction that satisfies a minimum threshold” [Wilson, 0076]. In the combination, the models of Wilson are the models trained on source tasks on which transfer learning is performed, as disclosed in Sinapov, and the threshold of Wilson is taken to be the “threshold correlation score” of the claim. With these mappings, making a selection of proxy task based on the speed by which a prediction satisfying a threshold is provided maps to the step of “selecting a proxy task choice” as recited by the claim.) The same motivation to combine applies. Claims 14 and 15 inherits limitations from claim 11 and recites additional limitations which are substantially similar to those recited by claims 6-7 and 8, respectively (with claim 14 combining limitations that appear in claims 6-7), so they are rejected by the same rationale. Claims 19 inherits limitations from claim 11 and recites additional limitations which are substantially similar to those recited by 6-7, so it is rejected by the same rationale. Claim(s) 9, 16, and 20 is/are rejected under 35 USC 103 as being unpatentable over Sinapov in view of Wang, further in view of Tomoyuki OKUDA et al. (Non-parametric Prediction Interval Estimate for Uncertainty Quantification of the Prediction of Road Pavement Deterioration, published 2018; hereafter, “Okuda”). Claim 9 Sinapov in view of Wang discloses the elements of the parent claim(s). It also discloses: [The method of claim 1, further comprising:] randomly sampling, by the one or more processors, the search space to find a model for testing variance ([Wang, section 3.4]: Wang discloses “randomly sampling K_{init} architectures from the search space” [Wang, section 3.4 first paragraph]. Any one of the K_{init} architectures thus sampled maps to the “model” of the claim. The examiner notes that “for testing variance” is a field of use limitation which is disclosed by the combination as proposed below.) The same motivation to combine applies. Sinapov does not distinctly disclose performing a bootstrapping procedure on a model. In other words, Sinapov in view of Wang does not distinctly disclose: running, by the one or more processors, training for a plurality of copies of the model for a reduced period of time of a total training time; and measuring, by the one or more processors, at least one of a score variance or smoothness of the plurality of copies of the model. Okuda is in the field of machine learning. It discloses a bootstrapping procedure for neural networks [Okuda, section III.B]. In particular, Sinapov in view of Wang and Okuda discloses: running, by the one or more processors, training for a plurality of copies of the model for a reduced period of time of a total training time; ([Okuda, sections III.B and IV.D]: Okuda discloses resampling the training set B times to produce B bootstrap samples [Okuda, section III.B part 2)] and, “[f]or each of the B [neural network] models, n_{ae} epochs are learned using each bootstrap sample” [Okuda, section III.B part 3)]. In other words, the B learning steps of [Okuda, section III.B part 3)] map to the “training for a plurality of copies of the model” of the claim. The examiner notes that, in the specific examples described in Okuda, n_{ae} is either 100 or 10 and is, in particular, less than (i.e., “reduced” relative to) the 400 training epochs used for the original model [Okuda, section IV.D].) and measuring, by the one or more processors, at least one of a score variance or smoothness of the plurality of copies of the model. ([Okuda, section III.B]: Okuda discloses computing a confidence interval using the bootstrapped models [Okuda, section III.B part 4)]. This confidence interval (or, alternatively, its width) falls under the broadest reasonable interpretation of a “score variance or smoothness” as recited by the claim.) Before the effective filing date of the invention, it would have been obvious to a person of ordinary skill in the art to combine the method of source task selection for neural architecture search as disclosed by Sinapov in view of Wang with the bootstrapping method disclosed by Okuda because it decreases the computational cost “to about 1/38 that of the usual bootstrap method” and because it avoids overestimating [Okuda, section 1 last paragraph], thereby resulting in an effective and efficient method of simulating the distribution of an estimator. Claims 16 and 20 inherit limitations from claims 11 and 17, respectively, and recite additional limitations which are substantially similar to those recited by claims 9, so they are rejected by the same rationale. Claim(s) 10 is/are rejected under 35 USC 103 as being unpatentable over Sinapov in view of Wang and Okuda, further in view of Matthew BJONERUD et al. (US20190102835A1, published 2019-04-04; hereafter, “Bjonerud”) Claim 10 Sinapov in view of Wang and Okuda discloses the elements of the parent claim(s). It does not distinctly disclose: [The method of claim 9, further comprising:] determining, by the one or more processors, that the score variance or smoothness meets a threshold; and outputting one or more instructions to lower the score variance or smoothness. Bjonerud is in the field of machine learning. Moreover, Sinapov in view of Wang, Okuda, and Bjonerud discloses: [The method of claim 9, further comprising:] determining, by the one or more processors, that the score variance or smoothness meets a threshold; and outputting one or more instructions to lower the score variance or smoothness. ([Bjonerud, 0134]: Bjonerud discloses that if “the variance exceeds the variance thresholds in input 710 then artificial intelligence system 730 modifies its decision criteria to reduce the variance” [Bjonerud, 0134]. In the combination, the input of Bjonerud corresponds to the intervals produced by the bootstrapping method disclosed in Okuda. The variance thresholds of Bjonerud map to the “threshold” of the claim, and modification of decision criteria to reduce variance maps to the “instructions to lower the score variance or smoothness” of the claim.) Before the effective filing date of the invention, it would have been obvious to a person of ordinary skill in the art to combine the method of source task selection for neural architecture search as disclosed by Sinapov in view of Wang and Okuda with reduction of variances as described by Bjonerud because reducing variances means having more certainty in model estimates, thereby resulting in a more robust system. Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to Shishir AGRAWAL whose telephone number is +1 703-756-1183. The examiner can normally be reached Monday through Thursday, 08:30-14:30 Pacific Time. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Alexey SHMATOV can be reached on +1 571-270-3428. The fax phone number for the organization where this application or proceeding is assigned is +1 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at +1 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call +1 800-786-9199 (IN USA OR CANADA) or +1 571-272-1000. /S.A./Examiner, Art Unit 2123 /ALEXEY SHMATOV/Supervisory Patent Examiner, Art Unit 2123
Read full office action

Prosecution Timeline

Feb 23, 2023
Application Filed
Apr 23, 2026
Non-Final Rejection mailed — §103, §112
Jun 15, 2026
Interview Requested
Jun 25, 2026
Examiner Interview Summary
Jun 25, 2026
Applicant Interview (Telephonic)
Jul 07, 2026
Response Filed
Sep 09, 2026
Final Rejection mailed — §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12725051
RANKING DATA SLICES USING MEASURES OF INTEREST
4y 7m to grant Granted Sep 01, 2026
Study what changed to get past this examiner. Based on 1 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
8%
Grant Probability
24%
With Interview (+15.4%)
4y 0m (~4m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 24 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month