Prosecution Insights
Last updated: August 17, 2026
Application No. 18/504,117

TRAINING-FREE ARCHITECTURE SEARCH FOR EFFICIENT VISION TRANSFORMERS VIA APPROXIMATED ATTENTION STATISTICS

Non-Final OA §101§103§112
Filed
Nov 07, 2023
Examiner
ALSHAHARI, SADIK AHMED
Art Unit
2121
Tech Center
2100 — Computer Architecture & Software
Assignee
Qualcomm Incorporated
OA Round
1 (Non-Final)
38%
Grant Probability
At Risk
1-2
OA Rounds
1y 8m
Est. Remaining
79%
With Interview

Examiner Intelligence

Grants only 38% of cases
38%
Career Allowance Rate
17 granted / 45 resolved
-17.2% vs TC avg
Strong +41% interview lift
Without
With
+41.3%
Interview Lift
resolved cases with interview
Typical timeline
4y 5m
Avg Prosecution
17 currently pending
Career history
64
Total Applications
across all art units

Statute-Specific Performance

§101
29.5%
-10.5% vs TC avg
§103
45.0%
+5.0% vs TC avg
§102
5.7%
-34.3% vs TC avg
§112
16.5%
-23.5% vs TC avg
Black line = Tech Center average estimate • Based on career data from 45 resolved cases

Office Action

§101 §103 §112
DETAILED ACTION Status of Claims Claim(s) 1-28 are pending and are examined herein. Claim(s) 1-28 are rejected under 35 U.S.C. §§ 101 and 103. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Information Disclosure Statement The information disclosure statement IDS(s) submitted on February 06, 2024 and March 03, 2025 are in compliance with the provisions of 37 CFR 1.97 and have been considered by the examiner. Claim Interpretation The following is a quotation of 35 U.S.C. 112(f): (f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph: An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked. As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph: (A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function; (B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and (C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function. Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function. Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function. Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitation(s) are: means found in claims [22-26]. Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof. If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. When considering subject matter eligibility under 35 U.S.C. 101, it must be determined whether the claim is directed to one of the four statutory categories of invention, i.e., process, machine, manufacture, or composition of matter (Step 1). If the claim does fall within one of the statutory categories, the second step in the analysis is to determine whether the claim is directed to a judicial exception (Step 2A). The Step 2A analysis is broken into two prongs. In the first prong (Step 2A, Prong 1), it is determined whether or not the claims recite a judicial exception (e.g., mathematical concepts, mental processes, certain methods of organizing human activity). If it is determined in Step 2A, Prong 1 that the claims recite a judicial exception, the analysis proceeds to the second prong (Step 2A, Prong 2), where it is determined whether or not the claims integrate the judicial exception into a practical application. If it is determined at step 2A, Prong 2 that the claims do not integrate the judicial exception into a practical application, the analysis proceeds to determining whether the claim is a patent-eligible application of the exception (Step 2B). If an abstract idea is present in the claim, any element or combination of elements in the claim must be sufficient to ensure that the claim integrates the judicial exception into a practical application, or else amounts to significantly more than the abstract idea itself. Applicant is advised to consult MPEP 2106 for more details of the analysis. Under Step 1 analysis, Claims 1-7 recite an apparatus (representing a machine); Claims 8-14 recite a processor-implemented (representing a process); Claims 15-21 recite a non-transitory storage medium (representing an article of manufacture); Claims 22-28 recite an apparatus (representing a machine). Therefore, each set of the claims falls into one of the four statutory categories (i.e., process, machine, article of manufacture, or composition of matter). Claim(s) 1-28 are rejected under 35 U.S.C. 101 because the claimed invention is directed to a judicial exception (i.e., a law of nature, a natural phenomenon, or an abstract idea) without significantly more, and hence is not patent-eligible subject matter. Regarding Claim 1, Step 2A Prong 1: The claim recites an abstract idea enumerated in the 2019 PEG. generate a set of transformer model candidates for a target device, each transformer model candidate of the set of transformer model candidates being initialized with random weights; (An abstract idea of a mental process. Examiner’s note: the “generating” step, as drafted, and under its broadest reasonable interpretation (BRI), covers concepts that can be practically performed in the human mind with the aid of pen and paper. The generating steps broadly defines an architecture search of candidate models involving the steps of generating a list of candidate models for specific hardware, where each model is assigned randomly initialized parameter value (weights). This process of generating a list of candidate models without any technical implementation or details of the generating process encompasses concepts that can be performed by a human with the aid of pen and paper. See MPEP § 2106.04(a)(2)(III).) randomly sample a set of data samples to produce random data samples for inputting at each transformer model candidate; (An abstract idea of a mental process. Examiner’s note: the “sampling” step, as drafted, and under its broadest reasonable interpretation (BRI), covers concepts that can be practically performed in the human mind. The process of randomly sampling a set of data that will be used as inputs to the generated list of candidate models. This covers concepts performed by a human with the aid of pen and paper. See MPEP § 2106.04(a)(2) (III).) compute an attention confidence score for each transformer model candidate based on the random data samples and the random weights; (An abstract idea of a mental process and/or mathematical concepts. Examiner’s note: the “computing” step, as drafted, and under its broadest reasonable interpretation (BRI), covers concepts that fall under the mental process and mathematical concept. This step involves calculating attention confidence score for each candidate model using a mathematical equation. See Specification [0059]. The claim does not define the technical implementation of the transformer model evaluated using the attention confidence score. This broadly covers a mathematical calculation that can be performed in the human mind with the aid of pen and paper. See MPEP § 2106.04(a)(2)(I) & (III).) select a transformer model candidate for the target device based on the attention confidence score. (An abstract idea of a mental process. Examiner’s note: the “selecting” step, as drafted, and under its broadest reasonable interpretation (BRI), covers concepts that can be practically performed in the human mind. The claimed step defines a decision-making process that can be performed in the human mind. For example, selecting the model based on the calculation/analysis of confidence score and/or performance metrics identified for each candidate model can be manually determined. This is an evaluation and judgment process that can be performed mentally without a computer. See MPEP § 2106.04(a)(2) (III).) Step 2A Prong 2: Under this prong, we evaluate whether the claim recites additional elements that integrate the abstract idea into a practical application by considering the claim as a whole. The judicial exception is not integrated into a practical application. Additional Elements Analysis: The claim recite the additional element such as: “An apparatus comprising: at least one memory; and at least one processor coupled to the at least one memory, the at least one processor configured to:” (This amounts to no more than merely reciting the words "apply it" (or an equivalent) with the judicial exception, or merely including instructions to implement an abstract idea on a computer, or merely using a computer as a tool to perform an abstract idea, as discussed in MPEP § 2106.05(f). In other words, the claim invokes computer and/or other machinery in its ordinary capacity merely as a tool to perform the abstract idea.) Additionally, the recitation of “generate a set of transformer model candidates for a target device,” (This amounts to no more than invoking computer or other machinery in its ordinary capacity to perform an existing process, as discussed in MPEP § 2106.05(f). Examiner’s Note: The high-level recitation of generating/building transformer model for target device represent a high-level machine learning operation that is generic and conventional.) Step 2B: Under this prong, the claim must include additional elements that amount to significantly more than the judicial exception. These elements must not be well-understood, routine, or conventional in the relevant field. When viewed individually and as an ordered combination, the claim does not include any such additional elements that are sufficient to amount to significantly more (i.e., inventive concept). Additional Elements Analysis: As explained above, the claimed additional elements merely represents generic computer components configured to execute computer instructions to execute conventional models and perform the abstract ideas. As described in MPEP § 2106.05(f), additional elements that invoke computers or other machinery merely as a tool to perform an existing process will generally not amount to significantly more than a judicial exception. Therefore, claim 1 does not recite patent-eligible subject matter. Regarding Claim 2, Step 2A Prong 1: Claim 2, which incorporates the rejection of claim 1, recites further limitation such as: determine one or more performance metric for each transformer model candidate, in which the transformer model candidate is selected based on the one or more performance metric for each transformer model candidate. (That is part of the abstract idea recited in claim 1. This merely defines the process of determining performance metric for each model candidate. The performance metric can be determined/calculated based on evaluation that can be practically performed in the human mind and/or with physical aid pen and paper. Additionally, the selection is performed based on the performance metric evaluation (i.e., comparison and analysis). These steps fall under the mental processes – concepts performed in the human mind (including an observation, evaluation, judgment, opinion) (see MPEP § 2106.04(a)(2), subsection III).) Step 2A Prong 2: The judicial exception is not integrated into a practical application. process the random data samples using each transformer model candidate; (This amounts to no more than merely reciting the words "apply it" (or an equivalent) with the judicial exception, or merely including instructions to implement an abstract idea on a computer, or merely using a computer as a tool to perform an abstract idea, as discussed in MPEP § 2106.05(f). Examiner’s Note: this represents a generic computer function i.e., high-level computer instructions of applying input data to a machine learning model.) Step 2B: the claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As explained above, the additional elements identified above do not provide significantly more than the abstract idea. Processing input data as input to a machine learning model represents a generic computer component performing generic computer functions at a high level of generality and does not meaningful limit the claim. See MPEP § 2106.05. Therefore, claim 2 is ineligible. Regarding Claim 3, Step 2A Prong 1: Claim 3, which incorporates the rejection of claim 2, recites further limitation such as: in which the one or more performance metric is compared to one or more threshold and the transformer model candidate is selected based on the comparing. (That is part of the abstract idea recited in claim 2. The claim merely introduces a threshold that is used in the evaluation of selecting the transformer model candidate. This threshold comparison is an act of evaluating information that can be practically performed in the human mind. This process falls under the mental processes category of abstract idea i.e., concepts performed in the human mind (including an observation, evaluation, judgment, opinion) (see MPEP § 2106.04(a)(2), subsection III).) Step 2A Prong 2: The claim does not recite additional element that integrates the judicial exception into a practical application. Step 2B: The claim does not recite additional elements that amount to significantly more than the judicial exception). Therefore, claim 3 is ineligible. Regarding Claim 4, Step 2A Prong 1: Claim 4, which incorporates the rejection of claim 1, recites further limitation such as: compute the attention confidence score as an average of a maximum attention weight based on the random weights and the random data samples. (That is part of the abstract idea recited in claim 1. Examiner’s note: the “computing” step, as drafted, and under its broadest reasonable interpretation (BRI), covers concepts that can be practically performed in the human mind with the aid of pen and paper. But for the recitation of a processor, that is not other than using a computer component to perform the abstract idea. This step is directed to a mathematical calculation that can be performed by an individual to determine the performance metric of each candidate model. See MPEP § 2106.04(a)(2)(I) & (III).) Step 2A Prong 2: The judicial exception is not integrated into a practical application. The recitation of at least one processor is configured to compute the attention confidence score amounts to no more than using/invoking a computer component to perform the abstract idea. This limitation does not limit the claim as it simply applies the abstract idea using computer components. See MPEP § 2106.05(f). Step 2B: the claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As explained above, the additional element identified above does not provide significantly more than the abstract idea. The additional element represents the implementation of the abstract idea on a computer. See MPEP § 2106.05(d). Therefore, claim 4 is ineligible. Regarding Claim 5, Step 2A Prong 1: Claim 5, which incorporates the rejection of claim 1, recites further limitation such as: select the transformer model candidate based on an attention variance for each transformer model candidate. (That is part of the abstract idea recited in claim 1. Examiner’s note: the “selecting” step, as drafted, and under its broadest reasonable interpretation (BRI), covers concepts that can be practically performed in the human mind with the aid of pen and paper. But for the recitation of a processor, that is not other than using a computer component to perform the abstract idea. This step is directed to a mathematical calculation that can be performed by an individual to determine the attention variance of each candidate model. This selection based on the computed attention variance broadly encompasses evaluation and judgment that can be performed in the human mind with the aid of pen and paper. See MPEP § 2106.04(a)(2)(I) & (III).) Step 2A Prong 2: The judicial exception is not integrated into a practical application. The recitation of at least one processor is configured to select the transformer model based on the computed attention variance amounts to no more than using/invoking a computer component to perform the abstract idea. This limitation does not limit the claim as it simply applies the abstract idea using computer components. See MPEP § 2106.05(f). Step 2B: the claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As explained above, the additional element identified above does not provide significantly more than the abstract idea. The additional element represents the implementation of the abstract idea on a computer. See MPEP § 2106.05(d). Therefore, claim 5 is ineligible. Regarding Claim 6, Step 2A Prong 1: Claim 6, which incorporates the rejection of claim 1, doesn’t recite an abstract idea. Step 2A Prong 2: The judicial exception is not integrated into a practical application. in which the random data samples comprise images that are not artificially generated to include random pixel values. (This amounts to generally linking the use of a judicial exception to a particular technological environment or field of use, as discussed in MPEP § 2106.05(h). The limitation merely defines the type of data being used. The claim does not define the technical implementation of processing such data and merely defines the data at high-level of generality. Additionally, obtaining/collecting random data comprising an image represents a generic computer function e.g., mere data gathering in conjunction with an abstract idea. See MPEP § 2106.05(g).) Step 2B: the claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As explained above, the additional element identified above does not provide significantly more than the abstract idea. The recited type of data being used in the process does not meaningfully limit the claim as it merely links the judicial exception to a field of use. Additionally, collecting these type of data represents routine data gathering. See MPEP § 2106.05(d). Accordingly, this does not provide an inventive concept. Therefore, claim 6 is ineligible. Regarding Claim 7, Step 2A Prong 1: Claim 7, which incorporates the rejection of claim 1, doesn’t recite an abstract idea. Step 2A Prong 2: The judicial exception is not integrated into a practical application. in which the selected transformer model candidate comprises a vision transformer or a language model. (This amounts to generally linking the use of a judicial exception to a particular technological environment or field of use, as discussed in MPEP § 2106.05(h).) Step 2B: the claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As explained above, the additional element identified above does not provide significantly more than the abstract idea. The recited type of model being selected does not meaningfully limit the claim as it merely links the judicial exception to a particular technological environment or field of use. Accordingly, this does not provide an inventive concept. Therefore, claim 7 is ineligible. Regarding Claim 8, The claim recites similar limitations as corresponding claim 1. Therefore, the same analysis (subject matter eligibility analysis) that was utilized for claim 1, as described above, is equally applicable to claim 8. The only difference is that claim 1 is drawn to an apparatus and claim 8 is drawn to a method. Therefore, claim 8 is ineligible. Regarding Claim 9, The claim recites similar limitations as corresponding claim 2. Therefore, the same subject matter eligibility analysis (including the abstract idea) that was utilized for claim 2, as described above, is equally applicable to claim 9. Therefore, claim 9 is ineligible. Regarding Claim 10, The claim recites similar limitations as corresponding claim 3. Therefore, the same subject matter eligibility analysis (including the abstract idea) that was utilized for claim 3, as described above, is equally applicable to claim 10. Therefore, claim 10 is ineligible. Regarding Claim 11, The claim recites similar limitations as corresponding claim 4. Therefore, the same subject matter eligibility analysis (including the abstract idea) that was utilized for claim 4, as described above, is equally applicable to claim 11. Therefore, claim 11 is ineligible. Regarding Claim 12, The claim recites similar limitations as corresponding claim 5. Therefore, the same subject matter eligibility analysis (including the abstract idea) that was utilized for claim 5, as described above, is equally applicable to claim 12. Therefore, claim 12 is ineligible. Regarding Claim 13, The claim recites similar limitations as corresponding claim 6. Therefore, the same subject matter eligibility analysis (including the abstract idea) that was utilized for claim 6, as described above, is equally applicable to claim 13. Therefore, claim 13 is ineligible. Regarding Claim 14, The claim recites similar limitations as corresponding claim 7. Therefore, the same subject matter eligibility analysis (including the abstract idea) that was utilized for claim 7, as described above, is equally applicable to claim 14. Therefore, claim 14 is ineligible. Regarding Claim 15, The claim recites similar limitations as corresponding claim 1. Therefore, the same analysis (subject matter eligibility analysis) that was utilized for claim 1, as described above, is equally applicable to claim 15. The only difference is that claim 1 is drawn to an apparatus, and claim 15 is drawn to a non-transitory computer-readable medium. The recitation of “a non-transitory computer-readable medium having program code recorded thereon, the program code executed by a processor” merely defines computer component and instructions to implement a judicial exception, and hence the claimed additional elements listed above are merely generic elements and the implementation of the elements merely amount to no more than instructions to apply the abstract idea using generic computer components. Therefore, the additional elements do not integrate the judicial exception into a practical application or amount to significantly more. See MPEP 2106.05(f). Therefore, claim 15 is ineligible. Regarding Claim 16, The claim recites similar limitations as corresponding claim 2. Therefore, the same subject matter eligibility analysis (including the abstract idea) that was utilized for claim 2, as described above, is equally applicable to claim 16. Therefore, claim 16 is ineligible. Regarding Claim 17, The claim recites similar limitations as corresponding claim 3. Therefore, the same subject matter eligibility analysis (including the abstract idea) that was utilized for claim 3, as described above, is equally applicable to claim 17. Therefore, claim 17 is ineligible. Regarding Claim 18, The claim recites similar limitations as corresponding claim 4. Therefore, the same subject matter eligibility analysis (including the abstract idea) that was utilized for claim 4, as described above, is equally applicable to claim 18. Therefore, claim 18 is ineligible. Regarding Claim 19, The claim recites similar limitations as corresponding claim 5. Therefore, the same subject matter eligibility analysis (including the abstract idea) that was utilized for claim 5, as described above, is equally applicable to claim 19. Therefore, claim 19 is ineligible. Regarding Claim 20, The claim recites similar limitations as corresponding claim 6. Therefore, the same subject matter eligibility analysis (including the abstract idea) that was utilized for claim 6, as described above, is equally applicable to claim 20. Therefore, claim 20 is ineligible. Regarding Claim 21, The claim recites similar limitations as corresponding claim 7. Therefore, the same subject matter eligibility analysis (including the abstract idea) that was utilized for claim 7, as described above, is equally applicable to claim 21. Therefore, claim 21 is ineligible. Regarding Claim 22, The claim recites similar limitations as corresponding claim 1. Therefore, the same analysis (subject matter eligibility analysis) that was utilized for claim 1, as described above, is equally applicable to claim 22. Therefore, claim 22 is ineligible. Regarding Claim 23, The claim recites similar limitations as corresponding claim 2. Therefore, the same subject matter eligibility analysis (including the abstract idea) that was utilized for claim 2, as described above, is equally applicable to claim 23. Therefore, claim 23 is ineligible. Regarding Claim 24, The claim recites similar limitations as corresponding claim 3. Therefore, the same subject matter eligibility analysis (including the abstract idea) that was utilized for claim 3, as described above, is equally applicable to claim 24. Therefore, claim 24 is ineligible. Regarding Claim 25, The claim recites similar limitations as corresponding claim 4. Therefore, the same subject matter eligibility analysis (including the abstract idea) that was utilized for claim 4, as described above, is equally applicable to claim 25. Therefore, claim 25 is ineligible. Regarding Claim 26, The claim recites similar limitations as corresponding claim 5. Therefore, the same subject matter eligibility analysis (including the abstract idea) that was utilized for claim 5, as described above, is equally applicable to claim 26. Therefore, claim 26 is ineligible. Regarding Claim 27, The claim recites similar limitations as corresponding claim 6. Therefore, the same subject matter eligibility analysis (including the abstract idea) that was utilized for claim 6, as described above, is equally applicable to claim 27. Therefore, claim 27 is ineligible. Regarding Claim 28, The claim recites similar limitations as corresponding claim 7. Therefore, the same subject matter eligibility analysis (including the abstract idea) that was utilized for claim 7, as described above, is equally applicable to claim 28. Therefore, claim 28 is ineligible. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claim(s) 1-2, 4, 7-9, 11, 14-16, 18, 21-23, 25, and 28 are rejected under 35 U.S.C. 103 as being unpatentable over Serianni et al., (IDS: "Training-free Neural Architecture Search for RNNs and Transformers." (2023)) in view of Sarah et al., (Pub. No.: US 20220035878 A1), hereinafter, Serianni in view of Sarah. Regarding Claim 1, Serianni discloses the following: An apparatus comprising: at least one memory; and at least one processor coupled to the at least one memory, the at least one processor configured to: (Serianni, [Abstract] “First, we develop a new training free metric, named hidden covariance, that predicts the trained performance of an RNN architecture and significantly outperforms existing training-free metrics. We experimentally evaluate the effectiveness of the hidden covariance metric on the NAS-Bench-NLP bench mark. Second, we find that the current search space paradigm for transformer architectures is not optimized for training-free neural architecture search. Instead, a simple qualitative analysis can effectively shrink the search space to the best performing architectures... Ultimately, our analysis shows that the architecture search space and the training-free metric must be developed together in order to achieve effective results.” [P. 2, Section: 1, Col. 1] “we apply existing training-free metrics and create our own metrics for RNNs and BERT-based transformers with language modeling tasks. Our main contributions are: • We develop a new training-free metric for RNN architectures, called “hidden covariance,” which significantly outperforms existing metrics on NAS-Bench-NLP.” [P. 13, Co. 1] “All transformer architectures were trained on TPUv2s with 8 cores and 64 GB of memory, using Google Collaboratory. The entire process of pretraining and finetuning our benchmark took approximately 25 TPU days. Evaluation of training free metrics occurred on 2.8 GHz Intel Cascade Lake processors with either 16 or 32 cores and 32 GB of memory.”) generate a set of transformer model candidates ..., each transformer model candidate of the set of transformer model candidates being initialized with random weights; (Serianni, [P. 5, Section: 4.2.1] “To allow for simpler implementation of the FlexiBERT search space and the utilization of absolute positional encoding, we keep the hidden dimension constant across all encoder layers. In total, this search space encompasses 10,621,440 different transformer architectures.” [P. 5, Section: 4.2.3] “We pretrain a random sample of 500 architectures from the FlexiBERT subspace using ELECTRA with the OpenWebText corpus, consisting of 38 GB of tokenized text data from 8,013,769 documents.” [P. 14, Section: B Ablation Studies] “Our evaluation of training-free metrics on both NAS-Bench-NLP and our NAS BERT bench mark requires random initialization of architectures, and many metrics require a mini-batch of input data, which we randomly sampled from respective datasets. To investigate the impact of initialization weights and input data, we conduct a series ablation studies for the training-free metrics on both benchmarks. Figures 7 and 8 show how the various training free metrics evaluated on 10 architectures from NAS-Bench and our NAS BERT benchmark each differ with 10 different initialization weights. Over all, initialization weight has minimal impact on the evaluations of training-free metrics, and the metrics’ scores are well distinguished between different architectures.” [p. 4, Section:3.3] “Given a minibatch of inputs fed into the network, the metric calculates the similarity of the activations within the initialized network between each input using their Hamming distance.”) [Examiner’s Note: Serianni teaches a training-free Neural Architecture Search (NAS) that generates a set of transformer model candidates (e.g., 500 FlexiBERT architectures sampled from the search space), the network search includes initialized with random weights, as part of the training-free metrics.] randomly sample a set of data samples to produce random data samples for inputting at each transformer model candidate; (Serianni, [P. 14, Section: B Ablation Studies] “Our evaluation of training-free metrics on both NAS-Bench-NLP and our NAS BERT bench mark requires random initialization of architectures, and many metrics require a mini-batch of input data, which we randomly sampled from respective datasets. To investigate the impact of initialization weights and input data, we conduct a series ablation studies for the training-free metrics on both benchmarks. Figures 7 and 8 show how the various training free metrics evaluated on 10 architectures from NAS-Bench and our NAS BERT benchmark each differ with 10 different initialization weights. Over all, initialization weight has minimal impact on the evaluations of training-free metrics, and the metrics’ scores are well distinguished between different architectures.” [P. 16, Section: B Ablation Studies] “Figure 10: Ablation study showing the effect of different minibatch inputs on training-free metrics, evaluated using transformer architectures from our NAS BERT benchmark. 10 architectures were sampled from the benchmark, one in each decile range of test loss (e.g., 0-10%,10-20%,...,90-100%). The same 10 minibatches of size 128, randomly selected from the OpenWebText dataset, were used for each architecture and metric.” [p. 4, Section:3.3] “Given a minibatch of inputs fed into the network, the metric calculates the similarity of the activations within the initialized network between each input using their Hamming distance.” [p. 4, Section: 3.6] “For transformer-specific metrics, we look into current transformer pruning literature... the attention heads of a trained transformer encoder block by computing the “confidence” of a head using a sample mini batch of input tokens. Confident heads attend their output highly to a single token, and, hypothetically, are more important to the transformer’s task.”) [Examiner’s Note: Serianni teaches that the NAS involves randomly sampling a mini-batch of input data from a dataset and inputting the sample data into each of the network architectures.] compute an attention confidence score for each transformer model candidate based on the random data samples and the random weights; (Serianni, [P. 4, Section: 3.6. Col. 1] “For transformer-specific metrics, we look into current transformer pruning literature... propose pruning the attention heads of a trained transformer encoder block by computing the “confidence” of a head using a sample mini batch of input tokens. Confident heads attend their output highly to a single token, and, hypothetically, are more important to the transformer’s task... These three attention scores are summarized by: Confidence: A h X = 1 N ∑ n = 1 N | max ⁡ A t t h x n | ..., where X   =   x n n = 1 N is a minibatch of N inputs, L is the loss function of the model, and A t t h and σ h are an attention head and its softmax respectively. Weexpand these scores into an metric for the entire network by averaging over all H attention heads: A ( X )   =   ∑ h = 1 H 1 H A t t h ( X ) .” [P. 7, Figure 2] “Figure 2: Plots of training-free metrics evaluated on 500 architectures randomly sampled from the FlexiBERT search space, against GLUE score of the pretrained and finetuned architecture.” [P. 8, Section: 7] “Evaluating the training-free metrics on our benchmark, our proposed Attention Confidence metric performs the best.”) [Examiner’s Note: Serianni defines and uses attention confidence metric for each network architecture, which involves averaging across all attention heads. The training-free NAS metrics experiments performed by the initialization of the random weights and data samples.] and select a transformer model candidate .. based on the attention confidence score. (Serianni, [Abstract] “Second, we find that the current search space paradigm for transformer architectures is not optimized for training-free neural architecture search. Instead, a simple qualitative analysis can effectively shrink the search space to the best performing architectures.” [P. 7, Figure 2] “Figure 2: Plots of training-free metrics evaluated on 500 architectures randomly sampled from the FlexiBERT search space, against GLUE score of the pretrained and finetuned architecture.” [P. 8, Section: 7] “Evaluating the training-free metrics on our benchmark, our proposed Attention Confidence metric performs the best.”) While Serianni teaches the proposed training-free neural architecture search (NAS) for RNNs and transformers and defines training-free metrics (i.e., attention confidence metric) that can be used as tools to speed up NAS algorithms to identify the best performing network architecture, Serianni is salient and does not appear to explicitly recite that: Generating the set of transformer model candidates specifically for a target device; and Selecting the transformer model candidate for the target device. However, Serianni in view of Sarah teaches the limitations: generate a set of transformer model candidates for a target device, each transformer model candidate of the set of transformer model candidates being initialized with random weights; (Sarah, [0016] “The present disclosure provides an ML architecture search system that discovers ML architectures that are applicable to specified AI/ML application domains and are optimized for specified hardware platforms in significantly less time...” [0037]-[0038] “The MLAS function 200 includes a population initializer 201, an architecture search type engine 202, a multi-objective candidate generator (MOCG) 203, and a performance metric evaluator 204. The elements of the MLAS function 200 may operate as follows. The population initializer 201 initializes a population of candidate ML architectures as candidate solutions to the MOEA.” [0040] “the initial population of candidate ML architectures is generated using the information from one or more previous optimal ML architectures. The previous optimal ML architectures may be previously discovered ML architectures that were determined to solve a same or similar AI/ML task, are within a same or similar AI/ML domain, and/or are associated with the same or similar HPI (e.g., were deployable on a same or similar hardware platform). “ [0045] “The MOCG 203 generates candidate ML architectures from those generated initially. The MOCG 203 includes one or more optimization algorithms 230 (also referred to as “optimizers 230”, “tuning algorithms 230” or “tuners 230”) that help generate ML architectures, which should be optimal with respect to the specified performance metrics.” [0050] “For eNSGA-II 230, the MOCG 203 starts with a parent population that is representation of different subnets. In one example, 50 subnets may be randomly selected as the parent population from 1019 possible combinations within an ML architecture search space that is based on the performance metrics: accuracy and latency.” Further see [0108]-[0109].) select a transformer model candidate for the target device based on the attention confidence score. (Sarah, [Abstract] “The present disclosure is related to framework for automatically and efficiently finding machine learning (ML) architectures that are optimized to one or more specified performance metrics and/or hardware platforms.” [0046] “The optimizer(s) 230 may find a set of ML parameters that yield an optimal ML architecture that minimizes the loss function on given independent data, and selects a set of model parameters for an ML model.” [0063] “the system 100 finds ML architecture candidates and displays them to the user via the GUI 300 who can then select and download the ML architecture which best fits their needs using the graphical object 307.” [0069] “FIG. 5 shows that as the search for optimal ML architectures progresses, the system 100 not only finds higher-performing ML architectures with lower latency and higher accuracy than the random search method, and does so with a significant increase in algorithmic efficiency.” Further see [0105].) [Examiner’s Note: Sarah teaches the process of generating and selecting the optimal ML architecture for specified hardware platforms (i.e., target devices) performance metric evaluator which include performance metric and proxy metrics (i.e., attention score).] Serianni and Sarah are from the same field of endeavor and their disclosure generally relates to (Neural Architecture Search (NAS)). Accordingly, at the effective filing date, it would have been prima facie obvious to one ordinarily skill in the art to modify the combination of Serianni and Sarah to incorporate the ML architecture search approach as taught by Sarah. One would have been motivated to make such a combination in order to improve the search such that very high-performance ML architectures are found in significantly less time than other approaches. Doing so would reduce resource consumption while improving AI/ML model performance (Sarah [0017]). Regarding Claim 2, Serianni in view of Sarah teaches the elements of claim 1 as outlined above, and further teaches: process the random data samples using each transformer model candidate; and determine one or more performance metric for each transformer model candidate, in which the transformer model candidate is selected based on the one or more performance metric for each transformer model candidate. (Sarah, [0051]-[0052] “The performance metric evaluator 204 evaluates specified performance metrics of the identified ML architectures. The performance metric evaluator 204 may compute actual performance metrics 241 or may use proxy function(s) 242 to approximate the performance metrics... For the actual performance metrics 241, the performance metric evaluator 204 collects measurements of the performance metrics using actual data. For example, if accuracy was specified as one of the performance metrics, the performance metric evaluator 204 would train or otherwise operate the optimal ML model using a predefined dataset and compute its accuracy... The proxy function(s) 242 include any function that takes one or more variables, ML parameters, data, and/or the like as inputs, and produces an output that is a replacement, substitute, stand-in, surrogate, or representation of the inputs. The proxy functions 242 used by the performance metric evaluator 204 are used to approximate model performance for the optimal ML architecture.” [0102] “At operation 804, the MLAS function 200 (or the MOCG 203) determines, based on the search, a set of optimal ML architectures from the population. Here, the set of optimal ML architectures are ML architectures in the population that fit a set or combination of ML parameters and HPI included in the ML configuration better than other ML architectures in the population (e.g., goodness of fit, fitness criteria, etc.). Additionally or alternatively, the set of optimal ML architectures are ML architectures in the population have better (predicted) performance metrics in comparison with other ML architectures in the population. At operation 805, the MLAS function 200 (or the performance metric evaluator 204) evaluates performance metrics of each optimal ML architecture in the set of optimal ML architectures. This evaluation may be done using actual performance metrics measure from operating the optimal ML architectures using test data, or predicted using proxy functions.”) [Examiner’s Note: Sarah teaches the process of determining performance metrics for each model architecture using performance metric evaluator and selects the optimal model architecture based on these metrics.] Regarding Claim 4, Serianni in view of Sarah teaches the elements of claim 1 as outlined above, and further teaches: in which the at least one processor is further configured to compute the attention confidence score as an average of a maximum attention weight based on the random weights and the random data samples. (Serianni, [P. 4, Section: 3.6. Col. 1] “For transformer-specific metrics, we look into current transformer pruning literature... propose pruning the attention heads of a trained transformer encoder block by computing the “confidence” of a head using a sample mini batch of input tokens. Confident heads attend their output highly to a single token, and, hypothetically, are more important to the transformer’s task... These three attention scores are summarized by: Confidence: A h X = 1 N ∑ n = 1 N | max ⁡ A t t h x n | ..., where X   =   x n n = 1 N is a minibatch of N inputs, L is the loss function of the model, and A t t h and σ h are an attention head and its softmax respectively. Weexpand these scores into an metric for the entire network by averaging over all H attention heads: A ( X )   =   ∑ h = 1 H 1 H A t t h ( X ) .” [P. 7, Figure 2] “Figure 2: Plots of training-free metrics evaluated on 500 architectures randomly sampled from the FlexiBERT search space, against GLUE score of the pretrained and finetuned architecture.” [P. 8, Section: 7] “Evaluating the training-free metrics on our benchmark, our proposed Attention Confidence metric performs the best.”) [Examiner’s Note: Serianni defines and uses attention confidence metric for each network architecture, which involves averaging across all attention heads. The training-free NAS metrics experiments performed by the initialization of the random weights and data samples.] Regarding Claim 7, Serianni in view of Sarah teaches the elements of claim 1 as outlined above, and further teaches: in which the selected transformer model candidate comprises a vision transformer or a language model. (Sarah, [0027] “The ML config. 105 can also include an appropriately formatted dataset (or a reference to such a dataset). Here, an appropriately formatted dataset refers to a dataset that corresponds to the provided supernet, and/or the specified AI/ML task and/or AI/ML domain. For example, a dataset that would be used for the NLP domain would likely be different than a dataset used for the computer vision domain.” Serianni also states [P. 2, Section: 1] “In this work, we apply existing training-free metrics and create our own metrics for RNNs and BERT-based transformers with language modeling tasks.”) Regarding Claim 8, The claim recites substantially similar limitations as corresponding claim 1 and is rejected for similar reasons as claim 1 using similar teachings and rationale. Claim 1 is directed to an apparatus, and claim 8 is directed to a processor-implemented method. Serianni in view of Sarah also discloses the process of the neural architecture search being implemented on a computer/processors. Regarding Claim 9, The claim recites substantially similar limitations as corresponding claim 2 and is rejected for similar reasons as claim 2 using similar teachings and rationale. Regarding Claim 11, The claim recites substantially similar limitations as corresponding claim 4 and is rejected for similar reasons as claim 4 using similar teachings and rationale. Regarding Claim 14, The claim recites substantially similar limitations as corresponding claim 7 and is rejected for similar reasons as claim 7 using similar teachings and rationale. Regarding Claim 15, The claim recites substantially similar limitations as corresponding claim 1 and is rejected for similar reasons as claim 1 using similar teachings and rationale. Claim 1 is directed to an apparatus, and claim 15 is directed to non-transitory computer-readable medium having program code recorded thereon, the program code executed by a processor. Serianni in view of Sarah also discloses computer-executable instructions, such as program code, software modules, and/or functional processes. See [0144]. Regarding Claim 16, The claim recites substantially similar limitations as corresponding claim 2 and is rejected for similar reasons as claim 2 using similar teachings and rationale. Regarding Claim 18, The claim recites substantially similar limitations as corresponding claim 4 and is rejected for similar reasons as claim 4 using similar teachings and rationale. Regarding Claim 21, The claim recites substantially similar limitations as corresponding claim 7 and is rejected for similar reasons as claim 7 using similar teachings and rationale. Regarding Claim 22, The claim recites substantially similar limitations as corresponding claim 1 and is rejected for similar reasons as claim 1 using similar teachings and rationale. Regarding Claim 23, The claim recites substantially similar limitations as corresponding claim 2 and is rejected for similar reasons as claim 2 using similar teachings and rationale. Regarding Claim 25, The claim recites substantially similar limitations as corresponding claim 4 and is rejected for similar reasons as claim 4 using similar teachings and rationale. Regarding Claim 28, The claim recites substantially similar limitations as corresponding claim 7 and is rejected for similar reasons as claim 7 using similar teachings and rationale. Claim(s) 3, 10, 17, and 24 are rejected under 35 U.S.C. 103 as being unpatentable over Serianni in view of Sarah as outlined above and further in view of Zhou et al., (Pub. No.: US 20240112027 A1). Regarding Claim 3, Serianni in view of Sarah teaches the elements of claim 1 as outlined above. While Serianni in view of Sarah defines the performance metrics assigned to candidate transformer model that are then used to determine and select the optimal candidate model, Serianni in view of Sarah does not appear to explicitly suggest: in which the one or more performance metric is compared to one or more threshold and the transformer model candidate is selected based on the comparing. However, Zhou, in combination with Serianni and Sarah, teaches the limitation: in which the one or more performance metric is compared to one or more threshold and the transformer model candidate is selected based on the comparing. (Zhou, [0037]-[0038] “As part of determining the final neural network 116, the system 100 generates and trains multiple candidate neural networks to perform the machine learning task and, evaluates the trained candidate neural networks based on certain performance metrics 110. The system 100 can then determine the final neural network 116 based on the performance metrics 110 for the trained candidate neural networks.” [0066] “the system 100 can return a final neural network 116 when a trained candidate neural network attains a performance score satisfying a pre-determined threshold.” [0077] “By thresholding the step times of the candidate neural networks, the system 100 can find a final neural network 116 that performs the machine learning task more efficiently than the baseline neural network.”) Accordingly, it would have been obvious to a person having ordinary skill in the art, before the effective filing date of the claimed invention, having the combination of Serianni, Sarah, and Zhou, to incorporate the methods for neural architecture search to determine the final neural network architecture as taught by Zhou. One would have been motivated to make such a combination in order to optimize the model accuracy of the trained architectures while still taking into account their computational efficiencies (Zhou [0008]). Regarding Claim 10, The claim recites substantially similar limitations as corresponding claim 3 and is rejected for similar reasons as claim 3 using similar teachings and rationale. Regarding Claim 17, The claim recites substantially similar limitations as corresponding claim 3 and is rejected for similar reasons as claim 3 using similar teachings and rational. Regarding Claim 24, The claim recites substantially similar limitations as corresponding claim 3 and is rejected for similar reasons as claim 3 using similar teachings and rational. Claim(s) 5, 12, 19, and 26 are rejected under 35 U.S.C. 103 as being unpatentable over Serianni in view of Sarah as outlined above and further in view of Zhou(1) et al., (IDS: “Training-free Transformer Architecture Search.” (2022)). Regarding Claim 5, Serianni in view of Sarah teaches the elements of claim 1 as outlined above, and further teaches: While Serianni in view of Sarah define the Synaptic diversity score as a metric used in the literature as training-free metric for NAS, Serianni in view of Sarah is salient and does not appear to explicitly teach: in which the at least one processor is further configured to select the transformer model candidate based on an attention variance for each transformer model candidate. However, Zhou(1), in combination with Serianni and Sarah, teaches the limitation: in which the at least one processor is further configured to select the transformer model candidate based on an attention variance for each transformer model candidate. (Zhou(1), [Pp. 4-5, Section: 3.2] “the rank of the weight parameters in the MSA module could be adopted as an indicator to evaluate the ViT architecture. Synaptic Diversity. For the MSA module, it is still computationally complex to directly measure the rank of its weight matrix and hinders practical applications. To accelerate the calculation of synaptic diversity in MSA module, we leverage the Nuclear-norm of the MSA’s weight matrix to approximate its rank as the diversity indicator. Theoretically, the Nuclear-norm of a weight matrix can be treated as an equivalent substitution for its rank, when the Frobenius-norm of the weight matrix meets certain conditions. Specifically, we denote the weight parameter matrix of an MSA module as W m . m indicates the m-th linear layer in an MSA module... To better estimate the synaptic diversity of MSA modules from one ViT network that the weights are randomly initialized, we further consider the aforementioned procedure on the gradient matrix ∂L/∂Wm (L is the loss function) of each MSA module. Overall, we define the synaptic diversity of the weight parameter in the l-th MSA module as follows: S D =   ∑ m ∂ L ∂ W m n u c ⊙ W m n u c .” [p. 5, Section: 3.4, Col. 2] “Given a specified parameter constraint, we first randomly sample 8,000 subnets on one ViT search space. Then, the synaptic diversity score of MSAs and the saliency score of MLPs are calculated as the evaluation rank of each subnet. Based on the calculated DSS-indicator scores of each ViT architecture, we pick the networks with the highest proxy value as the optimal one. Finally, we retrain the searched optimal network to obtain its final test accuracy.”) [Examiner’s Note: Under the BRI, the “attention variance” interpreted as measure of variance/diversity in the attention of the transformer. Zhou teaches the transformer architecture search (TAS) to find the optimal ViT architectures using Synaptic Diversity score measure that quantifies the diversity/variance of attention weight matrixes (heads) using a rank approximation. Accordingly, it would have been obvious to a person having ordinary skill in the art, before the effective filing date of the claimed invention, having the combination of Serianni, Sarah, and Zhou, to incorporate the Transformer Architecture Search (TAS) using synaptic diversity measure as taught by Zhou. One would have been motivated to make such a combination in order to achieve a competitive search performance and improve the search efficiency in searching ViT architectures (Zhou(1) [Abstract]). Regarding Claim 12, The claim recites substantially similar limitations as corresponding claim 5 and is rejected for similar reasons as claim 5 using similar teachings and rationale. Regarding Claim 19, The claim recites substantially similar limitations as corresponding claim 5 and is rejected for similar reasons as claim 5 using similar teachings and rational. Regarding Claim 26, The claim recites substantially similar limitations as corresponding claim 5 and is rejected for similar reasons as claim 5 using similar teachings and rational. Claim(s) 6, 13, 20, and 27 are rejected under 35 U.S.C. 103 as being unpatentable over Serianni in view of Sarah as outlined above and further in view of Mellor et al., (NPL: “Neural Architecture Search without Training.” (2021)). Regarding Claim 6, Serianni in view of Sarah teaches the elements of claim 1 as outlined above: While Serianni in view of Sarah teaches the randomly sampled data comprise images, Serianni in view of Sarah is salient on whether the random data samples comprise images that are not artificially generated to include random pixel values. However, it would have been obvious in view of Mellor. Hereinafter, Mellor, in combination with Serianni and Sarah, teaches: in which the random data samples comprise images that are not artificially generated to include random pixel values. (Mellor, [p. 4, Section: 3] “We compute KH for a random subset of NAS-Bench 201 ...and NDS-DARTS ... networks at initialization for a mini-batch of CIFAR-10 images.” [p. 5, Section: 3] “Figure 3. (a)-(i): Plots of our score for randomly sampled untrained architectures inNAS-Bench-201, NAS-Bench-101, NDS-Amoeba, NDS-DARTS,NDS-ENAS, NDS-NASNet, NDS-PNAS against validation accuracy when trained. The inputs when computing the score and the validation accuracy for each plot are from CIFAR-10 except for (b) and (c) which use CIFAR-100 and ImageNet16-120 respectively. (j): We include a plot from NDS-DARTS on ImageNet (121 networks provided) to illustrate that the score extends to more challenging datasets. We use a mini-batch from ImageNette2 which is a strict subset of ImageNet with only 10 classes. In all cases there is a noticeable correlation between the score for an untrained network and the final accuracy when trained.” [P. 6, Section: 3.1] “How important are the images used to compute the score? Since our approach relies on randomly sampling a single mini-batch of data, it is reasonable to question whether different mini-batches result in different scores. To determine whether our method is dependent on mini batches, we randomly select 10 architectures from different CIFAR-100 accuracy percentiles in NAS-Bench-201 and compute the score separately for 20 random CIFAR-100 mini-batches. The resulting box-and-whisker plot is given in Figure 5(top-left): the ranking of the scores is reasonably robust to the specific choice of images. In Figure 5(top right) we compute our score using normally distributed random inputs; this has little impact on the general trend. This suggests our score captures a property of the network architecture, rather than something data-specific... Figure 5. Ablation experiments showing the effect on our score using different CIFAR-100 mini-batches (top-left), random normally distributed input images (top-right),”) [Examiner’s Note: Mellor explicitly defines the use of randomly sampled mini-batches of real-world images from different image sources (e.g., CIFAR-10, CIFAR-100, and ImageNet) as inputs to score candidate architectures. The paper study shows the difference between these samples and the normally distributed random inputs. These randomly sampled mini-batches correspond to the “images that are not artificially generated to include random pixel values.”] Accordingly, it would have been obvious to a person having ordinary skill in the art, before the effective filing date of the claimed invention, having the combination of Serianni, Sarah, and Mellor, to incorporate the Neural Architecture Search without Training approach with randomly sampled mini-batch of data as taught by Mellor. One would have been motivated to make such a combination in order to allows us search for powerful networks without any training in a matter of seconds on a single GPU, and verify its effectiveness on NAS-Bench-101, NAS Bench-201, NATS-Bench, and Network Design Spaces (Mellor [Abstract]). Regarding Claim 13, The claim recites substantially similar limitations as corresponding claim 6 and is rejected for similar reasons as claim 6 using similar teachings and rationale. Regarding Claim 20, The claim recites substantially similar limitations as corresponding claim 6 and is rejected for similar reasons as claim 6 using similar teachings and rational. Regarding Claim 27, The claim recites substantially similar limitations as corresponding claim 6 and is rejected for similar reasons as claim 6 using similar teachings and rational. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure: (Pub. No.: US 20220318595 A1) – “Sharath Nittur Sridhar” relates to “Methods, systems, articles of manufacture and apparatus to improve neural architecture searches.” (Pub. No.: US 20220012572 A1) – “Pin-Yu Chen” relates to “Neural network architecture construction” Discloses a process that involves calibrating a selected candidate neural network component described in FIG. 2 for deploying the constructed neural network architecture on a target hardware platform. (Pub. No.: US 20230289276 A1) – “Kozhaya; Joseph” relates to “Intelligently optimized machine learning models.” Discloses a method that search machine learning model repository for candidate machine learning models in a target environment. NPL: Abdelfattah, Mohamed S., et al. "Zero-cost proxies for lightweight NAS." (2021). Any inquiry concerning this communication or earlier communications from the examiner should be directed to SADIK ALSHAHARI whose telephone number is (703)756-4749. The examiner can normally be reached Monday - Friday, 9 a.m. 6 p.m. ET. Examiner interviews are available via telephone, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Li Zhen can be reached on (571) 272-3768. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /S.A.A./Examiner, Art Unit 2121 /Li B. Zhen/Supervisory Patent Examiner, Art Unit 2121
Read full office action

Prosecution Timeline

Nov 07, 2023
Application Filed
Jul 14, 2026
Non-Final Rejection mailed — §101, §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12699892
NEURAL NETWORK METHOD AND APPARATUS
5y 9m to grant Granted Aug 04, 2026
Patent 12682625
TRAINING METHOD AND APPARATUS, DIALOGUE PROCESSING METHOD AND SYSTEM, AND MEDIUM
5y 1m to grant Granted Jul 14, 2026
Patent 12682259
Methods for Compressing a Neural Network
4y 4m to grant Granted Jul 14, 2026
Patent 12632738
DELTA-SIGMA MODULATION NEURONS FOR HIGH-PRECISION TRAINING OF MEMRISTIVE SYNAPSES IN DEEP NEURAL NETWORKS
4y 11m to grant Granted May 19, 2026
Patent 12596930
SENSOR COMPENSATION USING BACKPROPAGATION
4y 9m to grant Granted Apr 07, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
38%
Grant Probability
79%
With Interview (+41.3%)
4y 5m (~1y 8m remaining)
Median Time to Grant
Low
PTA Risk
Based on 45 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month