Prosecution Insights
Last updated: October 02, 2026
Application No. 18/411,616

NEURAL NETWORK SEARCH METHOD AND RELATED DEVICE

Non-Final OA §101§102§103§112
Filed
Jan 12, 2024
Priority
Jul 15, 2021 — CN 202110803202.X +1 more
Examiner
MAHARAJ, DEVIKA S
Art Unit
Tech Center
Assignee
Huawei Technologies Co., Ltd.
OA Round
1 (Non-Final)
57%
Grant Probability
Moderate
1-2
OA Rounds
1y 9m
Est. Remaining
64%
With Interview

Examiner Intelligence

Grants 57% of resolved cases
57%
Career Allowance Rate
50 granted / 88 resolved
-3.2% vs TC avg
Moderate +8% lift
Without
With
+7.7%
Interview Lift
resolved cases with interview
Typical timeline
4y 6m
Avg Prosecution
17 currently pending
Career history
111
Total Applications
across all art units

Statute-Specific Performance

§101
29.1%
-10.9% vs TC avg
§103
47.6%
+7.6% vs TC avg
§102
9.9%
-30.1% vs TC avg
§112
10.8%
-29.2% vs TC avg
Black line = Tech Center average estimate • Based on career data from 88 resolved cases

Office Action

§101 §102 §103 §112
DETAILED ACTION 1. This communication is in response to the Application No. 18/411,616 filed on January 12, 2024 and preliminary amendments filed on January 30, 2024 in which Claims 1-20 are presented for examination. Notice of Pre-AIA or AIA Status 2. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Information Disclosure Statement 3. The information disclosure statements submitted on 07/26/2024, 09/05/2024, 01/06/2025, 01/15/2025 are in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statements are being considered by the examiner. Claim Rejections - 35 USC § 112 4. The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. 5. Claims 11-13 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. The term “positive impact” in Claims 11-12 is a relative term which renders the claim indefinite. The term “positive impact” is not defined by the claim, the specification does not provide a standard for ascertaining the requisite degree, and one of ordinary skill in the art would not be reasonably apprised of the scope of the invention. Although Applicant’s specification Par. [0047-0048] discloses calculating the “positive impact” on performance, these are merely embodiments of such a calculation and do not limit/clarify the broad limitation of “positive impact” on performance – Applicant is encouraged to amend the claims to provide further details as to how the “positive impact” on performance is calculated/determined in context of the instant claim limitations. This applies to Claims 11-12 and Claim 13 by virtue of dependency. Claim Rejections - 35 USC § 101 6. 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. 7. Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Regarding Claim 1: Step 1: Claim 1 is a method type claim. Therefore, Claims 1-13 are directed to either a process, machine, manufacture, or composition of matter. 2A Prong 1: If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation by mathematical calculation but for the recitation of generic computer components, then it falls within the “Mathematical Concepts” grouping of abstract ideas. […] the plurality of operators are obtained by sampling a plurality of candidate operators comprised in a first search space (mathematical process – sampling a plurality of candidate operators comprised in a first search space may be performed by mathematical process utilizing random sampling – Applicant’s specification Par. [0011] further describes the mathematical sampling methods which may be performed to obtain the plurality of candidate operators) selecting a target neural network from the plurality of candidate neural networks based on performance of the plurality of candidate neural networks (mental process – selecting a target neural network from the plurality of candidate neural networks based on performance may be performed manually by a user observing/analyzing the plurality of candidate neural networks and their respective performance and accordingly using judgement/evaluation to select a target neural network from the plurality based on said analysis of performance (i.e., a user may select the network with the highest/best associated performance as compared to a threshold)) 2A Prong 2: This judicial exception is not integrated into a practical application. Additional elements: obtaining a plurality of candidate neural networks, wherein at least one candidate neural network in the plurality of candidate neural networks comprises a target transformer layer, the target transformer layer comprises a target attention head, comprising a plurality of operators […] (Adding insignificant extra-solution activity to the judicial exception – see MPEP 2106.05(g)) 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Additional elements: obtaining a plurality of candidate neural networks, wherein at least one candidate neural network in the plurality of candidate neural networks comprises a target transformer layer, the target transformer layer comprises a target attention head, comprising a plurality of operators […] (MPEP 2106.05(d)(II) indicates that merely “Receiving or transmitting data over a network” is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer) For the reasons above, Claim 1 is rejected as being directed to an abstract idea without significantly more. This rejection applies equally to dependent claims 2-13. The additional limitations of the dependent claims are addressed below. Regarding Claim 2: Step 2A Prong 1: See the rejection of Claim 1 above, which Claim 2 depends on. Step 2A Prong 2 & Step 2B: wherein the target attention head is constructed based on the plurality of operators and an arrangement relationship between the plurality of operators determined in a sampling manner (Field of Use – limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that the target attention head is constructed based on the plurality of operators and an arrangement relationship between the operators does not integrate the exception into a practical application nor amount to significantly more – See MPEP 2106.05(h)) Accordingly, under Step 2A Prong 2 and Step 2B, these additional elements do not integrate the abstract idea into practical application because they do not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 1. Regarding Claim 3: Step 2A Prong 1: See the rejection of Claim 1 above, which Claim 3 depends on. […] the plurality of operators are used to perform an operation on a data processing result of the first linear transformation layer (mathematical process –performing an operation on a data processing result of the first linear transformation layer may be performed by mathematical process utilizing a mathematical equation/formula for performing an operation (e.g., SoftMax, square root, transpose, dot multiplication, cosine similarity, etc. as supported by Applicant’s specification Par. [0018]) on a data processing result of the first linear transformation layer) Step 2A Prong 2 & Step 2B: wherein the target attention head further comprises a first linear transformation layer to process an input vector of the target attention head by using a target transformation matrix (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner' s note: high level recitation of applying a trained machine learning model with previously determined data without significantly more. This cannot provide an inventive concept) Accordingly, under Step 2A Prong 2 and Step 2B, these additional elements do not integrate the abstract idea into practical application because they do not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 1. Regarding Claim 4: Step 2A Prong 1: See the rejection of Claim 3 above, which Claim 4 depends on. Step 2A Prong 2 & Step 2B: wherein the target transformation matrix comprises only X transformation matrices, X is a positive integer less than or equal to 4, and a quantity of X is determined in a sampling manner (Field of Use – limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that the target transformation matrix comprises only X transformation matrices where x is less than or equal to 4 and is determined in a sampling manner does not integrate the exception into a practical application nor amount to significantly more – See MPEP 2106.05(h)) Accordingly, under Step 2A Prong 2 and Step 2B, these additional elements do not integrate the abstract idea into practical application because they do not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 1. Regarding Claim 5: Step 2A Prong 1: See the rejection of Claim 3 above, which Claim 5 depends on. Step 2A Prong 2 & Step 2B: wherein a size of the input vector of the target attention head and a size of an output vector of the target attention head are the same (Field of Use – limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that a size of the input vector and a size of an output vector of the target attention head are the same does not integrate the exception into a practical application nor amount to significantly more – See MPEP 2106.05(h)) Accordingly, under Step 2A Prong 2 and Step 2B, these additional elements do not integrate the abstract idea into practical application because they do not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 1. Regarding Claim 6: Step 2A Prong 1: See the rejection of Claim 1 above, which Claim 6 depends on. Step 2A Prong 2 & Step 2B: wherein a quantity of operators comprised in the target attention head is less than a preset value (Field of Use – limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that a quantity of operators comprised in the target attention head is less than a preset value does not integrate the exception into a practical application nor amount to significantly more – See MPEP 2106.05(h)) Accordingly, under Step 2A Prong 2 and Step 2B, these additional elements do not integrate the abstract idea into practical application because they do not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 1. Regarding Claim 7: Step 2A Prong 1: See the rejection of Claim 1 above, which Claim 7 depends on. Step 2A Prong 2 & Step 2B: wherein the at least one candidate neural network comprises a plurality of network layers connected in series, the plurality of network layers comprise the target transformer layer, and a location of the target transformer layer in the plurality of network layers is determined in a sampling manner (Field of Use – limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that the at least one candidate neural network comprises a plurality of network layers in series, a target transformer layer, and the location of the layer is determined by sampling does not integrate the exception into a practical application nor amount to significantly more – See MPEP 2106.05(h)) Accordingly, under Step 2A Prong 2 and Step 2B, these additional elements do not integrate the abstract idea into practical application because they do not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 1. Regarding Claim 8: Step 2A Prong 1: See the rejection of Claim 1 above, which Claim 8 depends on. Step 2A Prong 2 & Step 2B: wherein the at least one candidate neural network comprises the plurality of network layers connected in series, the plurality of network layers comprise the target transformer layer and a target network layer, and the target network layer comprises a convolutional layer (Field of Use – limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that the at least one candidate neural network comprises a plurality of network layers in series, a target transformer layer, and a target network layer comprising a convolutional layer does not integrate the exception into a practical application nor amount to significantly more – See MPEP 2106.05(h)) Accordingly, under Step 2A Prong 2 and Step 2B, these additional elements do not integrate the abstract idea into practical application because they do not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 1. Regarding Claim 9: Step 2A Prong 1: See the rejection of Claim 1 above, which Claim 9 depends on. Step 2A Prong 2 & Step 2B: wherein a location of the target network layer in the plurality of network layers is determined in a sampling manner (Field of Use – limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that the location of the target network layer in the plurality of network layers is determined by sampling does not integrate the exception into a practical application nor amount to significantly more – See MPEP 2106.05(h)) Accordingly, under Step 2A Prong 2 and Step 2B, these additional elements do not integrate the abstract idea into practical application because they do not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 1. Regarding Claim 10: Step 2A Prong 1: See the rejection of Claim 1 above, which Claim 10 depends on. Step 2A Prong 2 & Step 2B: wherein a convolution kernel in the convolutional layer is obtained by sampling convolution kernels of a plurality of sizes comprised in a second search space (Field of Use – limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that a convolution kernel is obtained by sampling convolution kernels of a plurality of sizes comprised in a second search space does not integrate the exception into a practical application nor amount to significantly more – See MPEP 2106.05(h)) Accordingly, under Step 2A Prong 2 and Step 2B, these additional elements do not integrate the abstract idea into practical application because they do not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 1. Regarding Claim 11: Step 2A Prong 1: See the rejection of Claim 1 above, which Claim 11 depends on. constructing the target attention head in the target candidate neural networks (mathematical process – constructing the target attention head in the target candidate neural networks may be performed by mathematical process utilizing the subsequent ‘determining’ step below to determine replacement operators based on positive impact on performance and replacing the corresponding operators in the attention head with the replacement operators – See Applicant’s specification Par. [0047-0048] which details this process) determining replacement operators from M candidate operators of the plurality of candidate operators based on positive impact on performance of the first neural network when target operators in the first attention head are replaced with the M candidate operators in the first search space; and replacing the target operators in the first attention head with the replacement operators, to obtain the target attention head, wherein M is a positive integer (mathematical process – determining replacement operators from M candidate operators based on positive impact on performance of the first neural network when operators are replaced may be performed by mathematical process as described by Applicant’s specification Par. [0047-0048] and correspondingly replacing the target operators in the first attention head (by mathematical substitution) with the replacement operators to obtain the respective target attention head) Step 2A Prong 2 & Step 2B: wherein the plurality of candidate neural networks comprise a target candidate neural network (Field of Use – limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that the plurality of candidate neural networks comprise a target candidate neural network does not integrate the exception into a practical application nor amount to significantly more – See MPEP 2106.05(h)) obtaining a first neural network, wherein the first neural network comprises a first transformer layer, comprising a first attention head, and a plurality of operators comprised in the first attention head are obtained by sampling the plurality of candidate operators comprised in the first search space (MPEP 2106.05(d)(II) indicates that merely “Receiving or transmitting data over a network” is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer) Accordingly, under Step 2A Prong 2 and Step 2B, these additional elements do not integrate the abstract idea into practical application because they do not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 1. Regarding Claim 12: Step 2A Prong 1: See the rejection of Claim 11 above, which Claim 12 depends on. in response to a target operator of the target operators is located at a target operator location of a second neural network, determining, based on an operator that is in each of a plurality of trained second neural networks and that is located at the target operator location and performance of the plurality of trained second neural networks, and/or an occurrence frequency of the operator that is in each trained second neural network and that is located at the target operator location, the positive impact on the performance of the first neural network when the target operators in the first attention head are replaced with the M candidate operators in the first search space (mathematical process – determining, based on an operator that is in each of a plurality of trained second neural networks and located at the target operator location and performance of the plurality of networks and/or occurrence frequency of the operator, the positive impact on performance of the first neural network when the target operators are replaced may be performed by mathematical process utilizing an equation for calculating/determining the positive impact, as supported by Applicant’s specification Par. [0047-0048]) Step 2A Prong 2 & Step 2B: Accordingly, under Step 2A Prong 2 and Step 2B, these additional elements do not integrate the abstract idea into practical application because they do not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 1. Regarding Claim 13: Step 2A Prong 1: See the rejection of Claim 11 above, which Claim 13 depends on. performing parameter initialization on the target candidate neural network based on the first neural network, to obtain an initialized target candidate neural network, wherein an updatable parameter in the initialized target candidate neural network is obtained by performing parameter sharing on an updatable parameter at a same location in the first neural network (mental process/mathematical process – performing parameter initialization on the target candidate neural network based on the first neural network to obtain an initialized target candidate neural network and performing parameter sharing on an updatable parameter may be performed manually by a user observing/analyzing the parameters and accordingly using judgement/evaluation to initialize and update parameters (with the aid of pen and paper). Alternatively, the initialization and updating of parameters may be performed by mathematical process) Step 2A Prong 2 & Step 2B: training the target candidate neural network on which parameter initialization is performed, to obtain performance of the target candidate neural network (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner' s note: high level recitation of training a machine learning model with previously determined data without significantly more. This cannot provide an inventive concept) Accordingly, under Step 2A Prong 2 and Step 2B, these additional elements do not integrate the abstract idea into practical application because they do not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 1. Regarding Claim 14: Step 1: Claim 14 is a method type claim. Therefore, Claims 14-16 are directed to either a process, machine, manufacture, or composition of matter. 2A Prong 1: If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation by mathematical calculation but for the recitation of generic computer components, then it falls within the “Mathematical Concepts” grouping of abstract ideas. […] the plurality of operators are obtained by sampling a plurality of candidate operators comprised in a first search space (mathematical process – sampling a plurality of candidate operators comprised in a first search space may be performed by mathematical process utilizing random sampling – Applicant’s specification Par. [0011] further describes the mathematical sampling methods which may be performed to obtain the plurality of candidate operators) 2A Prong 2: This judicial exception is not integrated into a practical application. Additional elements: receiving, from a device side, a performance requirement indicating a performance requirement of a neural network (Adding insignificant extra-solution activity to the judicial exception – see MPEP 2106.05(g)) obtaining, from a plurality of candidate neural networks based on the performance requirement, a target neural network that meets the performance requirement, wherein at least one candidate neural network in the plurality of candidate neural networks comprises a target transformer layer, the target transformer layer comprises a target attention head comprising a plurality of operators […] (Adding insignificant extra-solution activity to the judicial exception – see MPEP 2106.05(g)) sending, to the device side, the target neural network (Adding insignificant extra-solution activity to the judicial exception – see MPEP 2106.05(g)) 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Additional elements: receiving, from a device side, a performance requirement indicating a performance requirement of a neural network (MPEP 2106.05(d)(II) indicates that merely “Receiving or transmitting data over a network” is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer) obtaining, from a plurality of candidate neural networks based on the performance requirement, a target neural network that meets the performance requirement, wherein at least one candidate neural network in the plurality of candidate neural networks comprises a target transformer layer, the target transformer layer comprises a target attention head comprising a plurality of operators […] (MPEP 2106.05(d)(II) indicates that merely “Receiving or transmitting data over a network” is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer) sending, to the device side, the target neural network (MPEP 2106.05(d)(II) indicates that merely “Receiving or transmitting data over a network” is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer) For the reasons above, Claim 14 is rejected as being directed to an abstract idea without significantly more. This rejection applies equally to dependent claims 15-16. The additional limitations of the dependent claims are addressed below. Regarding Claim 15: Step 2A Prong 1: See the rejection of Claim 14 above, which Claim 15 depends on. Step 2A Prong 2 & Step 2B: wherein the performance requirement comprises at least one of data processing precision, a model size, or an implemented task type (Field of Use – limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that the performance requirement comprises at least one of data processing precision, a model size, or an implemented task type does not integrate the exception into a practical application nor amount to significantly more – See MPEP 2106.05(h)) Accordingly, under Step 2A Prong 2 and Step 2B, these additional elements do not integrate the abstract idea into practical application because they do not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 14. Regarding Claim 16: Step 2A Prong 1: See the rejection of Claim 14 above, which Claim 16 depends on. Step 2A Prong 2 & Step 2B: wherein the target attention head is constructed based on the plurality of operators and an arrangement relationship between the plurality of operators determined in a sampling manner (Field of Use – limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that the target attention head is constructed based on the plurality of operators and an arrangement relationship between the plurality of operators determined by sampling does not integrate the exception into a practical application nor amount to significantly more – See MPEP 2106.05(h)) Accordingly, under Step 2A Prong 2 and Step 2B, these additional elements do not integrate the abstract idea into practical application because they do not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 14. Independent Claim 17 recites substantially the same limitations as Claim 1, in the form of an apparatus, including generic computer components. The claim is also directed to performing mental processes/mathematical calculations without significantly more, therefore it is rejected under the same rationale. For the reasons above, Claim 17 is rejected as being directed to an abstract idea without significantly more. This rejection applies equally to dependent claims 18-20. The additional limitations of the dependent claims are addressed below. Claim 18 recites substantially the same limitations as Claim 2, in the form of an apparatus, including generic computer components. The claim is also directed to performing mental processes/mathematical calculations without significantly more, therefore it is rejected under the same rationale. Claim 19 recites substantially the same limitations as Claim 3, in the form of an apparatus, including generic computer components. The claim is also directed to performing mental processes/mathematical calculations without significantly more, therefore it is rejected under the same rationale. Claim 20 recites substantially the same limitations as Claim 4, in the form of an apparatus, including generic computer components. The claim is also directed to performing mental processes/mathematical calculations without significantly more, therefore it is rejected under the same rationale. Claim Rejections - 35 USC § 102 8. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. 9. Claims 1-3, 6-10, and 14-19 are rejected under 35 U.S.C. 102(a)(2) as being anticipated by So et al. (hereinafter So) (US PG-PUB 20220383119). Regarding Claim 1, So teaches a method of neural network search (So, Par. [0004], “This specification describes a neural architecture search system implemented as computer programs on one or more computers in one or more locations that determines a network architecture for a neural network that is configured to perform a particular machine learning task.”, thus, methods (See So claims 14-20) of neural network search are disclosed), comprising: obtaining a plurality of candidate neural networks (So, Par. [0044], “Generally, the system 100 determines the architecture for the neural network by repeatedly modifying the multiple sub-model architectures, thereby generating a set of candidate architectures for the neural network, and evaluating the performance of the neural network having each candidate architecture in the set on the task.”, therefore, a plurality of candidate neural networks are obtained), wherein at least one candidate neural network in the plurality of candidate neural networks comprises a target transformer layer, the target transformer layer comprises a target attention head (So, Par. [0085], “Generally, to apply the attention mechanism, the attention sub-layer 420 uses one or more attention heads. Each attention head generates a set of queries Q, a set of keys K, and a set of values V, and then applies a particular variant of query-key-value (QKV) attention using the queries, keys, and values to generate an output.” & Par. [0088], “Each attention head is configured to transform the original queries, and keys, and values using learned transformations and then apply an attention mechanism to the transformed queries, keys, and values. Each attention head will generally learn different transformations from each other attention head.”, thus, at least one candidate network comprises a target transformer layer (attention sub-layer) which comprises a target attention head), comprising a plurality of operators, and the plurality of operators are obtained by sampling a plurality of candidate operators comprised in a first search space (So, Par. [0071], “To generate the population of candidate neural network architectures, the system can first sample a candidate sub-model architecture from the search space data and, for the sampled candidate sub-model architecture, select a candidate primitive neural network operation included in the candidate sub-model architecture with uniform randomness. The system can then determine a candidate neural network architecture that includes at least the sampled candidate sub-model architecture and configured to perform at least the selected candidate primitive neural network operation.”, therefore, the attention head may comprise a plurality of operators (primitives – see also So Par. [0060] for the explicit recitation of operators) and the plurality of operators are obtained by sampling a plurality of candidate operators in the search space); and selecting a target neural network from the plurality of candidate neural networks based on performance of the plurality of candidate neural networks (So, Par. [0070], “Once the search process is complete, the measure of fitness for each candidate architecture can also be used to determine the final architecture of the neural network for performing the machine learning task, for example a candidate architecture having the best measure of fitness may be selected as the final architecture.”, thus, selecting a target neural network from the plurality of candidates based on performance (fitness – See Par. [0069] which discloses that the metric of fitness measures the performance) of the plurality of candidates is disclosed). Regarding Claim 2, So teaches the method according to claim 1, wherein the target attention head is constructed based on the plurality of operators and an arrangement relationship between the plurality of operators determined in a sampling manner (So, Par. [0097], “In the described attention neural network 450, however, each attention head additionally applies a respective depth-wise convolution function in order to generate the attention head-specific queries, keys, or values which are subsequently used by the attention mechanism to generate initial outputs for the attention sub-layer. In other words, the attention head first applies a learned query Q linear transformation to each original query to generate a transformed query 408, and then applies a learned depth-wise convolution transformation to the transformed query to generate the attention head-specific query 409 for each original query. The use of convolution transformation increases the representation power of the attention neural network. The attention head-specific keys and values are similarly generated by using both the learned linear transformations and the learned depth-wise convolution transformations.”, thus the target attention head is constructed based on the plurality of operators (implementing a particular operation) and an arrangement relationship between the plurality of operators determined in a sampling manner (See So Par. [0071-0075] which describe the sampling process)). Regarding Claim 3, So teaches the method according to claim 1, wherein the target attention head further comprises a first linear transformation layer to process an input vector of the target attention head by using a target transformation matrix, and the plurality of operators are used to perform an operation on a data processing result of the first linear transformation layer (So, Par. [0097], “In the described attention neural network 450, however, each attention head additionally applies a respective depth-wise convolution function in order to generate the attention head-specific queries, keys, or values which are subsequently used by the attention mechanism to generate initial outputs for the attention sub-layer. In other words, the attention head first applies a learned query Q linear transformation to each original query to generate a transformed query 408, and then applies a learned depth-wise convolution transformation to the transformed query to generate the attention head-specific query 409 for each original query.”, thus, the target attention head may comprise a first linear transformation layer to process an input vector using a target transformation matrix, hence applying a linear transformation. Then, a plurality of operators are used to perform an operation (such as convolution) on a data processing result of the first linear transformation layer (transformed query)). Regarding Claim 6, So teaches the method according to claim 1, wherein a quantity of operators comprised in the target attention head is less than a preset value (So, Par. [0060], “For example, the neural network search system 100 can output data specifying the final neural network architecture 150 to the user that submitted the initial architecture data 102. For example, the architecture data can specify the neural network operators that are part of the neural network, the connectivity between the neural network operations, and the operations performed by the neural network operators.”, thus, the quantity of operators comprised in the target attention head may be less than a user-specified preset value). Regarding Claim 7, So teaches the method according to claim 1, wherein the at least one candidate neural network comprises a plurality of network layers connected in series, the plurality of network layers comprise the target transformer layer (So, Claim 1, “an attention neural network configured to perform the machine learning task, the attention neural network comprising one or more layers, each layer comprising an attention sub-layer and a feed-forward sub-layer, the attention sub-layer configured to: […]”, thus, at least one candidate neural network comprises a plurality of network layers connected in series, the plurality of network layers comprising the target transformer layer (sub-attention layer) – this is better depicted by So Figure 4), and a location of the target transformer layer in the plurality of network layers is determined in a sampling manner (So, Par. [0071], “To generate the population of candidate neural network architectures, the system can first sample a candidate sub-model architecture from the search space data and, for the sampled candidate sub-model architecture, select a candidate primitive neural network operation included in the candidate sub-model architecture with uniform randomness. The system can then determine a candidate neural network architecture that includes at least the sampled candidate sub-model architecture and configured to perform at least the selected candidate primitive neural network operation. To create new candidate neural network architectures during the search process, the system can repeatedly modifying architectures in the set of candidate architectures by applying one or more mutations to the candidate architecture.”, thus, the location of the target transformer layer in the plurality of network layers, according to the neural network architecture, may be determined by sampling). Regarding Claim 8, So teaches the method according to claim 1, wherein the at least one candidate neural network comprises the plurality of network layers connected in series, the plurality of network layers comprise the target transformer layer and a target network layer (So, Claim 1, “an attention neural network configured to perform the machine learning task, the attention neural network comprising one or more layers, each layer comprising an attention sub-layer and a feed-forward sub-layer, the attention sub-layer configured to: […]”, thus, at least one candidate neural network comprises a plurality of network layers connected in series, the plurality of network layers comprising the target transformer layer (sub-attention layer) and a target network layer (attention sub-layer comprising one or more depth-wise convolution layers) – this is better depicted by So Figure 4), and the target network layer comprises a convolutional layer (So, Par. [0085], “In particular, the attention sub-layer 420 includes one or more depth-wise convolution layers 422, e.g., one or more depth-wise convolution layers for each attention head.”, thus, the target network layer may comprise a convolutional layer). Regarding Claim 9, So teaches the method according to claim 8, wherein a location of the target network layer in the plurality of network layers is determined in a sampling manner (So, Par. [0071], “To generate the population of candidate neural network architectures, the system can first sample a candidate sub-model architecture from the search space data and, for the sampled candidate sub-model architecture, select a candidate primitive neural network operation included in the candidate sub-model architecture with uniform randomness. The system can then determine a candidate neural network architecture that includes at least the sampled candidate sub-model architecture and configured to perform at least the selected candidate primitive neural network operation. To create new candidate neural network architectures during the search process, the system can repeatedly modifying architectures in the set of candidate architectures by applying one or more mutations to the candidate architecture”, thus, the location of the target network layer in the plurality of network layers, according to the neural network architecture, may be determined by sampling). Regarding Claim 10, So teaches the method according to claim 8, wherein a convolution kernel in the convolutional layer is obtained by sampling convolution kernels of a plurality of sizes comprised in a second search space (So, Par. [0098], “In the example of FIG. 5A, the depth-wise convolution layer is a depth-wise 2-D convolution layer having a convolution kernel of size 3×1, where 3 is the width and 1 is the height, although in other examples, the depth-wise convolution layer may have a smaller or larger convolution kernel size e.g., a 1×1 kernel or a 5×1 kernel.”, thus, a convolution kernel in the convolutional layer is obtained by sampling kernels of a plurality of sizes in a second search space). Regarding Claim 14, So teaches a method of model providing (So, Par. [0060], “After the search process has terminated, e.g., after a specified amount of time has elapsed since the beginning of the search process or after the performance of the neural network having a determined architecture has reached a specified threshold, the neural network search system 100 can then output final architecture data 150 of a neural network. For example, the neural network search system 100 can output data specifying the final neural network architecture 150 to the user that submitted the initial architecture data 102.”, thus, methods of model providing are disclosed), comprising: receiving, from a device side, a performance requirement indicating a performance requirement of a neural network (So, Par. [0054], “The system 100 also obtains training data 104 for training a neural network to perform the particular task and, in some cases, a validation set for evaluating the performance of the neural network on the particular task.” & Par. [0056], “The system can receive the initial neural network architecture data 102 and the training data 104 in any of a variety of ways. For example, the system can receive the training data as an upload from a remote user of the system over a data communication network, e.g., using an application programming interface (API) made available by the system, and randomly divide the uploaded data into the training data and the validation set. As another example, the system can receive an input from a user specifying which data that is already maintained by the system, or another system that is accessible by the system, should be used as the initial neural network architecture data 102, the training data 104, or both”, thus, a performance requirement (initial architecture data, training set, and validation set for evaluating performance) may be received from device side (user)) obtaining, from a plurality of candidate neural networks based on the performance requirement, a target neural network that meets the performance requirement (So, Par. [0044], “Generally, the system 100 determines the architecture for the neural network by repeatedly modifying the multiple sub-model architectures, thereby generating a set of candidate architectures for the neural network, and evaluating the performance of the neural network having each candidate architecture in the set on the task.” & Par. [0060], “After the search process has terminated, e.g., after a specified amount of time has elapsed since the beginning of the search process or after the performance of the neural network having a determined architecture has reached a specified threshold, the neural network search system 100 can then output final architecture data 150 of a neural network.”, therefore, a plurality of candidate neural networks are obtained based on the performance requirement (evaluating the performance of the network based on received data) and a target neural network that meets the performance requirement may be obtained as the final neural network architecture), wherein at least one candidate neural network in the plurality of candidate neural networks comprises a target transformer layer, the target transformer layer comprises a target attention head (So, Par. [0085], “Generally, to apply the attention mechanism, the attention sub-layer 420 uses one or more attention heads. Each attention head generates a set of queries Q, a set of keys K, and a set of values V, and then applies a particular variant of query-key-value (QKV) attention using the queries, keys, and values to generate an output” & Par. [0088], “Each attention head is configured to transform the original queries, and keys, and values using learned transformations and then apply an attention mechanism to the transformed queries, keys, and values. Each attention head will generally learn different transformations from each other attention head.”, thus, at least one candidate network comprises a target transformer layer (attention sub-layer) which comprises a target attention head) comprising a plurality of operators, and the plurality of operators are obtained by sampling a plurality of candidate operators comprised in a first search space (So, Par. [0071], “To generate the population of candidate neural network architectures, the system can first sample a candidate sub-model architecture from the search space data and, for the sampled candidate sub-model architecture, select a candidate primitive neural network operation included in the candidate sub-model architecture with uniform randomness. The system can then determine a candidate neural network architecture that includes at least the sampled candidate sub-model architecture and configured to perform at least the selected candidate primitive neural network operation.”, therefore, the attention head may comprise a plurality of operators (primitives – see also So Par. [0060] for the explicit recitation of operators) and the plurality of operators are obtained by sampling a plurality of candidate operators in the search space); and sending, to the device side, the target neural network (So, Par. [0060], “For example, the neural network search system 100 can output data specifying the final neural network architecture 150 to the user that submitted the initial architecture data 102. For example, the architecture data can specify the neural network operators that are part of the neural network, the connectivity between the neural network operations, and the operations performed by the neural network operators.”, thus, the target neural network (final neural network architecture) may be sent to the device side (user)). Regarding Claim 15, So teaches the method according to claim 14, wherein the performance requirement comprises at least one of data processing precision, a model size, or an implemented task type (So, Par. [0054], “The system 100 also obtains training data 104 for training a neural network to perform the particular task and, in some cases, a validation set for evaluating the performance of the neural network on the particular task.”, thus, the performance requirement comprises at least one of an implemented task type, as the obtained training/validation data (used to evaluate performance) may be task specific). Regarding Claim 16, So teaches the method according to claim 14, wherein the target attention head is constructed based on the plurality of operators and an arrangement relationship between the plurality of operators determined in a sampling manner (So, Par. [0097], “In the described attention neural network 450, however, each attention head additionally applies a respective depth-wise convolution function in order to generate the attention head-specific queries, keys, or values which are subsequently used by the attention mechanism to generate initial outputs for the attention sub-layer. In other words, the attention head first applies a learned query Q linear transformation to each original query to generate a transformed query 408, and then applies a learned depth-wise convolution transformation to the transformed query to generate the attention head-specific query 409 for each original query. The use of convolution transformation increases the representation power of the attention neural network. The attention head-specific keys and values are similarly generated by using both the learned linear transformations and the learned depth-wise convolution transformations.”, thus the target attention head is constructed based on the plurality of operators (implementing a particular operation) and an arrangement relationship between the plurality of operators determined in a sampling manner (See So Par. [0071-0075] which describe the sampling process)). Regarding Claim 17, So teaches an apparatus for neural network search (So, Par. [0004], “This specification describes a neural architecture search system implemented as computer programs on one or more computers in one or more locations that determines a network architecture for a neural network that is configured to perform a particular machine learning task.”, thus, an apparatus for neural network search is disclosed), comprising: at least one processor; and one or more memories coupled to the at least one processor and storing programming instructions for execution by the at least one processor (So, Claim 1, “A system for performing a machine learning task on a network input to generate a network output, the system comprising one or more computers and one or more storage devices storing instructions that, when executed by the one or more computers, cause the one or more computers to implement: […]”, therefore, an apparatus/system comprising at least one processor and one or more memories storing programming instructions for execution by the at least one processor is disclosed) to cause the apparatus to: […] The rest of the claim language in Claim 17 recites substantially the same limitations as Claim 1, in the form of an apparatus, therefore it is rejected under the same rationale. Claim 18 recites substantially the same limitations as Claim 2 in the form of an apparatus, therefore it is rejected under the same rationale. Claim 19 recites substantially the same limitations as Claim 3 in the form of an apparatus, therefore it is rejected under the same rationale. Claim Rejections - 35 USC § 103 10. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. 11. Claims 4 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over So et al. (hereinafter So) (US PG-PUB 20220383119), in view of Park et al. (hereinafter Park) (“SANVis: Visual Analytics for Understanding Self-Attention Networks”). Regarding Claim 4, So teaches the method according to claim 3. So does not explicitly disclose wherein the target transformation matrix comprises only X transformation matrices, X is a positive integer less than or equal to 4, and a quantity of X is determined in a sampling manner. However, Park teaches wherein the target transformation matrix comprises only X transformation matrices, X is a positive integer less than or equal to 4, and a quantity of X is determined in a sampling manner (Park, Pg. 2, Section 2. Preliminaries, “At each attention head, we transform encoded word vectors into three matrices of a query, a key, and a value, Q ∈ RT×dq, K ∈ RT×dk, and V ∈RT×dv, respectively, for h times, which in turn generated h×3 matrices, using the linear transformation and compute the attention-weighted combinations of value vectors as √ Attention(Q,K,V)= Softmax QKT dmodel V MultiHeadAttention= Concat(head1,...,headh)WO (1) where headi = Attention QWQ i ,KWKi ,VWV i , and WQ i , WV i and WKi indicate the linear transformation matrices at the i-th head. In multi-head self-attention, which consists of h parallel attention heads, transformation matrices of each head are randomly initialized, and then each set is used to project input vectors onto a different repre sentation subspace.”, therefore, the target transformation matrix may comprise three matrices for each attention head. Further, the quantity of matrices may be determined in a sampling manner, depending on h (the number of heads in the multi-head self-attention) to obtain hx3 matrices – i.e., for 1 attention head, there would be 3 respective transformation matrices). It would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of claim 3, as disclosed by So to include wherein the target transformation matrix comprises only X transformation matrices, X is a positive integer less than or equal to 4, and a quantity of X is determined in a sampling manner, as disclosed by Park. One of ordinary skill in the art would have been motivated to make this modification to enable the use of sampled transformation matrices which project input vectors into different representation subspaces hence preventing overfitting and correspondingly improving generalization and model robustness (Park, Pg. 2, Section 3. Preliminaries: Self-Attention Networks, “In multi-head self-attention, which consists of h parallel attention heads, transformation matrices of each head are randomly initialized, and then each set is used to project input vectors onto a different representation subspace. For this reason, every attention head is allowed to have different attention shapes and patterns. This characteristic encourages each head differently to attend adjacent words or linguistics relation words.”). Claim 20 recites substantially the same limitations as Claim 4 in the form of an apparatus, therefore it is rejected under the same rationale. 12. Claim 5 is rejected under 35 U.S.C. 103 as being unpatentable over So et al. (hereinafter So) (US PG-PUB 20220383119), in view of Hua et al. (hereinafter Hua) (US PG-PUB 20200257961). Regarding Claim 5, So teaches the method according to claim 3. So does not explicitly disclose wherein a size of the input vector of the target attention head and a size of an output vector of the target attention head are the same. However, Hua teaches wherein a size of the input vector of the target attention head and a size of an output vector of the target attention head are the same (Hua, Par. [0024], “Generally, a cell is a fully convolutional neural network that is configured to receive a cell input and to generate a cell output. In some implementations, the cell output may have a same dimension as the cell input, e.g., the same height (H), width (W), and depth (F). For example, a cell may receive a feature map as input and generate an output feature map having the same dimension as the input feature map.”, thus, the size of the input vector and the size of the output vector may be the same). It would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of claim 3, as disclosed by So to include wherein a size of the input vector of the target attention head and a size of an output vector of the target attention head are the same, as disclosed by Hua. One of ordinary skill in the art would have been motivated to make this modification to enable the input vector and output vector to be the same size, hence providing a consistent representation space which avoids mismatch between dimensionality (Hua, Par. [0024], “Generally, a cell is a fully convolutional neural network that is configured to receive a cell input and to generate a cell output. In some implementations, the cell output may have a same dimension as the cell input, e.g., the same height (H), width (W), and depth (F). For example, a cell may receive a feature map as input and generate an output feature map having the same dimension as the input feature map.”) 13. Claims 11-13 is rejected under 35 U.S.C. 103 as being unpatentable over So et al. (hereinafter So) (US PG-PUB 20220383119), in view of Chen et al. (hereinafter Chen) (“GLiT: Neural Architecture Search for Global and Local Image Transformer”). Regarding Claim 11, So teaches the method according to claim 1, wherein the plurality of candidate neural networks comprise a target candidate neural network (So, Par. [0006], “The described techniques allow for a neural network architecture search system to construct an open-ended search space that is composed of largely degenerate neural network architecture components from any of a variety of existing neural network architectures, and thereafter automatically and effectively determine a final architecture from the search space for a neural network for performing a given machine learning task.”, thus, the plurality of candidate neural networks may comprise a target/final neural network); the obtaining the plurality of candidate neural networks comprises: constructing the target attention head in the target candidate neural networks (So, Par. [0097], “In the described attention neural network 450, however, each attention head additionally applies a respective depth-wise convolution function in order to generate the attention head-specific queries, keys, or values which are subsequently used by the attention mechanism to generate initial outputs for the attention sub-layer. In other words, the attention head first applies a learned query Q linear transformation to each original query to generate a transformed query 408, and then applies a learned depth-wise convolution transformation to the transformed query to generate the attention head-specific query 409 for each original query. The use of convolution transformation increases the representation power of the attention neural network. The attention head-specific keys and values are similarly generated by using both the learned linear transformations and the learned depth-wise convolution transformations”, thus the target attention head is constructed), the constructing the target attention head in the target candidate neural network comprising: obtaining a first neural network, wherein the first neural network comprises a first transformer layer, comprising a first attention head (So, Par. [0085], “Generally, to apply the attention mechanism, the attention sub-layer 420 uses one or more attention heads. Each attention head generates a set of queries Q, a set of keys K, and a set of values V, and then applies a particular variant of query-key-value (QKV) attention using the queries, keys, and values to generate an output.” & Par. [0088], “Each attention head is configured to transform the original queries, and keys, and values using learned transformations and then apply an attention mechanism to the transformed queries, keys, and values. Each attention head will generally learn different transformations from each other attention head.”, thus, a first neural network is obtained and comprises a target transformer layer (attention sub-layer) which comprises a target attention head), and a plurality of operators comprised in the first attention head are obtained by sampling the plurality of candidate operators comprised in the first search space (So, Par. [0071], “To generate the population of candidate neural network architectures, the system can first sample a candidate sub-model architecture from the search space data and, for the sampled candidate sub-model architecture, select a candidate primitive neural network operation included in the candidate sub-model architecture with uniform randomness. The system can then determine a candidate neural network architecture that includes at least the sampled candidate sub-model architecture and configured to perform at least the selected candidate primitive neural network operation.”, therefore, the attention head may comprise a plurality of operators (primitives – see also So Par. [0060] for the explicit recitation of operators) and the plurality of operators are obtained by sampling a plurality of candidate operators in the search space); and While So Par. [0071-0075] discloses sampling candidate model architectures from search space data and modifying architectures by applying one or more mutations to the candidate architecture, including modifying/deleting/inserting candidate operation parameters associated with the candidate primitive neural network operation, So does not explicitly disclose: determining replacement operators from M candidate operators of the plurality of candidate operators based on positive impact on performance of the first neural network when target operators in the first attention head are replaced with the M candidate operators in the first search space; and replacing the target operators in the first attention head with the replacement operators, to obtain the target attention head, wherein M is a positive integer. However, Chen teaches determining replacement operators from M candidate operators of the plurality of candidate operators based on positive impact on performance of the first neural network when target operators in the first attention head are replaced with the M candidate operators in the first search space; and replacing the target operators in the first attention head with the replacement operators, to obtain the target attention head, wherein M is a positive integer (Chen, Pgs. 4-5, “Constructing Multi-head Global-Local Module. Given global and local sub-modules, next question is how to combine them. We construct the global-local module by replacing several heads in the MHA with local sub-modules. For example, if there are N=3 heads in the MHA, we can keep one MHA head (head0) unchanged and replace two heads (head1 and head2) with our local sub-module. […] As we can see, the ratio of global and local sub-modules has obvious influence on the performance. Simply replacing all self-attention heads by convolution heads will cause a huge performance drop due to the lack of global information. On the other hand, the network with 1 self-attention head and 2 convolution heads in every global-local module performs the best among all models, improving 1.8% Top-1 accuracy compared with the baseline model. The performance variation with different ratios between self-attention and convolution heads demonstrates that introducing local information brings more performance gains only with proper global-local ratio”, therefore, replacement operators of a plurality of operators may be determined based on positive impact on performance when the target operators are replaced. Further, the target operators may be replaced with the replacement operators to obtain the target attention head. The replacement is further shown by Chen Table 1 on Pg. 5). It would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of claim 1, as disclosed by So to include determining replacement operators from M candidate operators of the plurality of candidate operators based on positive impact on performance of the first neural network when target operators in the first attention head are replaced with the M candidate operators in the first search space; and replacing the target operators in the first attention head with the replacement operators, to obtain the target attention head, wherein M is a positive integer, as disclosed by Chen. One of ordinary skill in the art would have been motivated to make this modification to improve performance of the model by replacing operators based on positive impact on performance (Chen, Pg. 5, “Simply replacing all self-attention heads by convolution heads will cause a huge performance drop due to the lack of global information. On the other hand, the network with 1 self-attention head and 2 convolution heads in every global-local module performs the best among all models, improving 1.8% Top-1 accuracy compared with the baseline model. The performance variation with different ratios between self-attention and convolution heads demonstrates that introducing local information brings more performance gains only with proper global-local ratio.”). Regarding Claim 12, So in view of Chen teaches the method according to claim 11, further comprising: in response to a target operator of the target operators is located at a target operator location of a second neural network, determining, based on an operator that is in each of a plurality of trained second neural networks and that is located at the target operator location and performance of the plurality of trained second neural networks, and/or an occurrence frequency of the operator that is in each trained second neural network and that is located at the target operator location, the positive impact on the performance of the first neural network when the target operators in the first attention head are replaced with the M candidate operators in the first search space (Chen, Pgs. 4-5, “Constructing Multi-head Global-Local Module. Given global and local sub-modules, next question is how to combine them. We construct the global-local module by replacing several heads in the MHA with local sub-modules. For example, if there are N=3 heads in the MHA, we can keep one MHA head (head0) unchanged and replace two heads (head1 and head2) with our local sub-module. […] As we can see, the ratio of global and local sub-modules has obvious influence on the performance. Simply replacing all self-attention heads by convolution heads will cause a huge performance drop due to the lack of global information. On the other hand, the network with 1 self-attention head and 2 convolution heads in every global-local module performs the best among all models, improving 1.8% Top-1 accuracy compared with the baseline model. The performance variation with different ratios between self-attention and convolution heads demonstrates that introducing local information brings more performance gains only with proper global-local ratio” & Pg. 5, “The search space of proposed global-local block includes the high-level global-local sub-module distribution and low-level detailed architecture of each sub-module. At the high-level, we aim to search the distribution of convolution and self-attention heads over all global-local blocks. At the low-level, we search the detailed architecture of all sub modules. Table 2 summarizes the high-level and low-level search space implemented in this paper”, therefore, in response to a target operator location, a positive impact on performance when the operators are replaced is determined– See Chen Tables 1 & 2 on Pg. 5). The reasons of obviousness have been noted in the rejection of Claim 11 above and applicable herein. Regarding Claim 13, So in view of Chen teaches the method according to claim 11, further comprising: performing parameter initialization on the target candidate neural network based on the first neural network, to obtain an initialized target candidate neural network, wherein an updatable parameter in the initialized target candidate neural network is obtained by performing parameter sharing on an updatable parameter at a same location in the first neural network (So, Par. [0032], “The neural architecture search system 100 is a system that obtains initial neural network architecture data 102, generates a neural architecture search space 110 from the initial neural network architecture data 102, and subsequently determines an architecture 150 for a neural network by searching through the neural architecture search space 110. The search space 110 is composed of multiple primitive neural network operations 134, where each primitive neural network operation is associated with one or more tunable operation parameters 136. For each primitive neural network operation, the tunable operation parameters generally define how the operation should be applied. The tunable operation parameters can each have a predetermined set of possible values. Some operation parameters can have numerical values that may be any value within some range, e.g., real valued constant values, while other operation parameters that can only take one of a small number of possible values, e.g., integer index values”, thus, parameter initialization may be performed on the target/final neural network based on the first neural network, to obtain an initialized neural network, where an updatable (tunable) parameter in the initialized network is obtained by performing parameter sharing on an updatable parameter (tuning) at a same location in the first neural network (based on the neural architecture search sampling and mutations)); and training the target candidate neural network on which parameter initialization is performed, to obtain performance of the target candidate neural network (So, Par. [0054], “The system 100 also obtains training data 104 for training a neural network to perform the particular task and, in some cases, a validation set for evaluating the performance of the neural network on the particular task”, therefore, the target candidate neural network may be trained and validated to obtain/evaluate performance of the target network). Conclusion 14. Any inquiry concerning this communication or earlier communications from the examiner should be directed to Devika S Maharaj whose telephone number is (571)272-0829. The examiner can normally be reached Monday - Thursday 8:30am - 5:30pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Alexey Shmatov can be reached at (571)270-3428. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /DEVIKA S MAHARAJ/Examiner, Art Unit 2123
Read full office action

Prosecution Timeline

Jan 12, 2024
Application Filed
Jan 30, 2024
Response after Non-Final Action
Aug 25, 2026
Non-Final Rejection mailed — §101, §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12737598
NEURAL NETWORK PROCESSING WITH MODEL PINNING
4y 11m to grant Granted Sep 15, 2026
Patent 12694269
SELECTIVE REPORTING OF MACHINE LEARNING PARAMETERS FOR FEDERATED LEARNING
4y 0m to grant Granted Jul 28, 2026
Patent 12682205
DIFFERENTIAL EQUATIONS NETWORK
7y 9m to grant Granted Jul 14, 2026
Patent 12682215
FLEXIBLE MACHINE LEARNING
4y 1m to grant Granted Jul 14, 2026
Patent 12675689
MULTI-DOMAIN FEATURE ENHANCEMENT FOR TRANSFER LEARNING (FTL)
4y 4m to grant Granted Jul 07, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
57%
Grant Probability
64%
With Interview (+7.7%)
4y 6m (~1y 9m remaining)
Median Time to Grant
Low
PTA Risk
Based on 88 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month