Prosecution Insights
Last updated: October 02, 2026
Application No. 18/363,487

SINGLE SEARCH FOR ARCHITECTURES ON EMBEDDED DEVICES

Final Rejection §101§103§112
Filed
Aug 01, 2023
Priority
Oct 28, 2022 — provisional 63/420,511
Examiner
ILES, TYLER EDWARD
Art Unit
2122
Tech Center
2100 — Computer Architecture & Software
Assignee
Qualcomm Incorporated
OA Round
2 (Final)
56%
Grant Probability
Moderate
3-4
OA Rounds
5m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 56% of resolved cases
56%
Career Allowance Rate
5 granted / 9 resolved
+0.6% vs TC avg
Strong +67% interview lift
Without
With
+66.7%
Interview Lift
resolved cases with interview
Typical timeline
3y 7m
Avg Prosecution
10 currently pending
Career history
28
Total Applications
across all art units

Statute-Specific Performance

§101
27.7%
-12.3% vs TC avg
§103
52.0%
+12.0% vs TC avg
§102
12.8%
-27.2% vs TC avg
§112
7.4%
-32.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 9 resolved cases

Office Action

§101 §103 §112
CTNF 18/363,487 CTNF 100901 DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA. This action is in response to an application filed on August 1 st , 2023. Claims 1-30 are pending in the current application. The information disclosure statements have been considered. Claim Rejections - 35 USC § 112 07-30-02 AIA The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. 07-34-01 Claims 1-30 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. 07-34-03 AIA The term “ small ” in claim s 1, 11, and 21 is a relative term which renders the claim indefinite. The term “ small ” is not defined by the claim, the specification does not provide a standard for ascertaining the requisite degree, and one of ordinary skill in the art would not be reasonably apprised of the scope of the invention. The use of the relative term “small” renders the scope of the kernel indefinite and so the claims are rejected. Claims dependent on claims 1, 11, and 21 are rejected based on their dependence . 07-30-03-h AIA Claim Interpretation 07-30-03 AIA The following is a quotation of 35 U.S.C. 112(f): (f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph: An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. 07-30-05 The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked. As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph: (A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function; (B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and (C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function. Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function. Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function. Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Claim Rejections - 35 USC § 101 07-04-01 AIA 07-04 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claim(s) 1-30 is/are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Regarding claim 1, Under Step 1 of the Subject Matter Eligibility Test of Products and Processes, the claim is directed towards a process, which is one of the four statutory categories. Next, under a Step 2A Prong 1 Analysis, the claim recites the following limitations, which are interpreted to be, under the broadest reasonable interpretation, abstract ideas: performing gradient descent to evolve the largest super kernel to a small kernel corresponding to the search space to generate a range of kernel encodings (mathematical concept) identifying a subset of kernel encodings from the range of kernel encodings, for each layer of the super network, based on the gradient descent (mental process) determining a set of candidate architectures based on the subset of kernel encodings, each candidate architecture of the set of candidate architectures having a different model size (mental process) selecting a target model, from the set of architectures, based on meeting hardware specifications (mental process) Therefore, we have to examine the claim under Step 2A prong 2, which considers the additional elements within the claim. The claim’s additional elements are: generating an over-parameterized super network having a plurality of layers, the super network comprising a plurality of operator types, each layer of the plurality of layers comprising a largest super kernel corresponding to a search space and applying the target model. “Generating an over-parameterized super network having a plurality of layers, the super network comprising a plurality of operator types, each layer of the plurality of layers comprising a largest super kernel corresponding to a search space” is a limitation that is interpreted to be merely indicating the particular technological environment or field of use, and “generally links” the generation of a super network with the plurality of operator types and layers to the abstract idea. (See MPEP 2106.05(h)) The limitation, “applying the target model”, is interpreted to be mere instruction to apply a judicial exception, as it instructs to take a target model, and merely apply it to the abstract idea. (See MPEP 2106.05(f)) Therefore, these additional elements do not integrate the abstract idea into a practical application. The claim is directed to an abstract idea. Under a Step 2B analysis, the claim’s additional elements do not amount to significantly more than the judicial exception as explained above in Step 2A prong 2. Therefore, the claim is ineligible. Regarding claim 11, Under Step 1 of the Subject Matter Eligibility Test of Products and Processes, the claim is directed towards a machine, which is one of the four statutory categories. Next, under a Step 2A Prong 1 Analysis, the claim recites the following limitations, which are interpreted to be, under the broadest reasonable interpretation, abstract ideas: perform gradient descent to evolve the largest super kernel to a small kernel corresponding to the search space to generate a range of kernel encodings (mathematical concept) identify a subset of kernel encodings from the range of kernel encodings, for each layer of the super network, based on the gradient descent (mental process) determine a set of candidate architectures based on the subset of kernel encodings, each candidate architecture of the set of candidate architectures having a different model size (mental process) select a target model, from the set of architectures, based on meeting hardware specifications (mental process) Therefore, we have to examine the claim under Step 2A prong 2, which considers the additional elements within the claim. The claim’s additional elements are: at least one memory at least one processor coupled to the at least one memory generate an over-parameterized super network having a plurality of layers, the super network comprising a plurality of operator types, each layer of the plurality of layers comprising a largest super kernel corresponding to a search space and apply the target model. The limitation, to “generate an over-parameterized super network having a plurality of layers, the super network comprising a plurality of operator types, each layer of the plurality of layers comprising a largest super kernel corresponding to a search space”, is interpreted to be merely indicating the particular technological environment or field of use, and “generally links” the generation of a super network with the plurality of operator types and layers to the abstract idea. (See MPEP 2106.05(h)) The limitations, “at least one memory”, “at least one processor coupled to the at least one memory”, and to “apply the target model”, is interpreted to be mere instruction to apply a judicial exception, as it instructs to take a target model, at least one memory, and a processor coupled to the memory, and merely apply it to the abstract idea. (See MPEP 2106.05(f)) Therefore, these additional elements do not integrate the abstract idea into a practical application. The claim is directed to an abstract idea. Under a Step 2B analysis, the claim’s additional elements do not amount to significantly more than the judicial exception as explained above in Step 2A prong 2. Therefore, the claim is ineligible. Regarding claim 21, Under Step 1 of the Subject Matter Eligibility Test of Products and Processes, the claim is directed towards a machine, which is one of the four statutory categories. Next, under a Step 2A Prong 1 Analysis, the claim recites the following limitations, which are interpreted to be, under the broadest reasonable interpretation, abstract ideas: means for performing gradient descent to evolve the largest super kernel to a small kernel corresponding to the search space to generate a range of kernel encodings (mathematical concept) means for identifying a subset of kernel encodings from the range of kernel encodings, for each layer of the super network, based on the gradient descent (mental process) means for determining a set of candidate architectures based on the subset of kernel encodings, each candidate architecture of the set of candidate architectures having a different model size (mental process) means for selecting a target model, from the set of architectures, based on meeting hardware specifications (mental process) Therefore, we have to examine the claim under Step 2A prong 2, which considers the additional elements within the claim. The claim’s additional elements are: means for generating an over-parameterized super network having a plurality of layers, the super network comprising a plurality of operator types, each layer of the plurality of layers comprising a largest super kernel corresponding to a search space and means for applying the target model. The “means for generating an over-parameterized super network having a plurality of layers, the super network comprising a plurality of operator types, each layer of the plurality of layers comprising a largest super kernel corresponding to a search space” is a limitation that is interpreted to be merely indicating the particular technological environment or field of use, and “generally links” the generation of a super network with the plurality of operator types and layers to the abstract idea. (See MPEP 2106.05(h)) The limitation, “means for applying the target model” is interpreted to be mere instruction to apply a judicial exception, as it instructs to take a target model, and merely apply it. (See MPEP 2106.05(f)) Therefore, these additional elements do not integrate the abstract idea into a practical application. The claim is directed to an abstract idea. Under a Step 2B analysis, the claim’s additional elements do not amount to significantly more than the judicial exception as explained above in Step 2A prong 2. Therefore, the claim is ineligible. Regarding claims 2, 12, and 22, the claims recite “quantizing the target model.” The limitation, as drafted, is interpreted to be mere instructions to apply a judicial exception, as it instructs to apply quantization, a model optimization technique, to the model. (See MPEP 2106.05(f)) Therefore, the claims are rejected on the same basis as claims 1, 11, and 21. Regarding claims 3, 13, and 23, the claim recite “splitting convolution node activations of the target model by splitting each input tensor into a plurality of smaller tensors, the plurality of smaller tensors incorporated into the set of candidate architectures.” The limitation, as drafted, is interpreted to be mere instructions to apply a judicial exception, as it instructs to split each input tensor into a plurality of smaller tensors, and incorporate the tensor into a set of architectures, all to split convolution node activations of the target model. (See MPEP 2106.05(f)) Therefore, the claims are rejected on the same basis as claims 1, 11, and 21. Regarding claims 4, 14, and 24, the claims recite “the plurality of operating types comprises at least one convolutional neural network, at least one recurrent neural network, at least one transformer, and/or at least one multilayer perceptron (MLP).” The limitation, as drafted, merely indicates the technological environment and/or field of use, and “generally links” a CNN, RNN, transformers, and/or a MLP to the abstract idea. (See MPEP 2106.05(h)) Therefore, the claims are rejected on the same basis as claims 1, 11, and 21. Regarding claims 5, 15, and 25, the claims recite “the super network is based on a particular use case, a different super network constructed for each different use case.” The limitation, as drafted, is interpreted to be mere instructions to apply a judicial exception, as it instructs to construct a different super network for different use cases. (See MPEP 2106.05(f)) Therefore, the claims are rejected on the same basis as claims 1, 11, and 21. Regarding claims 6, 16, and 26, the claims recite “identifying the subset of kernel encodings comprises progressively increasing a selected kernel choice from a smallest inner core of the largest super kernel to the largest super kernel.” The limitation, as drafted, is interpreted to be, under the broadest reasonable interpretation, a “mental process”, which is a grouping of abstract idea. Therefore, the claims are rejected on the same basis as claims 1, 11, and 21. Regarding claims 7, 17, and 27, the claims recite “identifying the subset of kernel encodings comprises progressively reducing a selected kernel choice from the largest super kernel to an inner core of the largest super kernel.” The limitation, as drafted, is interpreted to be, under the broadest reasonable interpretation, a “mental process”, which is a grouping of abstract idea. Therefore, the claims are rejected on the same basis as claims 1, 11, and 21. Regarding claims 8, 18, and 28, the claims recite “applying the target model comprises training the target model.” The limitation, as drafted, is interpreted to be mere instructions to apply a judicial exception, as it instructs to train a model before merely applying it. (See MPEP 2106.05(f)) Therefore, the claims are rejected on the same basis as claims 1, 11, and 21. Regarding claims 9, 19, and 29, the claims recite “a one-shot neural architecture search (NAS).” The limitation, as drafted, merely indicates the technological environment and/or field of use, and “generally links” a one-shot neural architecture search to the abstract idea. (See MPEP 2106.05(h)) Therefore, the claims are rejected on the same basis as claims 1, 11, and 21. Regarding claims 10, 20, and 30, the claims recite “the one-shot NAS is based on at least one of: memory usage of a device that applies the target model, run time latency of the model running on the device, a model accuracy, and power consumption of the device.” The limitation, as drafted, merely indicates the technological environment and/or field of use, and “generally links” memory usage, run time latency, model accuracy, and power consumption to the abstract idea. (See MPEP 2106.05(h)) Therefore, the claims are rejected on the same basis as claims 9, 19, and 29. Claim Rejections - 35 USC § 103 07-06 AIA 15-10-15 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. 07-20-aia AIA The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. 07-21-aia AIA Claim (s) 1, 4, 5, 8-11, 14, 15, 18-21, 24, 25, and 28-30 are rejected under 35 U.S.C. 103 as being unpatentable over Han Shi et al. (Herein referred to as Shi) (Bridging the Gap between Sample-based and One-shot Neural Architecture Search with BONAS) in view of Andrew Brock et al. (Herein referred to as Brock) (SMASH: One-Shot Model Architecture Search through HyperNetworks) and in further view of Jiahui Yu et al. (Herein referred to as Yu) (BigNAS: Scaling Up Neural Architecture Search with Big Single-Stage Models) Regarding claim 1, Shi teaches a processor-implemented method, comprising: generating an over-parameterized super network having a plurality of layers, the super network comprising a plurality of operator types, each layer of the plurality of layers comprising a largest super kernel corresponding to a search space (“In sample-based NAS algorithms, an architecture is selected in each iteration for full training [25, 33, 35, 42], and is computationally expensive. To alleviate this problem, we select in each BO iteration a batch of k architectures {(Ai,Xi)}k i=1 with the top-k UCB scores, and then train them together as a super-network by weight-sharing… As illustrated in Figure 1, during the query phase, we construct the super-network with adjacency matrix ˆ A =A1||A2||...||Ak, and feature matrix ˆ X =X1||X2||...||Xk, where || denotes the logical OR operation.”, pg. 6, under “3.2 Efficient Estimation of Candidate Performance”; See also Figure 1 on pg. 3) (Backpropagation implicitly utilizes a plurality of layers, with the backpropagation leading to a multi-layer super network, which comprises a largest super kernel correspond to the search space, and a plurality of matrices, corresponding to operators.) performing gradient descent to evolve the largest super kernel to a small kernel corresponding to the search space (“Neural Architecture Optimization (NAO) [19], which finds the architecture in a continuous embedding space with gradient descent… and(iv)ASNG-NAS[1], which proposes a stochastic natural gradient method for the NAS problem.”, pg. 7, bottom paragraph; pg. 8, under “4.3 Open Domain Search”) and determining a set of candidate architectures (“Since a NAS-Bench 101 architecture is obtained by stacking multiple repeated cells, we only consider the embedding of such a cell. Graph connectivity is encoded by the adjacency matrix A. Individual operations are encoded as one-hot vectors, and then aggregated to form the feature matrix X”, pg. 4, first paragraph of “3.1.1 Representing Neural Networks using GCN”) However, Shi does not teach to generate a range of kernel encodings, nor identifying a subset of kernel encodings from the range of kernel encodings, for each layer of the super network, based on the gradient descent , nor selecting a target model, from the set of architectures, based on meeting hardware specifications; and applying the target model. Shi does teach determining a set of candidate architectures but doesn't determine them based on the kernel encodings, each candidate architecture of the set of candidate architectures having a different model size. Brock teaches to generate a range of kernel encodings (“SMASH comprises two core components: the method by which we sample architectures, and the method by which we sample weights for a given architecture. For the former, we develop a memory bank view of feed-forward networks that permits sampling complex, branching topologies, and encoding said topologies as binary vectors. For the latter, we employ a HyperNet [12] that learns to map directly from the binary architecture encoding to the weight space.”, pg. 3, first paragraph) (The topology encodings correspond to a range of encodings.) and identifying a subset of kernel encodings from the range of kernel encodings, for each layer of the super network, based on the gradient descent , (“For a given architecture, we find that SMASH validation performance is consistently highest when using the correct encoding tensor, suggesting that the HyperNet has indeed learned a passable mapping from architecture to weights… we posit that if the HyperNet learns a meaningful mapping W=H(c), then the classification error E=f(W,x)=f(H(c),x) can be backpropagated to find dE/dc, providing an approximate measure of the error with respect to the architecture itself. If this holds true, then perturbing the architecture according to the dE/dc vector (within the constraints of our scheme) should allow us to guide the architecture search through a gradient descent-like procedure… The conditional embedding c is a one-hot encoding of the memory banks we read and write at each layer.”, pg. 7, first and second paragraph; pg. 12, third paragraph) (The one-hot encoding, c, is present at each layer, corresponding to a range of kernel encodings for each layer of the super network. The goal then being to search via a gradient decent-like procedure for an architecture that yields a subset of tensor encodings with the lowest error.) Therefore, it would have considered obvious to one of ordinary skill in the art, prior to the filing date of the current application, to combine the super network of Shi with the identification of kernel encoding, as described in Brock. One of ordinary skill in the art would have been motivated to combine the two teaching, prior to the application’s filing date, as this allows for better validation performance, and better learning of weight from architecture , as disclosed in Brock. (“For a given architecture, we find that SMASH validation performance is consistently highest when using the correct encoding tensor, suggesting that the HyperNet has indeed learned a passable mapping from architecture to weights”, pg. 7, first paragraph) However, the combination does not teach selecting a target model, from the set of architectures, based on meeting hardware specifications; and applying the target model. Shi, as modified by Brock, does teach determining a set of candidate architectures but doesn't determine them based on the kernel encodings, each candidate architecture of the set of candidate architectures having a different model size. Yu specifically teaches the use of kernel encodings to determine sets of candidate architectures each candidate architecture of the set of candidate architectures having a different model size. (“we first find rough skeletons of good candidate networks. Specifically, in the coarse selection phase, we pre-define five input resolutions (network-wise, {192, 224, 256, 288, 320}), four depth configurations (stage-wise via global depth multipliers [33]), two channel configurations (stage-wise via global width multipliers [17]) and four kernel size configurations (stage-wise), and obtain all of their benchmarks (shown in Figure 8 on the left). Then under our interested latency budget, we perform a fine-grained grid search by varying its configurations (shown in Figure 8 on the right).“, pg. 12 and 13; See also Figure 8 on pg. 13) (The different network resolutions correspond to different model sizes.) selecting a target model, from the set of architectures, based on meeting hardware specifications; (“We select architectures using a simple coarse-to-fine selection method to find the most accurate model under the given resource constraints (e.g., FLOPs, memory footprint and/or runtime latency budgets on different devices).”, pg. 4, under “Architecture Search with Single-Stage Models” step 2) and applying the target model (“We presented a novel paradigm for neural architecture search by training a single-stage model, from which high-quality child models of different sizes can be induced for instant deployment without retraining or finetuning.”, pg. 14, under “Conclusion”) Therefore, it would have considered obvious to one of ordinary skill in the art, prior to the filing date of the current application, to combine the super network of Shi, as modified by Brock, with to respect to hardware limitations and model deployment of Yu. One of ordinary skill in the art would have been motivated to combine the two teaching, prior to the application’s filing date, as Yu’s frame work offers a better coverage over diverse deployment scenarios and varied resource budgets, as disclosed by Yu. (“Our BigNAS is able to handle a wider set of models (from 200 MFLOPs to 1 GFLOPs) and offers a better coverage over diverse deployment scenarios and varied resource budgets.”, pg. 4, first paragraph) Regarding claim 11, Shi teaches An apparatus, comprising: at least one memory; and at least one processor coupled to the at least one memory, (One would implicitly need these components to run the method of Shi.) the at least one processor configured to: generate an over-parameterized super network having a plurality of layers, the super network comprising a plurality of operator types, each layer of the plurality of layers comprising a largest super kernel corresponding to a search space (“In sample-based NAS algorithms, an architecture is selected in each iteration for full training [25, 33, 35, 42], and is computationally expensive. To alleviate this problem, we select in each BO iteration a batch of k architectures {(Ai,Xi)}k i=1 with the top-k UCB scores, and then train them together as a super-network by weight-sharing… As illustrated in Figure 1, during the query phase, we construct the super-network with adjacency matrix ˆ A =A1||A2||...||Ak, and feature matrix ˆ X =X1||X2||...||Xk, where || denotes the logical OR operation.”, pg. 6, under “3.2 Efficient Estimation of Candidate Performance”; See also Figure 1 on pg. 3) (Backpropagation implicitly utilizes a plurality of layers, with the backpropagation leading to a multi-layer super network, which comprises a largest super kernel correspond to the search space, and a plurality of matrices, corresponding to operators.) perform gradient descent to evolve the largest super kernel to a small kernel corresponding to the search space (“Neural Architecture Optimization (NAO) [19], which finds the architecture in a continuous embedding space with gradient descent… and(iv)ASNG-NAS[1], which proposes a stochastic natural gradient method for the NAS problem.”, pg. 7, bottom paragraph; pg. 8, under “4.3 Open Domain Search”) and determine a set of candidate architectures (“Since a NAS-Bench 101 architecture is obtained by stacking multiple repeated cells, we only consider the embedding of such a cell. Graph connectivity is encoded by the adjacency matrix A. Individual operations are encoded as one-hot vectors, and then aggregated to form the feature matrix X”, pg. 4, first paragraph of “3.1.1 Representing Neural Networks using GCN”) However, Shi does not teach to generate a range of kernel encodings, nor to identify a subset of kernel encodings from the range of kernel encodings, for each layer of the super network, based on the gradient descent , nor to select a target model, from the set of architectures, based on meeting hardware specifications; and applying the target model. Shi does teach to determine a set of candidate architectures but doesn't determine candidate architectures based on the kernel encodings, each candidate architecture of the set of candidate architectures having a different model size. Brock teaches to generate a range of kernel encodings (“SMASH comprises two core components: the method by which we sample architectures, and the method by which we sample weights for a given architecture. For the former, we develop a memory bank view of feed-forward networks that permits sampling complex, branching topologies, and encoding said topologies as binary vectors. For the latter, we employ a HyperNet [12] that learns to map directly from the binary architecture encoding to the weight space.”, pg. 3, first paragraph) (The topology encodings correspond to a range of encodings.) and identify a subset of kernel encodings from the range of kernel encodings, for each layer of the super network, based on the gradient descent (“For a given architecture, we find that SMASH validation performance is consistently highest when using the correct encoding tensor, suggesting that the HyperNet has indeed learned a passable mapping from architecture to weights… we posit that if the HyperNet learns a meaningful mapping W=H(c), then the classification error E=f(W,x)=f(H(c),x) can be backpropagated to find dE/dc, providing an approximate measure of the error with respect to the architecture itself. If this holds true, then perturbing the architecture according to the dE/dc vector (within the constraints of our scheme) should allow us to guide the architecture search through a gradient descent-like procedure… The conditional embedding c is a one-hot encoding of the memory banks we read and write at each layer.”, pg. 7, first and second paragraph; pg. 12, third paragraph) (The one-hot encoding, c, is present at each layer, corresponding to a range of kernel encodings for each layer of the super network. The goal then being to search via a gradient decent-like procedure for an architecture that yields a subset of tensor encodings with the lowest error.) Therefore, it would have considered obvious to one of ordinary skill in the art, prior to the filing date of the current application, to combine the super network of Shi with the identification of kernel encoding, as described in Brock. One of ordinary skill in the art would have been motivated to combine the two teaching, prior to the application’s filing date, as this allows for better validation performance, and better learning of weight from architecture , as disclosed in Brock. (“For a given architecture, we find that SMASH validation performance is consistently highest when using the correct encoding tensor, suggesting that the HyperNet has indeed learned a passable mapping from architecture to weights”, pg. 7, first paragraph) However, the combination does not teach that the step to select a target model, from the set of architectures, based on meeting hardware specifications; and apply the target model. Shi, as modified by Brock, does teach to determine a set of candidate architectures but doesn't determine the candidate architectures based on the kernel encodings, each candidate architecture of the set of candidate architectures having a different model size. Yu specifically teaches the use of kernel encodings to determine sets of candidate architectures each candidate architecture of the set of candidate architectures having a different model size. (“we first find rough skeletons of good candidate networks. Specifically, in the coarse selection phase, we pre-define five input resolutions (network-wise, {192, 224, 256, 288, 320}), four depth configurations (stage-wise via global depth multipliers [33]), two channel configurations (stage-wise via global width multipliers [17]) and four kernel size configurations (stage-wise), and obtain all of their benchmarks (shown in Figure 8 on the left). Then under our interested latency budget, we perform a fine-grained grid search by varying its configurations (shown in Figure 8 on the right).“, pg. 12 and 13; See also Figure 8 on pg. 13) (The different network resolutions correspond to different model sizes.) select a target model, from the set of architectures, based on meeting hardware specifications; (“We select architectures using a simple coarse-to-fine selection method to find the most accurate model under the given resource constraints (e.g., FLOPs, memory footprint and/or runtime latency budgets on different devices).”, pg. 4, under “Architecture Search with Single-Stage Models” step 2) and apply the target model (“We presented a novel paradigm for neural architecture search by training a single-stage model, from which high-quality child models of different sizes can be induced for instant deployment without retraining or finetuning.”, pg. 14, under “Conclusion”) Therefore, it would have considered obvious to one of ordinary skill in the art, prior to the filing date of the current application, to combine the super network of Shi, as modified by Brock, with to respect to hardware limitations and model deployment of Yu. One of ordinary skill in the art would have been motivated to combine the two teaching, prior to the application’s filing date, as Yu’s frame work offers a better coverage over diverse deployment scenarios and varied resource budgets, as disclosed by Yu. (“Our BigNAS is able to handle a wider set of models (from 200 MFLOPs to 1 GFLOPs) and offers a better coverage over diverse deployment scenarios and varied resource budgets.”, pg. 4, first paragraph) Regarding claim 21, Shi teaches an apparatus, comprising: means for generating an over-parameterized super network having a plurality of layers the super network comprising a plurality of operator types, each layer of the plurality of layers comprising a largest super kernel corresponding to a search space; (“In sample-based NAS algorithms, an architecture is selected in each iteration for full training [25, 33, 35, 42], and is computationally expensive. To alleviate this problem, we select in each BO iteration a batch of k architectures {(Ai,Xi)}k i=1 with the top-k UCB scores, and then train them together as a super-network by weight-sharing… As illustrated in Figure 1, during the query phase, we construct the super-network with adjacency matrix ˆ A =A1||A2||...||Ak, and feature matrix ˆ X =X1||X2||...||Xk, where || denotes the logical OR operation.”, pg. 6, under “3.2 Efficient Estimation of Candidate Performance”; See also Figure 1 on pg. 3) (Backpropagation implicitly utilizes a plurality of layers, with the backpropagation leading to a multi-layer super network, which comprises a largest super kernel correspond to the search space, and a plurality of matrices, corresponding to operators.) means for performing gradient descent to evolve the largest super kernel to a small kernel corresponding to the search space (“Neural Architecture Optimization (NAO) [19], which finds the architecture in a continuous embedding space with gradient descent… and(iv)ASNG-NAS[1], which proposes a stochastic natural gradient method for the NAS problem.”, pg. 7, bottom paragraph; pg. 8, under “4.3 Open Domain Search”) and means for determining a set of candidate architectures. (“Since a NAS-Bench 101 architecture is obtained by stacking multiple repeated cells, we only consider the embedding of such a cell. Graph connectivity is encoded by the adjacency matrix A. Individual operations are encoded as one-hot vectors, and then aggregated to form the feature matrix X”, pg. 4, first paragraph of “3.1.1 Representing Neural Networks using GCN”) However, Shi does not teach to generate a range of kernel encodings, nor means for identifying a subset of kernel encodings from the range of kernel encodings, for each layer of the super network, based on the gradient descent , nor means for selecting a target model, from the set of architectures, based on meeting hardware specifications; and means for applying the target model. Shi does teach means for determining a set of candidate architectures but doesn't determine them based on the kernel encodings, each candidate architecture of the set of candidate architectures having a different model size. Brock teaches to generate a range of kernel encodings (“SMASH comprises two core components: the method by which we sample architectures, and the method by which we sample weights for a given architecture. For the former, we develop a memory bank view of feed-forward networks that permits sampling complex, branching topologies, and encoding said topologies as binary vectors. For the latter, we employ a HyperNet [12] that learns to map directly from the binary architecture encoding to the weight space.”, pg. 3, first paragraph) (The topology encodings correspond to a range of encodings.) and means for identifying a subset of kernel encodings from the range of kernel encodings, for each layer of the super network, based on the gradient descent , (“For a given architecture, we find that SMASH validation performance is consistently highest when using the correct encoding tensor, suggesting that the HyperNet has indeed learned a passable mapping from architecture to weights… we posit that if the HyperNet learns a meaningful mapping W=H(c), then the classification error E=f(W,x)=f(H(c),x) can be backpropagated to find dE/dc, providing an approximate measure of the error with respect to the architecture itself. If this holds true, then perturbing the architecture according to the dE/dc vector (within the constraints of our scheme) should allow us to guide the architecture search through a gradient descent-like procedure… The conditional embedding c is a one-hot encoding of the memory banks we read and write at each layer.”, pg. 7, first and second paragraph; pg. 12, third paragraph) (The one-hot encoding, c, is present at each layer, corresponding to a range of kernel encodings for each layer of the super network. The goal then being to search via a gradient decent-like procedure for an architecture that yields a subset of tensor encodings with the lowest error.) Therefore, it would have considered obvious to one of ordinary skill in the art, prior to the filing date of the current application, to combine the super network of Shi with the identification of kernel encoding, as described in Brock. One of ordinary skill in the art would have been motivated to combine the two teaching, prior to the application’s filing date, as this allows for better validation performance, and better learning of weight from architecture , as disclosed in Brock. (“For a given architecture, we find that SMASH validation performance is consistently highest when using the correct encoding tensor, suggesting that the HyperNet has indeed learned a passable mapping from architecture to weights”, pg. 7, first paragraph) However, the combination does not teach means for selecting a target model, from the set of architectures, based on meeting hardware specifications; and means for applying the target model. Shi, as modified by Brock, does teach means for determining a set of candidate architectures but doesn't determine them based on the kernel encodings, each candidate architecture of the set of candidate architectures having a different model size. Yu specifically teaches the use of kernel encodings to determine sets of candidate architectures each candidate architecture of the set of candidate architectures having a different model size. (“we first find rough skeletons of good candidate networks. Specifically, in the coarse selection phase, we pre-define five input resolutions (network-wise, {192, 224, 256, 288, 320}), four depth configurations (stage-wise via global depth multipliers [33]), two channel configurations (stage-wise via global width multipliers [17]) and four kernel size configurations (stage-wise), and obtain all of their benchmarks (shown in Figure 8 on the left). Then under our interested latency budget, we perform a fine-grained grid search by varying its configurations (shown in Figure 8 on the right).“, pg. 12 and 13; See also Figure 8 on pg. 13) (The different network resolutions correspond to different model sizes.) means for selecting a target model, from the set of architectures, based on meeting hardware specifications; (“We select architectures using a simple coarse-to-fine selection method to find the most accurate model under the given resource constraints (e.g., FLOPs, memory footprint and/or runtime latency budgets on different devices).”, pg. 4, under “Architecture Search with Single-Stage Models” step 2) and means for applying the target model (“We presented a novel paradigm for neural architecture search by training a single-stage model, from which high-quality child models of different sizes can be induced for instant deployment without retraining or finetuning.”, pg. 14, under “Conclusion”) Therefore, it would have considered obvious to one of ordinary skill in the art, prior to the filing date of the current application, to combine the super network of Shi, as modified by Brock, with to respect to hardware limitations and model deployment of Yu. One of ordinary skill in the art would have been motivated to combine the two teaching, prior to the application’s filing date, as Yu’s frame work offers a better coverage over diverse deployment scenarios and varied resource budgets, as disclosed by Yu. (“Our BigNAS is able to handle a wider set of models (from 200 MFLOPs to 1 GFLOPs) and offers a better coverage over diverse deployment scenarios and varied resource budgets.”, pg. 4, first paragraph) Regarding claims 4, 14, and 24, Shi, as modified by Brock and Yu, teaches the method and apparatuses of claims 1, 11, and 21, as well as the plurality of operating types comprises at least one convolutional neural network, at least one recurrent neural network, at least one transformer, and/or at least one multilayer perceptron (MLP). (“In the search phase, we first use a graph convolutional network (GCN) [13] to produce embeddings for the neural architectures.”, pg. 2, third paragraph (Shi)) Regarding claims 5, 15, and 25, Shi, as modified by Brock and Yu, teaches the method and apparatuses of claims 1, 11, and 21, as well as the super network is based on a particular use case, a different super network constructed for each different use case. (“In the query phase, we construct a super-network from a batch of promising candidate architectures, and train them by uniform sampling. These candidates are then queried simultaneously based on the learned weight of the super-network.”, pg. 2, third paragraph (Shi)) (A different super network is constructed from different candidate architectures, which corresponds to different use cases.) Regarding claims 8, 18, and 28, Shi, as modified by Brock and Yu, teaches the method and apparatuses of claims 1, 11, and 21, as well as training the target model. (“To alleviate this problem, we select in each BO iteration a batch of k architectures {(Ai,Xi)}k i=1 with the top-k UCB scores, and then train them together as a super-network by weight-sharing.”, pg. 6, under “3.2 Efficient Estimation of Candidate Performance”) Regarding claims 9, 19, and 29, Shi, as modified by Brock and Yu, teaches the method and apparatuses of claims 1, 11, and 21, as well as, a one-shot neural architecture search (NAS). (“In line with our claim of "one-shot" model search, we keep our exploration of the SMASH design space to a minimum.”, pg. 13, fifth paragraph of “Appendix C: Experiment Details” (Brock)) Regarding claims 10, 20, and 30, Shi, as modified by Brock and Yu, teaches the method and apparatuses of claims 9, 19, and 29, as well as, the one-shot NAS is based on at least one of: memory usage of a device that applies the target model, run time latency of the model running on the device, a model accuracy, and power consumption of the device. (“Once an architecture is selected, we can obtain a child model by simply slicing the single-stage model for instant deployment w.r.t. the given constraints such as memory footprint and/or runtime latency.”, pg. 3, first paragraph (Yu)) 07-21-aia AIA Claim (s) 2, 12, and 22 are rejected under 35 U.S.C. 103 as being unpatentable over Shi, in view of Brock, in further view of Yu, and in further view of Song Han et al. (Herein referred to as Han) (DEEP COMPRESSION: COMPRESSING DEEP NEURAL NETWORKS WITH PRUNING, TRAINED QUANTIZATION AND HUFFMAN CODING) Regarding claims 2, 12, and 22, Shi, as modified by Brock and Yu, teaches the method and apparatuses of claims 1, 11, and 21, but does not explicitly teach quantizing the target model. Han teaches quantizing the target model. (“Suppose we have a layer that has 4 input neurons and 4 output neurons, the weight is a 4 × 4 matrix. On the top left is the 4 × 4 weight matrix, and on the bottom left is the 4 × 4 gradient matrix. The weights are quantized to 4 bins (denoted with 4 colors), all the weights in the same bin share the same value, thus for each weight, we then need to store only a small index into a table of shared weights.”, pg. 3, fourth paragraph) Therefore, it would have considered obvious to one of ordinary skill in the art, prior to the filing date of the current application, to combine the super network, candidate architectures and kernel encodings of Shi, as modified by Brock and Yu, with the network quantization of Han. One of ordinary skill in the art would have been motivated to combine the teaching, prior to the filing date of the current application, as quantization of weights enforces weight sharing and reduces the number of bits, as disclosed in Han. (“we quantize the weights to enforce weight sharing, finally, we apply Huffman coding. After the first two steps we retrain the network to fine tune the remaining connections and the quantized centroids. Pruning, reduces the number of connections by 9× to 13×; Quantization then reduces the number of bits that represent each connection from 32 to 5.”, pg. 1, Abstract) 07-21-aia AIA Claim (s) 3, 13, and 23 are rejected under 35 U.S.C. 103 as being unpatentable over Shi, in view of Brock, in further view of Yu, and in further view of Ashish Gondimalla et al. (Herein referred to as Gondimalla) (SparTen: A Sparse Tensor Accelerator for Convolutional Neural Networks) Regarding claims 3, 13, and 23, Shi, as modified by Brock and Yu, teaches the method and apparatuses of claims 1, 11, and 21, but does not explicitly teach splitting convolution node activations of the target model by splitting each input tensor into a plurality of smaller tensors, the plurality of smaller tensors incorporated into the set of candidate architectures. Gondimalla teaches splitting convolution node activations of the target model by splitting each input tensor into a plurality of smaller tensors, the plurality of smaller tensors incorporated into the set of candidate architectures. (“SparTen leverages its efficient inner join to confine the products for one output cell to one multiplier and distribute those for different output cells on different multipliers for parallelism. Consequently, SparTen avoids SCNN’s problems. Because SparTen’s inner join produces different output cells in separate compute units, SparTen applies to convolutions of any stride (and also to non-convolutional DNNs) and avoids SCNN’s cross bars.”, pg. 6, right column, second paragraph) (The separate compute unit correspond to smaller tensors, which can be easily configured to work with the candidate architectures of the modified Shi.) Therefore, it would have considered obvious to one of ordinary skill in the art, prior to the filing date of the current application, to combine the super network, candidate architectures and kernel encodings of Shi, as modified by Brock and Yu, with the separate tensors of Gondimalla. One of ordinary skill in the art would have been motivated to combine the teaching, prior to the filing date of the current application, as Gondimalla’s method requires just one address calculation, and allows for more room for input buffering, among other benefits, as Gondimalla discloses. (“Each of SparTen’s compute units produces only one output cell at a time (1) requiring just one address calculation for all the products in a chunk and (2) allowing more room for input buffering and hence fewer implicit barriers at input broadcast. Producing a single output cell also means that a chunk can capture many channels for smaller filters to achieve high compute-unit utilization without increasing output buffering as all the channels contribute to the same output cell. Because SparTen assigns one output cell per compute unit there is no underutilization due to input tiling.”, pg. 6, right column, right before “3.3 Greedy balancing”) 07-21-aia AIA Claim s 6, 7, 16, 17, 26, and 27 are rejected under 35 U.S.C. 103 as being unpatentable over Shi, in view of Brock, in further view of Yu, and in further view of Chris Ying et al. (Herein referred to as Ying) (NAS-Bench-101: Towards Reproducible Neural Architecture Search) Regarding claims 6, 16, and 26, Shi, as modified by Brock and Yu, teaches the method and apparatuses of claims 1, 11, and 21, but does not explicitly teach identifying the subset of kernel encodings comprises means for progressively increasing a selected kernel choice from a smallest inner core of the largest super kernel to the largest super kernel. Ying teaches identifying the subset of kernel encodings comprises means for progressively increasing a selected kernel choice from a smallest inner core of the largest super kernel to the largest super kernel. (“The goal of NAS algorithms is to find architectures that have high testing accuracy at epoch Emax. To do this, we repeatedly query the dataset at (A,Estop) pairs, where A is an architecture in the search space and Estop is an allowed number of epochs (Estop ∈ {4,12,36,108}).”, pg. 3, right column, bottom paragraph) (The search space is search progressively with the increased number of epochs.) Therefore, it would have considered obvious to one of ordinary skill in the art, prior to the filing date of the current application, to combine the super network, candidate architectures and kernel encodings of Shi, as modified by Brock and Yu, with the progressive search of a search space of Ying. One of ordinary skill in the art would have been motivated to combine the teaching, prior to the filing date of the current application, as this allows for the NAS algorithm to find architectures that have high testing accuracy, as disclosed in Ying. (“The goal of NAS algorithms is to find architectures that have high testing accuracy at epoch Emax. To do this, we repeatedly query the dataset at (A,Estop) pairs, where A is an architecture in the search space and Estop is an allowed number of epochs (Estop ∈ {4,12,36,108}).”, pg. 3, right column, bottom paragraph) Regarding claims 7, 17, and 27, Shi, as modified by Brock and Yu, teaches the method and apparatuses of claims 1, 11, and 21, but does not explicitly teach identifying the subset of kernel encodings comprises means for progressively reducing a selected kernel choice from the largest super kernel to an inner core of the largest super kernel. Ying teaches identifying the subset of kernel encodings comprises means for progressively reducing a selected kernel choice from the largest super kernel to an inner core of the largest super kernel. (“Because NAS-Bench-101 exhaustively evaluates a search space, it permits, for the first time, a comprehensive analysis of a NAS search space as a whole. We illustrate such potential by measuring search space properties relevant to architecture search… we restrict our search for neural net topologies to the space of small feedforward structures, usually called cells…”, pg. 1, right column, fourth paragraph; pg. 2, under “2.1. Architecture”) (The search is restricted to certain cells from the search space as a whole.) Therefore, it would have considered obvious to one of ordinary skill in the art, prior to the filing date of the current application, to combine the super network, candidate architectures and kernel encodings of Shi, as modified by Brock and Yu, with the reduction of a search space of Ying. One of ordinary skill in the art would have been motivated to combine the teaching, prior to the filing date of the current application, as this allows for exhaustive enumeration, as disclosed in Ying. (“this space of labeled DAGs grows exponentially in both V and L. In order to limit the size of the space to allow exhaustive enumeration, we impose the following constraints…”, pg. 2, left column, last paragraph) Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to Tyler E Iles whose telephone number is (571)272-5442. The examiner can normally be reached 9:00am - 5:00pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kakali Chaki can be reached at (571) 272-3719. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /T.E.I./ Patent Examiner, Art Unit 2122 /KAKALI CHAKI/ Supervisory Patent Examiner, Art Unit 2122 Application/Control Number: 18/363,487 Page 2 Art Unit: 2122 Application/Control Number: 18/363,487 Page 3 Art Unit: 2122 Application/Control Number: 18/363,487 Page 4 Art Unit: 2122 Application/Control Number: 18/363,487 Page 5 Art Unit: 2122 Application/Control Number: 18/363,487 Page 6 Art Unit: 2122 Application/Control Number: 18/363,487 Page 7 Art Unit: 2122 Application/Control Number: 18/363,487 Page 8 Art Unit: 2122 Application/Control Number: 18/363,487 Page 9 Art Unit: 2122 Application/Control Number: 18/363,487 Page 10 Art Unit: 2122 Application/Control Number: 18/363,487 Page 11 Art Unit: 2122 Application/Control Number: 18/363,487 Page 12 Art Unit: 2122 Application/Control Number: 18/363,487 Page 13 Art Unit: 2122 Application/Control Number: 18/363,487 Page 14 Art Unit: 2122 Application/Control Number: 18/363,487 Page 15 Art Unit: 2122 Application/Control Number: 18/363,487 Page 16 Art Unit: 2122 Application/Control Number: 18/363,487 Page 17 Art Unit: 2122 Application/Control Number: 18/363,487 Page 18 Art Unit: 2122 Application/Control Number: 18/363,487 Page 19 Art Unit: 2122 Application/Control Number: 18/363,487 Page 20 Art Unit: 2122 Application/Control Number: 18/363,487 Page 21 Art Unit: 2122 Application/Control Number: 18/363,487 Page 22 Art Unit: 2122 Application/Control Number: 18/363,487 Page 23 Art Unit: 2122 Application/Control Number: 18/363,487 Page 24 Art Unit: 2122 Application/Control Number: 18/363,487 Page 25 Art Unit: 2122 Application/Control Number: 18/363,487 Page 26 Art Unit: 2122 Application/Control Number: 18/363,487 Page 27 Art Unit: 2122 Application/Control Number: 18/363,487 Page 28 Art Unit: 2122 Application/Control Number: 18/363,487 Page 29 Art Unit: 2122 Application/Control Number: 18/363,487 Page 30 Art Unit: 2122 Application/Control Number: 18/363,487 Page 31 Art Unit: 2122 Application/Control Number: 18/363,487 Page 32 Art Unit: 2122 Application/Control Number: 18/363,487 Page 33 Art Unit: 2122 Application/Control Number: 18/363,487 Page 34 Art Unit: 2122
Read full office action

Prosecution Timeline

Aug 01, 2023
Application Filed
May 14, 2026
Non-Final Rejection mailed — §101, §103, §112
Jul 15, 2026
Applicant Interview (Telephonic)
Jul 15, 2026
Examiner Interview Summary
Jul 15, 2026
Response Filed
Sep 29, 2026
Final Rejection mailed — §101, §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12748978
EARLY STOPPING METHOD FOR NEURAL NETWORK USING UNLABELED DATA
3y 8m to grant Granted Sep 29, 2026
Patent 12664410
METHODS AND DEVICES FOR ACCELERATING A TRANSFORMER WITH A SPARSE ATTENTION PATTERN
4y 7m to grant Granted Jun 23, 2026
Patent 12619883
SYSTEMS AND METHODS FOR DETERMINING TIME-SERIES FEATURE IMPORTANCE OF A MODEL
4y 4m to grant Granted May 05, 2026
Study what changed to get past this examiner. Based on 3 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
56%
Grant Probability
99%
With Interview (+66.7%)
3y 7m (~5m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 9 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month