Prosecution Insights
Last updated: August 18, 2026
Application No. 17/887,021

ELECTRONIC DEVICE AND METHOD WITH SENSITIVITY-BASED QUANTIZED TRAINING AND OPERATION

Non-Final OA §103
Filed
Aug 12, 2022
Priority
Mar 15, 2022 — RE 10-2022-0032221
Examiner
NAULT, VICTOR ADELARD
Art Unit
2124
Tech Center
2100 — Computer Architecture & Software
Assignee
Samsung Electronics Co., Ltd.
OA Round
3 (Non-Final)
53%
Grant Probability
Moderate
3-4
OA Rounds
0m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 53% of resolved cases
53%
Career Allowance Rate
9 granted / 17 resolved
-2.1% vs TC avg
Strong +66% interview lift
Without
With
+65.7%
Interview Lift
resolved cases with interview
Typical timeline
3y 10m
Avg Prosecution
18 currently pending
Career history
46
Total Applications
across all art units

Statute-Specific Performance

§101
28.6%
-11.4% vs TC avg
§103
42.8%
+2.8% vs TC avg
§102
7.6%
-32.4% vs TC avg
§112
19.9%
-20.1% vs TC avg
Black line = Tech Center average estimate • Based on career data from 17 resolved cases

Office Action

§103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Continued Examination Under 37 CFR 1.114 A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 04/27/2026 has been entered. Remarks This Office Action is responsive to Applicants' Amendment filed on April 27, 2026, in which claims 1, 5, 12, and 13 are amended. Claim 2 has been cancelled. No claims have been newly added. Claims 1 and 3-21 are currently pending. Response to Arguments With regards to the objection to claim 13 for a minor informality, claim 13 has been amended to remove the minor informality, and thus the objection is withdrawn. With regards to the rejections of claims 1-4 and 6-20 under 35 U.S.C. 101 as directed towards abstract ideas, Applicant’s argument that the claims as amended overcome the rejections are persuasive. Independent claims 1 and 12 now recite a limitation of processing the layer with the low sensitivity lower than the predetermined threshold with a first precision by quantizing the layer during the at least one of the operation of backward propagation or the operation of updating the weight, which, although it recites a mathematical concept of quantization, integrates into a practical application at Step 2A, Prong 2 by improving a machine learning model with reduced accuracy degradation, which would otherwise result from uniform quantization. With regards to the rejections of claims 1-3, 6-8, 11, 14, and 18 under 35 U.S.C. 103 as being unpatentable over Bijalwan, in view of Shen, further in view of Xu, Applicant argues that at least claim 1 as amended is not taught by the prior combination of art, as “none of these references teaches or suggests applying sensitivity-based quantization differentially during backward propagation or weight updating”. Applicant’s arguments that the claims as amended overcome the rejections are persuasive, however the arguments are moot in view of a new grounds of rejection, necessitated by Applicant’s amendments to the claims, as presented below. Additionally, Applicant argues on page 12 of the Remarks that “Bijalwan teaches away from the use of mixed precision quantization. Bijalwan teaches a post-training quantization approach”. Examiner respectfully disagrees. Bijalwan states: (Bijalwan [0038]) “Quantization of the neural network model may be performed using two techniques - Post-training quantization and Quantization-aware training…Methods in the present disclosure preferably employ Quantization-aware training technique”. Prior Art The following references are used for prior art claim rejections: Bijalwan et al. (U.S. Patent Application Publication No. 2023/0281423), hereinafter Bijalwan Shen et al. (U.S. Patent Application Publication No. 2022/0129736), hereinafter Shen Xu et al. (U.S. Patent Application Publication No. 2021/0295168), hereinafter Xu Li et al. “Training Quantized Nets: A Deeper Understanding”, hereinafter Li Dong et al. “HAWQ-V2: Hessian Aware trace-Weighted Quantization of Neural Networks”, hereinafter Dong Liu et al. “Post-training Quantization with Multiple Points: Mixed Precision without Mixed Precision”, hereinafter Liu Micikevicius et al. “Mixed Precision Training”, hereinafter Micikevicius Alistarh et al. “The Convergence of Sparsified Gradient Methods” hereinafter Alistarh Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1, 3, 6-8, and 11 are rejected under 35 U.S.C. 103 as being unpatentable over Bijalwan, in view of Shen, further in view of Xu, further in view of Li. Regarding claim 1, Bijalwan teaches An electronic device comprising: ((Bijalwan [0028]) “In one embodiment, the system 100 comprises…at least one device such as a computing device 104”) a processor; ((Bijalwan [0009]) “The system comprises a memory and a processor that is coupled to the memory”) and a memory configured to store instructions executable by the processor, ((Bijalwan [0009]) “The system comprises a memory and a processor that is coupled to the memory”) wherein the processor is configured to, in response to the instructions being executed by the processor: generate, based on a determination of sensitivity of layers in a model to be trained, sensitivity results; ((Bijalwan [0055]) At block 408, the sensitivity evaluation module 114 normalizes the features sensitivity values and the weight sensitivity values of the plurality of layers independently and combines the sensitivity values into a union sensitivity list…The first row of the union sensitivity list illustrates that a weight sensitivity value corresponding to layer 1 indicated by "Layer_l_w" is 0.77) Shen teaches the following further limitations more explicitly than Bijalwan: and perform one or more operations [of a distributed training] on the model by applying quantization to a layer of the layers with a low sensitivity of the sensitivity results lower than a predetermined threshold, ((Shen [0028]) “when the value of the objective function corresponding to the first layer L1 is greater than the threshold, this indicates that the loss is small, and the processing unit 120 decides to quantize the first layer L1 with the second precision”, a small loss corresponds to a low sensitivity, Shen does not teach distributed training) wherein, to perform the application of the quantization, the processor is further configured to: process the layer with the low sensitivity lower than the predetermined threshold with a first precision by quantizing the layer … ((Shen [0028]) “when the value of the objective function corresponding to the first layer L1 is greater than the threshold, this indicates that the loss is small, and the processing unit 120 decides to quantize the first layer L1 with the second precision”, a small loss corresponds to a low sensitivity) and process, without quantization, a layer with a high sensitivity of the sensitivity results higher than or equal to the predetermined threshold with a second precision, higher than the first precision, … ((Shen [0028]) “when the values of the objective function corresponding to the second layer L2 and the third layer L3 is not greater than the threshold, this indicates that the loss is large, and the processing unit 120…does not quantize the second layer L2 and the third layer L3 (that is, the second layer L2 and the third layer L3 remain at the first precision)”, a large loss for a layer corresponds to a high sensitivity for the layer) At the time of filing, one of ordinary skill in the art would have motivation to combine Bijalwan and Shen by taking the electronic device for determining the sensitivity of layers in a model taught by Bijalwan and applying a threshold to layers based on sensitivity to determine whether to apply quantization, taught by Shen, as doing so imparts the benefit of optimizing the memory efficiency and speed of the model, increased by quantizing some layers, with respect to the accuracy of the model, increased by refraining from quantizing some layers, by quantizing the layers that are not sensitive enough to significantly decrease accuracy. Such a combination would be obvious. Xu teaches the following further limitations that neither Bijalwan nor Shen teaches: and perform one or more operations of a distributed training… ((Xu [0025]) “A distributed system can accelerate a training process by distributing the training process across multiple computing systems, which can be referred to as worker nodes. Training data can be split into multiple portions, with each portion to be processed by a worker node. Each worker node can perform the forward and backward propagation operations independently”) where the one or more operations include at least one of an operation of backward propagation of the layer or an operation of updating a weight of the model dependent on a calculated gradient ((Xu [0023]) “As part of the training process, each neural network layer can then perform a backward propagation process to adjust the set of weights at each neural network layer. Specifically, the highest neural network layer can receive the set of gradients and compute, in a backward propagation operation, a set of first data gradients and a set of first weight gradients based on applying the set of weights to the input data gradients in similar mathematical operations as the forward propagation operation. The highest neural network layer can adjust the set of weights of the layer based on the set of first weight gradients”) At the time of filing, one of ordinary skill in the art would have motivation to combine Bijalwan, Shen, and Xu by taking the electronic device for determining the sensitivity of layers in a model and applying quantization to layers in the model that are insensitive to quantization relative to a threshold, taught jointly by Bijalwan and Shen, and adding the method of distributed training taught by Xu, as Xu teaches: (Xu [0025]) “A distributed system can accelerate a training process by distributing the training process across multiple computing systems”. Such a combination would be obvious. Li teaches the following further limitations that neither nor Shen, nor Xu teaches and that Bijalwan does not teach explicitly: process the layer…by quantizing the layer… ((Li Pg. 8) “we only quantize the weights in the convolutional layers, but not linear layers, during training”) during the at least one of the operation of backward propagation or the operation of updating the weight; ((Li Pg. 3) “The deterministic rounding SGD maintains quantized weights with updates of the form: [Equation 4], where wb denotes the low-precision weights, which are quantized using Qd immediately after applying the gradient descent update”, Li Pg. 3, Equation 4 shows that the updated weight at the next time step is the result of quantizing the weight at the previous time step after a gradient descent update) PNG media_image1.png 28 533 media_image1.png Greyscale and process, without quantization, a layer…during the at least one of the operation of backward propagation or the operation of updating the weight, ((Li Pg. 8) “we only quantize the weights in the convolutional layers, but not linear layers, during training”, training neural networks includes weight update operations) and wherein the application of the quantization is applied during the operation of updating the weight ((Li Pg. 3) “The deterministic rounding SGD maintains quantized weights with updates of the form: [Equation 4], where wb denotes the low-precision weights, which are quantized using Qd immediately after applying the gradient descent update”, Li Pg. 3, Equation 4 shows that the updated weight at the next time step is the result of quantizing the weight at the previous time step after a gradient descent update) At the time of filing, one of ordinary skill in the art would have motivation to combine Bijalwan, Shen, Xu, and Li by taking the electronic device for determining the sensitivity of layers in a model, applying quantization to layers in the model that are insensitive to quantization relative to a threshold, and training the model with distributed training, taught jointly by Bijalwan, Shen, and Xu, and applying the quantization of layers during weight update operations, taught by Li, as Li teaches: (Li Pg. 2) “For training quantized NNs from scratch, many authors suggest maintaining a high-precision floating point copy of the weights while feeding quantized weights into backprop [5, 11, 3, 16], which results in good empirical performance. There are limitations in using such methods on low-power devices, however, where floating-point arithmetic is not always available or not desirable”, that is, that directly quantizing the weights during the update step without maintaining any high-precision floating point weights intermediately has advantages in a low-power device environment. Such a combination would be obvious. Regarding claim 3, Bijalwan, Shen, Xu, and Li jointly teach The electronic device of claim 1, wherein the distributed training one or more operations include: Xu further teaches: performing forward propagation moving from a first layer to a last layer of the model; ((Xu [0021]) “During training of a neural network, a first neural network layer can receive training input data, combine the training input data with the weights (e.g., by multiplying the training input data with the weights and then summing the products) to generate first output data for the neural network layer, and propagate the output data to a second neural network layer, in a forward propagation operation…The forward propagation operations can start at the first neural network layer and end at the highest neural network layer”) performing the backward propagation moving from the last layer to the first layer of the model; ((Xu [0023]) “As part of the training process, each neural network layer can then perform a backward propagation process to adjust the set of weights at each neural network layer…The backward propagation operations can start from the highest neural network layer and end at the first neural network layer”) determining a mean value of gradients calculated in each of a plurality of nodes used for the distributed training of the model; ((Xu [0025]) “Each worker node can exchange its set of weight gradients with other worker nodes, and average its set of weight gradients and the sets of weight gradients received from other worker nodes”) performing the updating of the weight of the model based on the mean value ((Xu [0025]) “Each computing node can have the same set of averaged weight gradients, and can then update a set of weights for each neural network layer based on the averaged weight gradients”) At the time of filing, one of ordinary skill in the art would have motivation to combine the method jointly taught by Bijalwan, Shen, Xu, and Li for the parent claim of claim 3, claim 1. No new embodiments are introduced, so the reason to combine is the same as for the parent claim. Regarding claim 6, Bijalwan, Shen, Xu, and Li jointly teach The electronic device of claim 1, wherein the processor is further configured to: Bijalwan further teaches: classify the sensitivity results of the layers into a plurality of levels; ((Bijalwan [0037]) “The grouping module 116 may cluster the plurality of layers within the input neural network model into a plurality of groups based on the union sensitivity list”) and train the model by applying quantization to each of the layers with a precision at a level corresponding to each of the plurality of levels ((Bijalwan [0037]) “In one embodiment, the grouping module 116 clusters the plurality of layers into a plurality of groups to quantize each group into a high precision format. In another embodiment, the grouping module 116 clusters the plurality of layers into another plurality of groups to quantize each group into a lower precision format”) At the time of filing, one of ordinary skill in the art would have motivation to combine the method jointly taught by Bijalwan, Shen, Xu, and Li for the parent claim of claim 6, claim 1. No new embodiments are introduced, so the reason to combine is the same as for the parent claim. Regarding claim 7, Bijalwan, Shen, Xu, and Li jointly teach The electronic device of claim 3, Xu further teaches: wherein the processor is further configured to train the model by applying quantization [to the layer with the low sensitivity lower than the predetermined threshold] in any one or any combination of the operations of the distributed training ((Xu [0065]) “In some embodiments, a DMA controller 1002 at the worker node 120-1 may perform one or more of the compression tasks. For example, the DMA controller 1002 may utilize a gradient compression engine (GCE) to perform sparsity analysis, gradient clipping, quantization, and/or compression”, Bijalwan teaches quantization of a layer with low sensitivity below a threshold) At the time of filing, one of ordinary skill in the art would have motivation to combine the method jointly taught by Bijalwan, Shen, Xu, and Li for the parent claim of claim 7, claim 3. No new embodiments are introduced, so the reason to combine is the same as for the parent claim. Regarding claim 8, Bijalwan, Shen, Xu, and Li jointly teach The electronic device of claim 7, wherein the processor is further configured to compress data used in any one or any combination of the operations of the distributed training ((Xu [0059]) “After uncompressed gradients 806 are computed by the transmitting worker node 802, but prior to transmission, the uncompressed gradients 806 are compressed by a compression module 812 at the transmitting worker node 802 to generate compressed gradients 808. The compressed gradients 808 are then transmitted from the transmitting worker node 802 to the receiving worker node 804”) At the time of filing, one of ordinary skill in the art would have motivation to combine the method jointly taught by Bijalwan, Shen, Xu, and Li for the parent claim of claim 8, claim 7. No new embodiments are introduced, so the reason to combine is the same as for the parent claim. Regarding claim 11, Bijalwan, Shen, Xu, and Li jointly teach The electronic device of claim 1, Bijalwan further teaches: wherein the model to be trained [in the distributed training] is pretrained, [before the distributed training,] with a precision without quantization ((Bijalwan [0036]) “The sensitivity evaluation module 114 receives data from the data acquisition module 230 and generates a union sensitivity list. In one embodiment, the sensitivity evaluation module 114 generates a base model from the input neural network model by representing the parameters of the input neural network model in high precision format and stores as base model 210”, a model in a high precision format corresponds to a model pretrained with a precision without quantization, Xu but not Bijalwan teaches distributed training) At the time of filing, one of ordinary skill in the art would have motivation to combine the method jointly taught by Bijalwan, Shen, Xu, and Li for the parent claim of claim 11, claim 1. No new embodiments are introduced, so the reason to combine is the same as for the parent claim. Claim 4 is rejected under 35 U.S.C. 103 as being unpatentable over Bijalwan, in view of Shen, further in view of Xu, further in view of Li, further in view of Dong. Regarding claim 4, Bijalwan, Shen, Xu, and Li jointly teach The electronic device of claim 1, Dong teaches the following further limitation that neither Bijalwan, nor Shen, nor Xu, nor Li explicitly teaches: wherein the processor is further configured to periodically determine training sensitivity of the layers for each training of the model, or for each epoch or each of one or more iterations performed during the training of the model ((Dong Pgs. 1-2)) “these searching methods can require a large amount of computational resources, are time-consuming, and, worst of all, the quality of quantization is very sensitive to the initialization of their search parameters and therefore unpredictable. This makes deployment of these methods in online learning scenarios especially challenging, as in these applications a new model is trained every few hours and needs to be quantized for efficient inference. To address these issues, recent work introduced HAWQ [7], a Hessian AWare Quantization framework. The main idea is to assign higher bit-precision to layers that are more sensitive, and lower bit-precision to less sensitive layers”, application of the method of determining training sensitivity of layers to online learning scenarios where a new model is trained and quantized every few hours corresponds to periodically determining training sensitivity of the layers for each training of the model) At the time of filing, one of ordinary skill in the art would have motivation to combine Bijalwan, Shen, Xu, Li, and Dong by taking the electronic device for determining the sensitivity of layers in a model and applying quantization to layers in the model that are insensitive to quantization relative to a threshold during weight updates, taught jointly by Bijalwan, Shen, Xu, and Li, and adding quantizing the layers, including determining the sensitivity of the layers to quantization, every time the model is trained, as taught by Dong, as doing so increases the effectiveness of the model by optimizing its accuracy relative to its size for the newly trained parameters. Such a combination would be obvious. Claim 9 is rejected under 35 U.S.C. 103 as being unpatentable over Bijalwan, in view of Shen, further in view of Xu, further in view of Li, further in view of Micikevicius. Regarding claim 9, Bijalwan, Shen, Xu, and Li jointly teach The electronic device of claim 1, Micikevicus teaches the following further limitation that neither Bijalwan, nor Shen, nor Xu, nor Li teaches: wherein the processor is further configured to train the model by scaling a gradient calculated in training the model ((Micikevicius Pg. 4) “One efficient way to shift the gradient values into FP16-representable range is to scale the loss value computed in the forward pass, prior to starting back-propagation. By chain rule back-propagation ensures that all the gradient values are scaled by the same amount…We trained a variety of networks with scaling factors ranging from 8 to 32K”) At the time of filing, one of ordinary skill in the art would have motivation to combine Bijalwan, Shen, Xu, Li, and Micikevicius by taking the electronic device for determining the sensitivity of layers in a model and applying quantization to layers in the model that are insensitive to quantization relative to a threshold during weight update operations, taught jointly by Bijalwan, Shen, Xu, and Li, and adding scaling a gradient calculated during model training, as taught by Micikevicius, as Micikevicius teaches (Micikevicius Pg. 3) “Scaling up the gradients will shift them to occupy more of the representable range and preserve values that are otherwise lost to zeros”. Such a combination would be obvious. Claim 10 is rejected under 35 U.S.C. 103 as being unpatentable over Bijalwan in view of Shen, further in view of Xu, further in view of Li, further in view of Alistarh. Regarding claim 10, Bijalwan, Shen, Xu, and Li jointly teach The electronic device of claim 3, Alistarh teaches the following further limitation that neither Bijalwan, nor Shen, nor Xu, nor Li teaches: wherein the processor is further configured to determine the mean value using "k" largest gradients of the gradients calculated in each of the plurality of nodes, or by applying a genetic algorithm to the gradients, ((Alistarh Pg. 3) “Strom [23], Dryden et al. [8] and Aji and Heafield [2] considered sparsifying the gradient updates by only applying the top K components, taken at at every node, in every iteration, for K corresponding to < 1% of the dimension, and accumulating the error”, Alistarh Pg. 5, Algorithm 1 shows that the TopK gradients are averaged) PNG media_image2.png 236 672 media_image2.png Greyscale where k is an integer ((Alistarh Pg. 6) “To further illustrate necessity, consider a dummy instance with two nodes, dimension 2, and K = 1”, 1 is an integer) At the time of filing, one of ordinary skill in the art would have motivation to combine Bijalwan, Shen, Xu, Li, and Alistarh by taking the electronic device of claim 3, taught jointly by Bijalwan, Shen, Xu, and Li, and adding determining the average gradient using the k largest gradients, taught by Alistarh, as doing so increases the efficiency of the distributed system by not utilizing network bandwidth for transmission of gradients of such small size that they provide only marginal additions to accuracy when considered. Such a combination would be obvious. Claims 12, 17, 19, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Bijalwan, in view of Shen, further in view of Li. Regarding claim 12, Bijalwan teaches An operating method, comprising: generating, based on a determination of sensitivity of layers in a model to be trained, sensitivity results; ((Bijalwan [0055]) At block 408, the sensitivity evaluation module 114 normalizes the features sensitivity values and the weight sensitivity values of the plurality of layers independently and combines the sensitivity values into a union sensitivity list…The first row of the union sensitivity list illustrates that a weight sensitivity value corresponding to layer 1 indicated by "Layer_l_w" is 0.77) Shen teaches the following further limitation more explicitly than Bijalwan: and training the model by applying quantization to a layer of the layers with a low sensitivity of the sensitivity results lower than a predetermined threshold, ((Shen [0028]) “when the value of the objective function corresponding to the first layer L1 is greater than the threshold, this indicates that the loss is small, and the processing unit 120 decides to quantize the first layer L1 with the second precision”, a small loss corresponds to a low sensitivity) and wherein the applying of the quantization comprises: processing the layer with the low sensitivity lower than the predetermined threshold with a first precision by quantizing the layer … ((Shen [0028]) “when the value of the objective function corresponding to the first layer L1 is greater than the threshold, this indicates that the loss is small, and the processing unit 120 decides to quantize the first layer L1 with the second precision”, a small loss corresponds to a low sensitivity) and processing, without quantization, a layer with a high sensitivity of the sensitivity results higher than or equal to the predetermined threshold with a second precision, higher than the first precision, … ((Shen [0028]) “when the values of the objective function corresponding to the second layer L2 and the third layer L3 is not greater than the threshold, this indicates that the loss is large, and the processing unit 120…does not quantize the second layer L2 and the third layer L3 (that is, the second layer L2 and the third layer L3 remain at the first precision)”, a large loss for a layer corresponds to a high sensitivity for the layer) At the time of filing, one of ordinary skill in the art would have motivation to combine Bijalwan and Shen by taking the method for determining the sensitivity of layers in a model taught by Bijalwan and applying a threshold to layers based on sensitivity to determine whether to apply quantization, taught by Shen, as doing so imparts the benefit of optimizing the memory efficiency and speed of the model, increased by quantizing some layers, with respect to the accuracy of the model, increased by refraining from quantizing some layers, by quantizing the layers that are not sensitive enough to significantly decrease accuracy. Such a combination would be obvious. Li teaches the following further limitations that Shen does not teach and that Bijalwan does not teach explicitly: wherein the training includes at least one of an operation of backward propagation of the layer or an operation of updating a weight of the model dependent on a calculated gradient ((Li Pg. 2) “the standard method for training neural networks is stochastic gradient descent (SGD)”, (Li Pg. 3) “The deterministic rounding SGD maintains quantized weights with updates…where wb denotes the low-precision weights, which are quantized using Qd immediately after applying the gradient descent update”) processing the layer…by quantizing the layer ((Li Pg. 8) “we only quantize the weights in the convolutional layers, but not linear layers, during training”) during the at least one of the operation of backward propagation or the operation of updating the weight; ((Li Pg. 3) “The deterministic rounding SGD maintains quantized weights with updates of the form: [Equation 4], where wb denotes the low-precision weights, which are quantized using Qd immediately after applying the gradient descent update”, Li Pg. 3, Equation 4 shows that the updated weight at the next time step is the result of quantizing the weight at the previous time step after a gradient descent update) and processing, without quantization, a layer…during the at least one of the operation of backward propagation or the operation of updating the weight, ((Li Pg. 8) “we only quantize the weights in the convolutional layers, but not linear layers, during training”, training neural networks includes weight update operations) and wherein the application of the quantization is applied during the operation of updating the weight ((Li Pg. 3) “The deterministic rounding SGD maintains quantized weights with updates of the form: [Equation 4], where wb denotes the low-precision weights, which are quantized using Qd immediately after applying the gradient descent update”, Li Pg. 3, Equation 4 shows that the updated weight at the next time step is the result of quantizing the weight at the previous time step after a gradient descent update) At the time of filing, one of ordinary skill in the art would have motivation to combine Bijalwan, Shen, and Li by taking the method for determining the sensitivity of layers in a model and applying quantization to layers in the model that are insensitive to quantization relative to a threshold, taught jointly by Bijalwan and Shen, and applying the quantization of layers during weight update operations of training, taught by Li, as Li teaches: (Li Pg. 2) “For training quantized NNs from scratch, many authors suggest maintaining a high-precision floating point copy of the weights while feeding quantized weights into backprop [5, 11, 3, 16], which results in good empirical performance. There are limitations in using such methods on low-power devices, however, where floating-point arithmetic is not always available or not desirable”, that is, that directly quantizing the weights during the update step without maintaining any high-precision floating point weights intermediately has advantages in a low-power device environment. Such a combination would be obvious. Regarding claim 17, Bijalwan, Shen, and Li jointly teach The operating method of claim 12, Bijalwan further teaches: the determining of the sensitivity comprises classifying the sensitivity results of the layers into a plurality of levels; ((Bijalwan [0037]) “The grouping module 116 may cluster the plurality of layers within the input neural network model into a plurality of groups based on the union sensitivity list”) and the training of the model comprises training the model by applying quantization to each of the layers with a precision at a level corresponding to each of the plurality of levels ((Bijalwan [0037]) “In one embodiment, the grouping module 116 clusters the plurality of layers into a plurality of groups to quantize each group into a high precision format. In another embodiment, the grouping module 116 clusters the plurality of layers into another plurality of groups to quantize each group into a lower precision format”) At the time of filing, one of ordinary skill in the art would have motivation to combine the electronic device jointly taught by Bijalwan, Shen, and Li for the parent claim of claim 17, claim 12. No new embodiments are introduced, so the reason to combine is the same as for the parent claim. Regarding claim 19, Bijalwan, Shen, and Li jointly teach The operating method of claim 12, Bijalwan further teaches: wherein the model to be trained is pretrained with a precision without quantization ((Bijalwan [0036]) “The sensitivity evaluation module 114 receives data from the data acquisition module 230 and generates a union sensitivity list. In one embodiment, the sensitivity evaluation module 114 generates a base model from the input neural network model by representing the parameters of the input neural network model in high precision format and stores as base model 210”, a model in a high precision format corresponds to a model pretrained with a precision without quantization) At the time of filing, one of ordinary skill in the art would have motivation to combine the electronic device jointly taught by Bijalwan, Shen, and Li for the parent claim of claim 19, claim 12. No new embodiments are introduced, so the reason to combine is the same as for the parent claim. Regarding claim 20, Claim 20 discloses a computer readable medium with instructions to perform the method of claim 12. All other limitations in claim 20 are substantially the same as those in claim 12, therefore the same rationale for rejection applies. Claims 13 and 15 are rejected under 35 U.S.C. 103 as being unpatentable over Bijalwan, in view of Shen, further in view of Li, further in view of Dong. Regarding claim 13, Bijalwan, Shen, and Li jointly teach The operating method of claim 12, Dong teaches the following further limitation that neither Bijalwan, nor Shen, nor Li explicitly teaches: wherein the generating of the sensitivity results comprises periodically determining training sensitivity of the layers for each training of the model, or for each epoch or each of one or more iterations performed during the training of the model ((Dong Pgs. 1-2)) “these searching methods can require a large amount of computational resources, are time-consuming, and, worst of all, the quality of quantization is very sensitive to the initialization of their search parameters and therefore unpredictable. This makes deployment of these methods in online learning scenarios especially challenging, as in these applications a new model is trained every few hours and needs to be quantized for efficient inference. To address these issues, recent work introduced HAWQ [7], a Hessian AWare Quantization framework. The main idea is to assign higher bit-precision to layers that are more sensitive, and lower bit-precision to less sensitive layers”, application of the method of determining training sensitivity of layers to online learning scenarios where a new model is trained and quantized every few hours corresponds to periodically determining training sensitivity of the layers for each training of the model) At the time of filing, one of ordinary skill in the art would have motivation to combine Bijalwan, Shen, Li, and Dong by taking the method for determining the sensitivity of layers in a model and applying quantization to layers in the model that are insensitive to quantization relative to a threshold during weight update operations, taught jointly by Bijalwan, Shen, and Li, and adding quantizing the layers, including determining the sensitivity of the layers to quantization, every time the model is trained, as taught by Dong, as doing so increases the effectiveness of the model by optimizing its accuracy relative to its size for the newly trained parameters. Such a combination would be obvious. Regarding claim 15, Bijalwan, Shen, and Li jointly teach The operating method of claim 12, Dong teaches the following further limitation that neither Bijalwan, nor Shen, nor Li explicitly teaches: wherein the determining of the sensitivity comprises periodically determining training sensitivity of the layers for each training of the model, or for each epoch or each of one or more iterations performed during the training of the model ((Dong Pgs. 1-2)) “these searching methods can require a large amount of computational resources, are time-consuming, and, worst of all, the quality of quantization is very sensitive to the initialization of their search parameters and therefore unpredictable. This makes deployment of these methods in online learning scenarios especially challenging, as in these applications a new model is trained every few hours and needs to be quantized for efficient inference. To address these issues, recent work introduced HAWQ [7], a Hessian AWare Quantization framework. The main idea is to assign higher bit-precision to layers that are more sensitive, and lower bit-precision to less sensitive layers”, application of the method of determining training sensitivity of layers to online learning scenarios where a new model is trained and quantized every few hours corresponds to periodically determining training sensitivity of the layers for each training of the model) At the time of filing, one of ordinary skill in the art would have motivation to combine Bijalwan, Shen, Li, and Dong by taking the method for determining the sensitivity of layers in a model and applying quantization to layers in the model that are insensitive to quantization relative to a threshold during weight update operations, taught jointly by Bijalwan, Shen, and Li, and adding quantizing the layers, including determining the sensitivity of the layers to quantization, every time the model is trained, as taught by Dong, as doing so increases the effectiveness of the model by optimizing its accuracy relative to its size for the newly trained parameters. Such a combination would be obvious. Claims 14 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Bijalwan, in view of Shen, further in view of Li, further in view of Xu. Regarding claim 14, Bijalwan, Shen, and Li jointly teach The operating method of claim 12, Xu teaches the following further limitations that neither Bijalwan, nor Shen, nor Li teaches: wherein the training of the model comprises performing, on the model, distributed training comprising operations of: ((Xu [0025]) “A distributed system can accelerate a training process by distributing the training process across multiple computing systems, which can be referred to as worker nodes. Training data can be split into multiple portions, with each portion to be processed by a worker node. Each worker node can perform the forward and backward propagation operations independently”) performing forward propagation moving from a first layer to a last layer of the model; ((Xu [0021]) “During training of a neural network, a first neural network layer can receive training input data, combine the training input data with the weights (e.g., by multiplying the training input data with the weights and then summing the products) to generate first output data for the neural network layer, and propagate the output data to a second neural network layer, in a forward propagation operation…The forward propagation operations can start at the first neural network layer and end at the highest neural network layer”) performing backward propagation moving from the last layer to the first layer of the model; ((Xu [0023]) “As part of the training process, each neural network layer can then perform a backward propagation process to adjust the set of weights at each neural network layer…The backward propagation operations can start from the highest neural network layer and end at the first neural network layer”) determining a mean value of gradients calculated in each of a plurality of nodes used for the distributed training of the model; ((Xu [0025]) “Each worker node can exchange its set of weight gradients with other worker nodes, and average its set of weight gradients and the sets of weight gradients received from other worker nodes”) and updating a weight of the model based on the mean value ((Xu [0025]) “Each computing node can have the same set of averaged weight gradients, and can then update a set of weights for each neural network layer based on the averaged weight gradients”) At the time of filing, one of ordinary skill in the art would have motivation to combine Bijalwan, Shen, Li, and Xu by taking the method for determining the sensitivity of layers in a model and applying quantization to layers in the model that are insensitive to quantization relative to a threshold during weight update operations, taught jointly by Bijalwan, Shen, and Li, and adding the method of distributed training taught by Xu, as Xu teaches: (Xu [0025]) “A distributed system can accelerate a training process by distributing the training process across multiple computing systems”. Such a combination would be obvious. Regarding claim 18, Bijalwan, Shen, Li, and Xu jointly teach The operating method of claim 14, Xu further teaches: wherein the training of the model comprises training the model by applying quantization [to the layer with the low sensitivity lower than the predetermined threshold] in any one or any combination of the operations the distributed training comprises ((Xu [0065]) “In some embodiments, a DMA controller 1002 at the worker node 120-1 may perform one or more of the compression tasks. For example, the DMA controller 1002 may utilize a gradient compression engine (GCE) to perform sparsity analysis, gradient clipping, quantization, and/or compression”, Bijalwan teaches quantization of a layer with low sensitivity below a threshold) At the time of filing, one of ordinary skill in the art would have motivation to combine the electronic device jointly taught by Bijalwan, Shen, Li, and Xu for the parent claim of claim 18, claim 14. No new embodiments are introduced, so the reason to combine is the same as for the parent claim. Claims 5 and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Bijalwan, in view of Shen, further in view of Li, further in view of Liu. Regarding claim 5, Bijalwan teaches An electronic device comprising: ((Bijalwan [0028]) “In one embodiment, the system 100 comprises…at least one device such as a computing device 104”) a processor; ((Bijalwan [0009]) “The system comprises a memory and a processor that is coupled to the memory”) and a memory configured to store instructions executable by the processor, ((Bijalwan [0009]) “The system comprises a memory and a processor that is coupled to the memory”) wherein the processor is configured to, in response to the instructions being executed by the processor: generate, based on a determination of sensitivity of layers in a model to be trained, sensitivity results; ((Bijalwan [0055]) At block 408, the sensitivity evaluation module 114 normalizes the features sensitivity values and the weight sensitivity values of the plurality of layers independently and combines the sensitivity values into a union sensitivity list…The first row of the union sensitivity list illustrates that a weight sensitivity value corresponding to layer 1 indicated by "Layer_l_w" is 0.77) Shen teaches the following further limitation more explicitly than Bijalwan: train the model by applying quantization to a layer of the layers with a low sensitivity of the sensitivity results lower than a predetermined threshold, ((Shen [0028]) “when the value of the objective function corresponding to the first layer L1 is greater than the threshold, this indicates that the loss is small, and the processing unit 120 decides to quantize the first layer L1 with the second precision”, a small loss corresponds to a low sensitivity) At the time of filing, one of ordinary skill in the art would have motivation to combine Bijalwan and Shen by taking the electronic device for determining the sensitivity of layers in a model taught by Bijalwan and applying a threshold to layers based on sensitivity to determine whether to apply quantization, taught by Shen, as doing so imparts the benefit of optimizing the memory efficiency and speed of the model, increased by quantizing some layers, with respect to the accuracy of the model, increased by refraining from quantizing some layers, by quantizing the layers that are not sensitive enough to significantly decrease accuracy. Such a combination would be obvious. Li teaches the following further limitations that Shen does not teach and that Bijalwan does not teach explicitly: where the training includes at least one of an operation of backward propagation of the layer or an operation of updating a weight of the model dependent on a calculated gradient ((Li Pg. 2) “the standard method for training neural networks is stochastic gradient descent (SGD)”, (Li Pg. 3) “The deterministic rounding SGD maintains quantized weights with updates…where wb denotes the low-precision weights, which are quantized using Qd immediately after applying the gradient descent update”) and wherein the application of the quantization is applied during the operation of updating the weight ((Li Pg. 3) “The deterministic rounding SGD maintains quantized weights with updates of the form: [Equation 4], where wb denotes the low-precision weights, which are quantized using Qd immediately after applying the gradient descent update”, Li Pg. 3, Equation 4 shows that the updated weight at the next time step is the result of quantizing the weight at the previous time step after a gradient descent update) At the time of filing, one of ordinary skill in the art would have motivation to combine Bijalwan, Shen, and Li by taking the electronic device for determining the sensitivity of layers in a model and applying quantization to layers in the model that are insensitive to quantization relative to a threshold, taught jointly by Bijalwan and Shen, and applying the quantization of layers during weight update operations of training, taught by Li, as Li teaches: (Li Pg. 2) “For training quantized NNs from scratch, many authors suggest maintaining a high-precision floating point copy of the weights while feeding quantized weights into backprop [5, 11, 3, 16], which results in good empirical performance. There are limitations in using such methods on low-power devices, however, where floating-point arithmetic is not always available or not desirable”, that is, that directly quantizing the weights during the update step without maintaining any high-precision floating point weights intermediately has advantages in a low-power device environment. Such a combination would be obvious. Liu teaches the following further limitations that neither Bijalwan, nor Shen, nor Li explicitly teaches: wherein. to perform the application of the quantization, the processor is further configured to: generate, based on a determination of channel-wise sensitivity of a tensor ((Liu Pg. 2) “The b-bit linear quantization amounts to approximate real numbers using the following quantization set Q…[.]Q denotes the nearest rounding operator w.r.t. Q…[.]Q can be generalized to higher dimensional tensors by first stretching them to one-dimensional vectors then applying Eq. 3”) used for the model, channel-wise sensitivity results; ((Liu Pg. 3) “For a layer L with d-dimensional input, we adopt a simple criterion, output error, to determine the target channels. Output error is the difference of the output of a channel before and after quantization”, output error of a channel in response to quantization corresponds to a determination of channel-wise sensitivity) process a channel with a low channel-wise sensitivity of the channel-wise sensitivity results lower than a second predetermined threshold with a first precision by applying quantization to the channel; ((Liu Pg. 1) “we propose multipoint quantization for post-training quantization, which can achieve the flexibility similar to mixed precision, but uses only a single precision level. The idea is to approximate a full-precision weight vector by a linear combination of multiple low-bit vectors. This allows us to use a larger number of low-bit vectors to approximate the weights of more important channels, while use less points to approximate the insensitive channels”, (Liu Pg. 3) “If e(w;w-hat;DL) is larger than a predefined threshold ϵ, we apply multipoint quantization to this channel”, applying multipoint quantization, which is higher precision, to a channel with error higher than a predefined threshold, while using normal, lower-precision quantization for insensitive channels, corresponds to applying quantization to a channel with sensitivity lower than a threshold) and process a channel with a high channel-wise sensitivity of the channel-wise sensitivity results higher than or equal to the second predetermined threshold with a second precision, higher than the first precision, without quantization ((Liu Pg. 3) “If e(w;w-hat;DL) is larger than a predefined threshold ϵ, we apply multipoint quantization to this channel”, therefore if the error e is above the threshold multi-point quantization, which approximates full-precision, is applied, which corresponds to processing the channel without quantization) At the time of filing, one of ordinary skill in the art would have motivation to combine Bijalwan, Shen, Li, and Liu by taking the electronic device for determining the sensitivity of layers in a model and applying quantization to layers in the model that are insensitive to quantization relative to a threshold during weight update operations, taught jointly by Bijalwan, Shen, and Li, and adding quantizing tensors, including quantizing channels with low sensitivity relative to a threshold but approximating full precision for channels with high error above a threshold, as taught by Liu, as Liu teaches (Liu Pg. 2) “There are two common configurations for post-training quantization, per-layer quantization and per-channel quantization. Per-layer quantization assigns the same K and B for all the weights in the same layer. Per-channel quantization is more fine-grained, and it uses different K and B for different channels. The latter can achieve higher precision, but it also requires more complicated hardware design…We propose multipoint quantization, which can be implemented with common operands on commodity hardware”. Such a combination would be obvious. Regarding claim 16, Bijalwan, Shen, and Li jointly teach The operating method of claim 12, wherein Liu teaches the following further limitations that neither Bijalwan, nor Shen, nor Li explicitly teaches: the determining of the sensitivity comprises generating, based on a determination of channel-wise sensitivity of a tensor used for the model ((Liu Pg. 2) “The b-bit linear quantization amounts to approximate real numbers using the following quantization set Q…[.]Q denotes the nearest rounding operator w.r.t. Q…[.]Q can be generalized to higher dimensional tensors by first stretching them to one-dimensional vectors then applying Eq. 3”), channel-wise sensitivity results, ((Liu Pg. 3) “For a layer L with d-dimensional input, we adopt a simple criterion, output error, to determine the target channels. Output error is the difference of the output of a channel before and after quantization”, output error of a channel in response to quantization corresponds to a determination of channel-wise sensitivity) and the training of the model comprises training the model by applying quantization to a channel with a low channel-wise sensitivity of the channel-wise sensitivity results lower than a second predetermined threshold ((Liu Pg. 1) “we propose multipoint quantization for post-training quantization, which can achieve the flexibility similar to mixed precision, but uses only a single precision level. The idea is to approximate a full-precision weight vector by a linear combination of multiple low-bit vectors. This allows us to use a larger number of low-bit vectors to approximate the weights of more important channels, while use less points to approximate the insensitive channels”, (Liu Pg. 3) “If e(w;w-hat;DL) is larger than a predefined threshold ϵ, we apply multipoint quantization to this channel”, applying multipoint quantization, which is higher precision, to a channel with error higher than a predefined threshold, while using normal, lower-precision quantization for insensitive channels, corresponds to applying quantization to a channel with sensitivity lower than a threshold) At the time of filing, one of ordinary skill in the art would have motivation to combine Bijalwan, Shen, Li, and Liu by taking the method for determining the sensitivity of layers in a model and applying quantization to layers in the model that are insensitive to quantization relative to a threshold during weight update operations, taught jointly by Bijalwan, Shen, and Li, and adding quantizing tensors, including quantizing channels with low sensitivity relative to a threshold, as taught by Liu, as Liu teaches (Liu Pg. 2) “There are two common configurations for post-training quantization, per-layer quantization and per-channel quantization. Per-layer quantization assigns the same K and B for all the weights in the same layer. Per-channel quantization is more fine-grained, and it uses different K and B for different channels. The latter can achieve higher precision, but it also requires more complicated hardware design…We propose multipoint quantization, which can be implemented with common operands on commodity hardware”. Such a combination would be obvious. Claim 21 is rejected under 35 U.S.C. 103 as being unpatentable over Bijalwan, in view of Shen, further in view of Xu, further in view of Li, further in view of Liu. Regarding claim 21, Bijalwan, Shen, Xu, and Li jointly teach The electronic device of claim 1, Liu teaches the following further limitations that neither Bijalwan, nor Shen, nor Xu, nor Li explicitly teaches: wherein the processor is further configured to: generate, based on a determination of channel-wise sensitivity of a tensor ((Liu Pg. 2) “The b-bit linear quantization amounts to approximate real numbers using the following quantization set Q…[.]Q denotes the nearest rounding operator w.r.t. Q…[.]Q can be generalized to higher dimensional tensors by first stretching them to one-dimensional vectors then applying Eq. 3”) used for the model, channel-wise sensitivity results; ((Liu Pg. 3) “For a layer L with d-dimensional input, we adopt a simple criterion, output error, to determine the target channels. Output error is the difference of the output of a channel before and after quantization”, output error of a channel in response to quantization corresponds to a determination of channel-wise sensitivity) process a channel with a low channel-wise sensitivity of the channel-wise sensitivity results lower than a second predetermined threshold with a first precision by applying quantization to the channel; ((Liu Pg. 1) “we propose multipoint quantization for post-training quantization, which can achieve the flexibility similar to mixed precision, but uses only a single precision level. The idea is to approximate a full-precision weight vector by a linear combination of multiple low-bit vectors. This allows us to use a larger number of low-bit vectors to approximate the weights of more important channels, while use less points to approximate the insensitive channels”, (Liu Pg. 3) “If e(w;w-hat;DL) is larger than a predefined threshold ϵ, we apply multipoint quantization to this channel”, applying multipoint quantization, which is higher precision, to a channel with error higher than a predefined threshold, while using normal, lower-precision quantization for insensitive channels, corresponds to applying quantization to a channel with sensitivity lower than a threshold) and process a channel with a high channel-wise sensitivity of the channel-wise sensitivity results higher than or equal to the second predetermined threshold with a second precision, higher than the first precision, without quantization ((Liu Pg. 3) “If e(w;w-hat;DL) is larger than a predefined threshold ϵ, we apply multipoint quantization to this channel”, therefore if the error e is above the threshold multi-point quantization, which approximates full-precision, is applied, which corresponds to processing the channel without quantization) At the time of filing, one of ordinary skill in the art would have motivation to combine Bijalwan, Shen, Xu, Li, and Liu by taking the electronic device for determining the sensitivity of layers in a model and applying quantization to layers in the model that are insensitive to quantization relative to a threshold during weight update operations, taught jointly by Bijalwan, Shen, Xu, and Li, and adding quantizing tensors, including quantizing channels with low sensitivity relative to a threshold but approximating full precision for channels with high error above a threshold, as taught by Liu, as Liu teaches (Liu Pg. 2) “There are two common configurations for post-training quantization, per-layer quantization and per-channel quantization. Per-layer quantization assigns the same K and B for all the weights in the same layer. Per-channel quantization is more fine-grained, and it uses different K and B for different channels. The latter can achieve higher precision, but it also requires more complicated hardware design…We propose multipoint quantization, which can be implemented with common operands on commodity hardware”. Such a combination would be obvious. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Krishnamoorthi “Quantizing deep convolutional networks for efficient inference: A whitepaper” teaches a variety of quantization techniques for neural networks. Liang “Post Training Mixed-Precision Quantization Based on Key Layers Selection” teaches a method of mixed-precision quantization where layers are ranked based on entropy and the most sensitive layers are chosen to have higher precision. Bleiweiss et al. (U.S. Patent Application Publication No. 2019/0205736) teaches a variety of techniques for accelerating the computation of neural networks, including quantization, compression, and distributed training. Any inquiry concerning this communication or earlier communications from the examiner should be directed to VICTOR A NAULT whose telephone number is (703) 756-5745. The examiner can normally be reached M - F, 12 - 8. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Miranda Huang can be reached at (571) 270-7092. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /V.A.N./Examiner, Art Unit 2124 /MIRANDA M HUANG/Supervisory Patent Examiner, Art Unit 2124
Read full office action

Prosecution Timeline

Show 1 earlier event
Jul 24, 2025
Non-Final Rejection mailed — §103
Oct 15, 2025
Response Filed
Jan 28, 2026
Final Rejection mailed — §103
Mar 26, 2026
Applicant Interview (Telephonic)
Mar 26, 2026
Examiner Interview Summary
Apr 27, 2026
Request for Continued Examination
Apr 29, 2026
Response after Non-Final Action
Jul 30, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12682277
TEMPORAL DRIFT DETECTION
4y 5m to grant Granted Jul 14, 2026
Patent 12651173
KNOWLEDGE DISCOVERY BASED ON INDIRECT INFERENCE OF ASSOCIATION
4y 11m to grant Granted Jun 09, 2026
Patent 12579429
DEEP LEARNING BASED EMAIL CLASSIFICATION
4y 2m to grant Granted Mar 17, 2026
Patent 12566953
AUTOMATED PROCESSING OF FEEDBACK DATA TO IDENTIFY REAL-TIME CHANGES
3y 9m to grant Granted Mar 03, 2026
Patent 12561563
AUTOMATED PROCESSING OF FEEDBACK DATA TO IDENTIFY REAL-TIME CHANGES
3y 10m to grant Granted Feb 24, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
53%
Grant Probability
99%
With Interview (+65.7%)
3y 10m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 17 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month