Prosecution Insights
Last updated: October 02, 2026
Application No. 17/576,914

CREATING AN ACCURATE LATENCY LOOKUP TABLE FOR NPU

Final Rejection §103§112
Filed
Jan 14, 2022
Priority
Nov 18, 2021 — provisional 63/281,068
Examiner
NYE, LOUIS CHRISTOPHER
Art Unit
2141
Tech Center
2100 — Computer Architecture & Software
Assignee
Samsung Electronics Co., Ltd.
OA Round
4 (Final)
29%
Grant Probability
At Risk
5-6
OA Rounds
0m
Est. Remaining
59%
With Interview

Examiner Intelligence

Grants only 29% of cases
29%
Career Allowance Rate
4 granted / 14 resolved
-26.4% vs TC avg
Strong +30% interview lift
Without
With
+30.0%
Interview Lift
resolved cases with interview
Typical timeline
4y 2m
Avg Prosecution
23 currently pending
Career history
37
Total Applications
across all art units

Statute-Specific Performance

§101
26.2%
-13.8% vs TC avg
§103
58.7%
+18.7% vs TC avg
§102
8.4%
-31.6% vs TC avg
§112
6.7%
-33.3% vs TC avg
Black line = Tech Center average estimate • Based on career data from 14 resolved cases

Office Action

§103 §112
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 1-7 rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Claim 1 recites the limitation “executing, by the NPU, an inference operation through only the auxiliary layer, measuring, by the host processing device based on hardware-level timing signals, an auxiliary-layer latency associated with the auxiliary layer;”. It is unclear if the “executing…” is a separate step from the “measuring…” or if the “measuring…” is intended to be part of the “executing…” step, and therefore the limitation renders the scope of claim 1 unclear. Thus, claim 1 fails to particularly point and distinctly claim the subject matter which the inventor regards as the invention. Examiner suggests amending this limitation of claim 1 to “executing, by the NPU, an inference operation through only the auxiliary layer; measuring, by the host processing device based on hardware-level timing signals, an auxiliary-layer latency associated with the auxiliary layer;”. Claims 2-7 incorporate by reference all the limitations of claim 1 and are rejected for similar reasons as above. Claim Rejections - 35 USC § 103 The text of those sections of Title 35, U.S. Code not included in this action can be found in a prior Office action. Claim(s) 8-12 is/are rejected under 35 U.S.C. 103 as being unpatentable over Ponomarev et al. (NPL from IDS: Latency Estimation Tool and Investigation of Neural Networks Inference on Mobile GPU, published Aug. 2021, hereinafter "Ponomarev") in view of Tan et al. (NPL: Efficient Execution of Deep Neural Networks on Mobile Devices with NPU, published May 2021, hereinafter “Tan”), further in view of Hermans et al. (NPL: Training and Analyzing Deep Recurrent Neural Networks, published Dec. 2013, hereinafter "Hermans") and Kim Y. et al. (NPL: CPU-Accelerator Co-Scheduling for CNN Acceleration at the Edge, published Nov. 2020, hereinafter “Kim Y.”). Regarding claim 8, Ponomarev teaches a method to estimate a latency of a layer of a neural network executed on a neural processing unit (NPU) (Ponomarev, Pg. 4, Bullets 4-5), the method comprising: measuring, by the host processing device, a total latency for the inference operation for the selected layer and the auxiliary layer based on hardware-level timing signals (Ponomarev, Page 3 Paragraph 2— “We study inference TensorFlow Lite models on mobile GPU and propose an open-source LETI tool allowing to reconstruct models from parametrization and model latency”, Page 4 Bullet 5 — “we evaluate latency of TFLite models on CPU, GPU, or NPU of Android devices”, and in Page 8 Section 4.1 Paragraph 2 — “After that, we fill all the required layers and compute total latency as the sum of corresponding blocks. The speed of single layer is measured directly with TensorFlow-Benchmark.” — teaches measuring, by the host processing device (on mobile devices), a total latency for the inference operation for the selected layer and auxiliary layer (computes total latency as the sum of corresponding blocks, where the corresponding blocks may be the corresponding layers) based on hardware-level timing signals (latency is evaluated on CPU, GPU, or NPU of devices, and thus measurement of latency is based on hardware-level timing)); modeling, by the host processing device, an overhead latency based on a linear regression model of an input size that input to the selected layer, and an output size that is output from the auxiliary layer, the model being fitted using measured NPU timing data obtained by executing the auxiliary layer on the NPU for a plurality of output data sizes (Ponomarev, Fig. 7 and Page 9 Section 4.2 Paragraph 1 – “The baseline is just to use FLOPs as a latency proxy. There are several ways how to fit it, for example, optimize least square error (linear regression).” & Page 8 Section 4.1 Paragraph 2 — “We use a single layer as a block. After that, we initialize inputs for each block with correct input tensor (for the first one, it is image shape: 224x224x3, for the second it is the shape of the first block’s output). After that, we convert these blocks as standalone TFLite models and deploy them on the device for evaluation. We evaluate each block’s inference time within 300 runs” – teaches modeling, by the host processing device, a latency for the inference operation that is associated with the layers by modeling the latency based on a linear regression of an input data size that is input to the selected layer and an output data size that is output from the auxiliary layer (linear regression model based on FLOPS, thus latency is based on complexity of selected layers), the model fitted using measured NPU timing data (as in Page 4 Bullet 5 latency is evaluated on NPU and in Fig. 7 model is fitted using predicted and measured timing data on devices) obtained by executing the auxiliary layer on the NPU for a plurality of output data sizes (uses single layer as block, evaluates latency of executing each block over 300 runs and thus for a plurality of output data sizes, thus obtaining NPU timing data by executing the auxiliary layer, which is one of the executed blocks, on the NPU for a plurality of output data sizes)); updating a latency lookup table (LUT) stored in memory to include the estimated latency, the LUT being used by a neural architecture search engine to configure NPU execution scheduling with reduced root-mean-square error between the estimated and the measured latencies (Ponomarev, Page 4 Bullet 4 – “we evaluate the latency of TFLite models on the CPU, GPU or NPU of the Android devices”, Page 8 Section 4.1 Paragraph 2 – “We use a single layer as a block. After that, we initialize inputs for each block with the correct input tensor (for the first one, it is image shape 224 x 224 x 3; for the second one, it is the shape of the first block’s output). After that, we convert these blocks as standalone TFLite models and deploy them on the device for evaluation. We evaluate each block’s inference time within 300 runs and put the value into the lookup table”, Fig. 2, and in Page 8 Table 2 – teaches updating a latency lookup table to include the estimated latency (evaluates each block’s inference time and puts value into LUT), the LUT being used by a neural architecture search engine to configure NPU execution scheduling (Fig. 2 shows the LUT being used by a NAS to configure execution scheduling) with reduced root-mean-square error between the estimated and the measured latencies (Table 2 shows the reduced error and relative error between the estimated and measured latencies)). Ponomarev fails to explicitly teach measuring, by the host processing device, an overhead latency for the inference operation; and subtracting, by the host processing device, the overhead latency from the total latency to generate an estimate of the latency of the layer. However, analogous to the field of the claimed invention, Tan teaches: executing, by the NPU, an inference operation over the selected layer and the auxiliary layer (Tan, Section 4.2.2 Paragraph 1 – “Since NPU has its own memory space, all data must be moved from the main memory to NPU before the model can be executed on NPU.” and in Section 4.2.2 Paragraph 5 – “The layer processing time for the 𝑙𝑡ℎ layer can be computed as 𝑡𝑝 𝑙 = (𝑇𝑙𝑎𝑠𝑡 𝑙 −𝑡𝑑 𝑙 ) − (𝑇𝑙𝑎𝑠𝑡 𝑙+1 −𝑡𝑑 𝑙+1).” – teaches executing, by the NPU (model executed on NPU), an inference operation over the selected layer and an auxiliary layer to obtain data associated with a model (performs an inference operation over a layer and an auxiliary layer to obtain data, such as processing time information, associated with a model)); measuring, by the host processing device, an overhead latency (Tan, Section 4.2.2 Paragraph 1 – “Since NPU has its own memory space, all data must be moved from the main memory to NPU before the model can be executed on NPU. Since the NPU processing time is short and a large amount of data is transmitted between the main memory and NPU, the data transmission time and the layer processing time may be at similar level. To better estimate the processing time, we cannot ignore this data transmission time, especially when many layers are run on NPU. However, the current SDK for NPU does not include tools for measuring the data transmission time or the layer processing time of different layers. They can only measure the processing time of running the whole DNN model. To address this problem, we propose a method to compute the layer processing time and the data transmission time, and use them to model the data processing time of running the DNNmodel with a layer combination” – teaches measuring, by the host processing device (NPU can only measure processing time of whole model, thus host processing device computes the layer processing time), an overhead latency for the inference operation (running the DNN model) wherein the overhead latency includes data transfer latency (data transmission time) between a dynamic random access memory of a host processing device and a static random access memory of the NPU (NPU has its own memory space, main memory moves data to NPU, computes data transmission time)); subtracting, by the host processing device, the overhead latency from the total latency to generate an estimate of the latency of the selected layer (Tan, Section 4.2.2, Paragraph 3 – “let 𝑡𝑑 𝑙 denote the data transmission time of moving the input data of the 𝑙𝑡ℎ layer from the main memory to NPU”, Section 4.2.2, Paragraph 4 – “let 𝑇𝑙𝑎𝑠𝑡 𝑙 denote the processing time from the 𝑙𝑡ℎ layer to the last layer of the DNN model and it can be computed as 𝑇𝑙𝑎𝑠𝑡 𝑙 = 𝑡𝑑 𝑙 + 𝑛 𝑖=𝑙 𝑡𝑝 𝑖 + 𝑡𝑟 𝑛..” and in Section 4.2.2 Paragraph 5 – “The layer processing time for the 𝑙𝑡ℎ layer can be computed as 𝑡𝑝 𝑙 = (𝑇𝑙𝑎𝑠𝑡 𝑙 −𝑡𝑑 𝑙 ) − (𝑇𝑙𝑎𝑠𝑡 𝑙+1 −𝑡𝑑 𝑙+1).” – teaches subtracting, by the host processing device, the overhead latency (data transmission time) from the total latency (𝑇𝑙𝑎𝑠𝑡 𝑙 and 𝑇𝑙𝑎𝑠𝑡 𝑙+1), to generate an estimate of the latency of the selected layer (determines latency of the lth layer)); Therefore, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the overhead and layer latency measurements of Tan with the total latency measurement, regression models, and look-up tables of Ponomarev to model the overhead latency and update the LUT for more accurate latency calculation. Doing so would better estimate the processing time of layer by accounting for the data transmission time or layer processing time of different layers (Tan, Section 4.2.2). The combination of Ponomarev and Tan fails to explicitly teach adding, by a host processing device, an auxiliary layer to a selected layer of the neural network. adding, by a host processing device, an auxiliary layer to a selected layer of the neural network (Hermans, Section 2.3, Paragraph 2 — “We add layers one by one and at all times an output layer only exists at the current top layer.” — teaches adding, by a host processing device (as in Section 3), an auxiliary layer to a selected layer of a neural network (adds layer one by one and at all times an output layer exist at the top layer)); Therefore, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the supporting layer of Hermans with the consecutive layer latency measurements of Ponomarev and Tan in order to measure the latency of an auxiliary layer and a selected layer. Doing so would enable measuring performance at each layer (Hermans, Section 3.2). The combination of Ponomarev, Tan, and Hermans fails to explicitly teach wherein the linear regression model includes a first coefficient applied to the input data size, a second coefficient applied to the output data size, and an intercept variable. However, analogous to the field of the claimed invention, Kim Y. teaches: wherein the linear regression model includes a first coefficient applied to the input data size, a second coefficient applied to the output data size, and an intercept variable (Kim Y., Figs. 6-9 and Section III, Subsection B 1) b) “Transfer Latency” Paragraph 1 – “The transfer latency is also linearly proportional to the size of data to be transferred. Thus, the transfer latency can be estimated by the following equation: Eq. (4) where SizeData indicates a total size of the data (in the number of the elements) to be transferred. The αtran and βtran are also determined by the linear regression analysis. The SizeData can be calculated as follows: Eq (5)” – teaches wherein the linear regression model includes a first coefficient to be applied to the input data size (SizeIFMs is the size of the input feature maps multiplied by a coefficient αtran, as in Eq. (4) SizeData is multiplied by αtran and in Eq. (5) SizeData includes SizeIFM), a second coefficient applied to the output data size (SizeOFM is the size of the output feature map multiplied by a coefficient αtran, as in Eq. (4) SizeData is multiplied by αtran and in Eq. (5) SizeData includes SizeOFM), and an intercept variable (in Eq. (4), βtran)), Therefore, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to incorporate the latency modeling using regressions of Kim Y. to the linear regression models fitted on NPU and data sizes of Ponomarev, Tan, and Hermans to create an accurate representation of latency based on a linear regression. Doing so would enable the latency regression model to be extended to other systems where an overlap exists between data transfer, computation, and coherence operations (Kim Y., Section III Subsection B 1) d)) and enables determining the relationships between latency and data size by leveraging the linear regression-based latency model (Kim, Section III Subsection B 2)). Regarding claim 9, the combination of Ponomarev, Tan, Hermans, and Kim Y. teach the method of claim 8, wherein modeling the overhead latency further comprises: determining a first size of data input to the selected layer, and a second size of data output from the auxiliary layer (Ponomarev, Page 8 Section 4.1 Paragraph 2 — “we initialize inputs for each block with correct input tensor (for the first one, it is image shape: 224x224x3, for the second it is the shape of the first block’s output)” – indicates determining an input data size to selected layer and output data size from proceeding layer); determining a linear regression model fitted using measured NPU timing data obtained by executing the auxiliary layer on the NPU for a plurality of output data sizes (Ponomarev, Fig. 7 and Page 9 Section 4.2 Paragraph 1 – “The baseline is just to use FLOPs as a latency proxy. There are several ways how to fit it, for example, optimize least square error (linear regression).” & Page 8 Section 4.1 Paragraph 2 — “We use a single layer as a block. After that, we initialize inputs for each block with correct input tensor (for the first one, it is image shape: 224x224x3, for the second it is the shape of the first block’s output). After that, we convert these blocks as standalone TFLite models and deploy them on the device for evaluation. We evaluate each block’s inference time within 300 runs” – teaches determining a linear regression model fitted using measured NPU timing data (as in Page 4 Bullet 5 latency is evaluated on NPU and in Fig. 7 model is fitted using predicted and measured timing data on devices) obtained by executing the auxiliary layer on the NPU for a plurality of output data sizes (uses single layer as block, evaluates latency of executing each block over 300 runs and thus for a plurality of output data sizes, thus obtaining NPU timing data by executing the auxiliary layer, which is one of the executed blocks, on the NPU for a plurality of output data sizes)) The combination of Ponomarev, Tan, and Hermans fail to explicitly teach determining a first value for a first coefficient, a second value for a second coefficient and a third value for an intercept variable using a linear regression model; and determining the overhead latency based on the input data size, the output data size, the first coefficient, the second coefficient and the third value. However, analogous to the field of the claimed invention, Kim Y. teaches: determining a first value for a first coefficient, a second value for a second coefficient and a third value for an intercept variable using a linear regression model (Kim Y., Figs. 6-9 and Section III, Subsection B Paragraph 1 – “For latency estimation, we use a linear regression-based methodology. In general linear regression, we use a following form of the equation: Eq. (1) where X and Y are explanatory and dependent variables, respectively. With the pairs of X and Y values, we determine α and β values through the linear regression training (line fitting). Based on the above form of the equation, we extend it to estimate accelerator and CPU latencies when processing a single CONV layer.” – teaches determining a first value for a first coefficient, a second value for a second coefficient and a third value for an intercept variable using a linear regression model); and determining the overhead latency based on the first size of data, the second size of data, the first coefficient, the second coefficient and the third value (Kim Y., Figs. 6-9 and Section III, Subsection B 1) b) “Transfer Latency” Paragraph 1 – “The transfer latency is also linearly proportional to the size of data to be transferred. Thus, the transfer latency can be estimated by the following equation: Eq. (4) where SizeData indicates a total size of the data (in the number of the elements) to be transferred. The αtran and βtran are also determined by the linear regression analysis. The SizeData can be calculated as follows: Eq (5)” – teaches determining the overhead latency (transfer latency, which may be aggregated with processing, or computation, latency as in Section III, Subsection B 1) d) “Latency Aggregation”) based on the input data size (in Eq. (5), SizeIFMs is the size of the input feature maps), the output data size (SizeOFM is the size of the output feature map), the coefficients and the third value (αtran and βtran)). Therefore, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to incorporate the latency modeling using regressions of Kim Y. to further modify the overhead latency, data sizes, and regression models fitted on NPU of Ponomarev, Tan, Hermans, and Kim Y. to create an accurate representation of overhead latency based on a linear regression. Doing so would enable the latency regression model to be extended to other systems where an overlap exists between data transfer, computation, and coherence operations (Kim Y., Section III Subsection B 1) d)) and enables determining the relationships between latency and data size by leveraging the linear regression-based latency model (Kim Y., Section III Subsection B 2)). Regarding claim 10, the combination of Ponomarev, Tan, Hermans, and Kim Y. teaches the method of claim 8, wherein the auxiliary layer comprises a convolutional Conv lx1 layer (Ponomarev, Page 4 Bullet 1 — “Firstly, we generate the set of parametrized architectures as in NAS-Benchmark. We verify uniqueness by the same hashing procedure as in Reference [10]. Thus, it is additional proof of the same parametrization and set of models. For NAS-Benchmark configuration at this step, we obtain 423,624 parametrized models/graphs with: max7vertices, max 9 edges with 3 possible layer values (except input and output): [‘conv3x3-bn-relu’, ‘conv1x1-bn-relu’, ‘maxpool3x3’]..” — teaches wherein the auxiliary layer comprises an average pooling layer, a convolutional Conv 1x1 layer, or a Conv 3x3 layer (three possible layer types, thus the added layer would be one of a pooling layer, Conv 1x1 layer, or a Conv 3x3 layer)). Regarding claim 11, the combination of Ponomarev, Tan, Hermans, and Kim Y. teaches the method of claim 8, wherein the neural processing unit comprises a first memory, wherein the host processing device is coupled to the neural processing unit, the host processing device comprising a second memory (Tan, Section 4.2.2 Paragraph 1 – “Since NPU has its own memory space, all data must be moved from the main memory to NPU before the model can be executed on NPU. Since the NPU processing time is short and a large amount of data is transmitted between the main memory and NPU, the data transmission time and the layer processing time may be at similar level.” – teaches wherein the neural processing unit comprises a first memory (NPU has its own memory space), wherein the host processing device is coupled to the neural processing unit (data moves from main memory to NPU, thus host processing device is coupled to NPU) and the host processing device comprises a second memory (main memory is memory of host processing device)), and wherein the overhead latency for the inference operation includes data processing by the host processing device and data transportation between the first memory of the neural processing unit and the second memory of the host processing device to execute the inference operation on the selected layer and the auxiliary layer of the neural network (Tan, Section 4.2.2 Paragraph 1 – “Since NPU has its own memory space, all data must be moved from the main memory to NPU before the model can be executed on NPU. Since the NPU processing time is short and a large amount of data is transmitted between the main memory and NPU, the data transmission time and the layer processing time may be at similar level. To better estimate the processing time, we cannot ignore this data transmission time, especially when many layers are run on NPU. However, the current SDK for NPU does not include tools for measuring the data transmission time or the layer processing time of different layers. They can only measure the processing time of running the whole DNN model. To address this problem, we propose a method to compute the layer processing time and the data transmission time, and use them to model the data processing time of running the DNNmodel with a layer combination”, Section 4.2.2 Paragraph 3 – “let 𝑡𝑑 𝑙 denote the data transmission time of moving the input data of the 𝑙𝑡ℎ layer from the main memory to NPU”, and in Section 4.2.2 Paragraph 5 – “The layer processing time for the 𝑙𝑡ℎ layer can be computed as 𝑡𝑝 𝑙 = (𝑇𝑙𝑎𝑠𝑡 𝑙 −𝑡𝑑 𝑙 ) − (𝑇𝑙𝑎𝑠𝑡 𝑙+1 −𝑡𝑑 𝑙+1).” – teaches wherein the overhead latency for the inference operation includes data processing by the host processing device (layer processing time) and data transportation between the first memory of the neural processing unit and the second memory of the host processing device to execute the inference operation on the selected layer and auxiliary layer of the neural network (computes data transmission time, which is the time to transmit data from main memory to NPU memory, of the inference operation over the consecutive layers 𝑡𝑑 𝑙 and 𝑡𝑑 𝑙+1)). Therefore, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the first and second memories, overhead measurements, and layer latency measurements of Tan to further modify the method of Ponomarev, Tan, Hermans, and Kim Y. to account for the latency of data transportation during the inference operation between the two layers. Doing so would better estimate the processing time of layer by accounting for the data transmission time or layer processing time of different layers (Tan, Section 4.2.2). Regarding claim 12, the combination of Ponomarev, Tan, Hermans, and Kim Y. teaches the method of claim 8, further comprising repeating a predetermined number of times executing the inference operation over the selected layer and the auxiliary layer (Ponomarev, Page 4 Bullet 4-5 — “we evaluate latency of TFLite models on CPU, GPU, or NPU of Android devices — with n >= 300 runs” — here, Ponomarev evaluates the total latency, layer by layer, repeated 300 times & in Page 8 Section 4.1 Paragraph 2 – “We use a single layer as a block. After that, we initialize inputs for each block with the correct input tensor (for the first one, it is image shape 224 x 224 x 3; for the second one, it is the shape of the first block’s output). After that, we convert these blocks as standalone TFLite models and deploy them on the device for evaluation. We evaluate each block’s inference time within 300 runs and put the value into the lookup table” – teaches executing an inference operation over selected layers), measuring the total latency for the inference operation for the selected layer and the auxiliary layer (Ponomarev, Page 8 Table 2 — “Latency, measured by direct method and as a sum of block’s latencies for mobile CPU” — shows that the latency is measured as a sum of each layer’s latency) Ponomarev fails to explicitly teach measuring the overhead latency for the inference operation that is associated with the auxiliary layer. However, analogous to the field of the claimed invention, Tan teaches: measuring the overhead latency for the inference operation that is associated with the auxiliary layer (Tan, Section 4.2.2 Paragraph 1 – “Since NPU has its own memory space, all data must be moved from the main memory to NPU before the model can be executed on NPU. Since the NPU processing time is short and a large amount of data is transmitted between the main memory and NPU, the data transmission time and the layer processing time may be at similar level. To better estimate the processing time, we cannot ignore this data transmission time, especially when many layers are run on NPU. However, the current SDK for NPU does not include tools for measuring the data transmission time or the layer processing time of different layers. They can only measure the processing time of running the whole DNN model. To address this problem, we propose a method to compute the layer processing time and the data transmission time, and use them to model the data processing time of running the DNNmodel with a layer combination”, Section 4.2.2 Paragraph 3 – “let 𝑡𝑑 𝑙 denote the data transmission time of moving the input data of the 𝑙𝑡ℎ layer from the main memory to NPU”, and in Section 4.2.2 Paragraph 5 – “The layer processing time for the 𝑙𝑡ℎ layer can be computed as 𝑡𝑝 𝑙 = (𝑇𝑙𝑎𝑠𝑡 𝑙 −𝑡𝑑 𝑙 ) − (𝑇𝑙𝑎𝑠𝑡 𝑙+1 −𝑡𝑑 𝑙+1).” – teaches measuring an overhead latency for the inference operation associated with the consecutive layer (𝑡𝑑 𝑙+1)). Therefore, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to incorporate the measuring of overhead latency for inference operations associated with layers of Tan to further modify the method of Ponomarev, Tan, Hermans, and Kim Y. in order to measure the overhead latency for an inference operation associated with the auxiliary layer of the model. Doing so would better estimate the processing time of layer by accounting for the data transmission time or layer processing time of different layers (Tan, Section 4.2.2). No Prior Art Rejection None of the prior art of record, alone or combined, fairly teaches or suggests claims 1-7. These claims would be allowable if the rejections under 35 U.S.C. 112(b) are resolved. Allowable Subject Matter Claims 13-19 allowed. None of the prior art of record, alone or combined, fairly teaches or suggests claims 13-19. Response to Arguments Applicant’s arguments, see pp. 7-8 of Remarks, filed 12 June 2026, with respect to the rejection(s) of claim(s) 8 under 35 U.S.C. 103 have been fully considered and are persuasive. Therefore, the rejection has been withdrawn. However, upon further consideration, a new ground(s) of rejection is made over Ponomarev in view of Tan, further in view of Hermans and Kim Y. (NPL: CPU-Accelerator Co-Scheduling for CNN Acceleration at the Edge, published Nov. 2020). Ponomarev teaches the amended limitations of claim 8 regarding “the model being fitted using measured NPU timing data obtained by executing…”. Kim Y. teaches the amended limitations of claim 8 regarding “wherein the linear regression model includes a first coefficient applied to the input data size…”. Applicant argues on pp. 7 of Remarks that Ponomarev fails to disclose the model fitted using measured NPU timing data obtained by executing the auxiliary layer for a plurality of output data sizes. Examiner respectfully disagrees. Ponomarev teaches on Page 4 Bullet 5 that latency is evaluated on NPU. Ponomarev at Page 8 Section 4.1 Paragraph 2 states that layers are used as blocks, and that each block’s inference time is evaluated within 300 runs. Thus every layer’s timing data (including the auxiliary layer’s timing data) is obtained by executing the layer for a plurality of output data sizes. Ponomarev shows the linear regression model fitted on NPU timing data at Fig. 7. Thus, Ponomarev teaches the model fitted on NPU timing data obtained by executing the auxiliary layer for a plurality of output data sizes. Applicant further argues on pp. 8-10 that Kim Y. fails to teach modeling the overhead latency and fitting a regression model to NPU timing data. Kim Y. is not relied upon for teaching or suggesting fitting a regression model to NPU timing data, as Ponomarev at Page 8 Section 4.1 Paragraph 2 and Fig. 7 teaches fitting a regression model to NPU timing data. Ponomarev further teaches at Page 8 Section 4.1 Paragraph 2 determining an output data size for a proceeding layer. Kim Y. at Section III, Subsection B 1) b) “Transfer Latency” teaches modeling a transfer latency based on input and output data sizes. The specification of the claimed invention at [0022] states that the overhead latency is associated with data processing and data transportation. Kim Y. states at Section III Subsection B 1) b) that transfer latency is associated with data transfer/transportation and is proportional to the size of data being transferred. It would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to incorporate the latency modeling using regressions based on data size of Kim Y. to the overhead latency, auxiliary layer output data size, and regression models of Ponomarev, Tan, and Hermans to create an accurate representation of overhead latency based on a linear regression. Doing so would enable the latency regression model to be extended to other systems where an overlap exists between data transfer, computation, and coherence operations (Kim Y., Section III Subsection B 1) d)) and enables determining the relationships between latency and data size by leveraging the linear regression-based latency model (Kim Y., Section III Subsection B 2)). In response to applicant's argument that the examiner's conclusion of obviousness is based upon improper hindsight reasoning, it must be recognized that any judgment on obviousness is in a sense necessarily a reconstruction based upon hindsight reasoning. But so long as it takes into account only knowledge which was within the level of ordinary skill at the time the claimed invention was made, and does not include knowledge gleaned only from the applicant's disclosure, such a reconstruction is proper. See In re McLaughlin, 443 F.2d 1392, 170 USPQ 209 (CCPA 1971). Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to LOUIS C NYE whose telephone number is 571-272-0636. The examiner can normally be reached Monday - Friday 9:00AM - 5:00PM. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, MATT ELL can be reached at 571-270-3264. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /LOUIS CHRISTOPHER NYE/Examiner, Art Unit 2141 /MATTHEW ELL/Supervisory Patent Examiner, Art Unit 2141
Read full office action

Prosecution Timeline

Show 3 earlier events
Aug 12, 2025
Final Rejection mailed — §103, §112
Nov 12, 2025
Response after Non-Final Action
Dec 03, 2025
Request for Continued Examination
Dec 10, 2025
Response after Non-Final Action
Mar 12, 2026
Non-Final Rejection mailed — §103, §112
Jun 12, 2026
Response Filed
Aug 18, 2026
Examiner Interview (Telephonic)
Aug 31, 2026
Final Rejection mailed — §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12725092
SYSTEM AND METHOD FOR DECENTRALIZED FEDERATED LEARNING
5y 2m to grant Granted Sep 01, 2026
Patent 12639577
SYSTEMS AND METHODS FOR SELF SUPERVISED MULTI-VIEW REPRESENTATION LEARNING FOR TIME SERIES
4y 8m to grant Granted May 26, 2026
Patent 12524683
METHOD FOR PREDICTING REMAINING USEFUL LIFE (RUL) OF AERO-ENGINE BASED ON AUTOMATIC DIFFERENTIAL LEARNING DEEP NEURAL NETWORK (ADLDNN)
3y 2m to grant Granted Jan 13, 2026
Study what changed to get past this examiner. Based on 3 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

5-6
Expected OA Rounds
29%
Grant Probability
59%
With Interview (+30.0%)
4y 2m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 14 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month