Prosecution Insights
Last updated: October 04, 2026
Application No. 18/328,631

NEURAL NETWORKS PROCESSING UNITS ACTIVATION SPARSITY REMOVAL

Final Rejection §101§102§103
Filed
Jun 02, 2023
Priority
Dec 10, 2020 — provisional 63/123,784 +1 more
Examiner
TRAN, TAN H
Art Unit
2141
Tech Center
2100 — Computer Architecture & Software
Assignee
Neuronix AI Labs Inc.
OA Round
2 (Final)
61%
Grant Probability
Moderate
3-4
OA Rounds
1m
Est. Remaining
94%
With Interview

Examiner Intelligence

Grants 61% of resolved cases
61%
Career Allowance Rate
195 granted / 320 resolved
+5.9% vs TC avg
Strong +33% interview lift
Without
With
+32.6%
Interview Lift
resolved cases with interview
Typical timeline
3y 6m
Avg Prosecution
46 currently pending
Career history
374
Total Applications
across all art units

Statute-Specific Performance

§101
13.4%
-26.6% vs TC avg
§103
59.8%
+19.8% vs TC avg
§102
16.5%
-23.5% vs TC avg
§112
6.3%
-33.7% vs TC avg
Black line = Tech Center average estimate • Based on career data from 320 resolved cases

Office Action

§101 §102 §103
Notice of Pre-AIA or AIA Status 1. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . DETAILED ACTION 2. This Office Action is sent in response to Applicant’s Communication received on 06/26/2026 for application number 18/328,631. Response to Amendments 3. The Amendment filed 06/26/2026 has been entered. Claims 1 and 5 have been amended. Claims 9-16 have been added. Claims 1-16 remain pending in the application. 4. Applicant’s amendment to claim 5 has been fully considered and is persuasive. The amendment provided to overcome the 112(b) rejection issued in the last office action is sufficient. The 35 U.S.C § 112(b) rejection of claim 5 is respectfully withdrawn. 5. Applicant’s amendment to claim 1 has been fully considered and is persuasive. The amendment provided to overcome the 101 rejection issued in the last office action is sufficient. The 35 U.S.C § 101 rejection of claims 1-8 is respectfully withdrawn. Response to Arguments Applicant argues that the cited references fail to teach the amended independent claim 1 and the new independent claim 9. However, the argument is moot since this is a newly presented limitation, thus changes the scope of the claim. However, newly found references, Chinya and Chinya ‘137, are applied. Claim Rejections - 35 USC § 101 6. 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 9-16 are rejected under 35 U.S.C. 101 because the claimed invention is directed to the abstract idea without significantly more. Step 1, the claims are directed to the statutory categories of a process. Claim 9: Step 2A Prong 1, Claim 9 recites, in part using an activation bit-map to identify non-zero activations (Mental process, a person can examine an activation bitmap and identify which indicators correspond to non-zero activations). Step 2A Prong 2, this judicial exception is not integrated into a practical application. The additional elements: a scalable deep neural network accelerator (sDNA) comprising multiple neural processing units (NPUs) (mere instructions to apply the exception using computer component). implementing a non-zero Activation jump algorithm (mere instructions to apply the exception using computing technology). an activation bit-map fetched from an activation map memory (mere data gathering and recited at a high level of generality, and thus are insignificant extra-solution activity). Step 2B, the claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception, either alone or in combination. The additional elements: a scalable deep neural network accelerator (sDNA) comprising multiple neural processing units (NPUs) (mere instructions to apply the exception using computer component). implementing a non-zero Activation jump algorithm (mere instructions to apply the exception using computing technology). an activation bit-map fetched from an activation map memory (mere data gathering and recited at a high level of generality, and thus are insignificant extra-solution activity). Claims 10-16 provide further limitations to the abstract idea (Mental processes and/or Mathematical concepts) as rejected in claim 9, however, they do not disclose any additional elements that would amount to a practical application or significantly more than an abstract idea (data gathering/insignificant extra-solution activity and/or generic computer component). Claim Rejections - 35 USC § 102 7. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale or otherwise available to the public before the effective filing date of the claimed invention. (a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. 8. Claim 9 is rejected under 35 U.S.C. 102(a)(1) as being anticipated by Chinya et al. (U.S. Patent Application Pub. No. US 20210042617 A1). Claim 9: Chinya teaches a method of activation sparsity removal (i.e. Method 480 may load activations and employ a tunable look-ahead window that skips activations that are zero … where sparse activations that have a zero value within a pre-specified tunable window length may be skipped during a load cycle for processing elements; para. [0046, 0047]) in a scalable deep neural network accelerator (sDNA) (i.e. an AI accelerator 148 that is dedicated to artificial intelligence (AI) and/or neural network (NN) processing … being scalable to operate with different neural network sizes and areas; para. [0057, 0107]) comprising multiple neural processing units (NPUs) (i.e. Process 100 may include dividing the neural network workload 102 and compressing the data of the workload 104 based on processing elements (PEs). For example, in the present example 16 PE0-PE15 are provided; para. [0019, 0057]), the method comprising: implementing a non-zero Activation jump algorithm (i.e. the lookahead examples 702, 704, 706 employ a lookahead technique for loading activations, to employ a tunable look-ahead window that skips activations that are zero within the specified window length … the sparsity activation pointer may be incremented by 1+Look-ahead Length from the current value; para. [0053, 0054]) using an activation bit-map (i.e. the sparsity decoder of a PE may first identify the sparsity bitmap (e.g., Bit 0-Bit 15) to determine which byte positions are non-zero; para. [0049]) fetched from an activation map memory (i.e. The activation data within a PE may include a sparsity bitmap register 814; para. [0054]) to identify non-zero activations (i.e. Illustrated processing block 482 identifies a decode operation 482. Illustrated processing block 484 identifies a lookahead window for a sparsity bitmap decode operation based on a current position in the bitmap. Illustrated processing block 486 determines if any of the sparsity bitmap values from the sparsity bitmap in the lookahead window are associated with a non-zero number; para. [0043, 0049]). Claim Rejections – 35 USC § 103 9. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. 10. Claims 1-4 and 10-12 are rejected under 35 U.S.C. 103 as being unpatentable over Chinya in view of Chinya ‘137 et al. (U.S. Patent Application Pub. No. US 20200228137 A1). Claim 1: Chinya teaches a method of activation sparsity removal (i.e. Method 480 may load activations and employ a tunable look-ahead window that skips activations that are zero … where sparse activations that have a zero value within a pre-specified tunable window length may be skipped during a load cycle for processing elements; para. [0046, 0047]) performed in a scalable deep neural network accelerator (sDNA) (i.e. an AI accelerator 148 that is dedicated to artificial intelligence (AI) and/or neural network (NN) processing … being scalable to operate with different neural network sizes and areas; para. [0057, 0107]) comprising a plurality of neural network processing units (NPUs) (i.e. Process 100 may include dividing the neural network workload 102 and compressing the data of the workload 104 based on processing elements (PEs). For example, in the present example 16 PE0-PE15 are provided; para. [0019, 0057]), each comprising an activation map memory (i.e. FIGS. 8A and 8B illustrate a layout of compressed data and the reconstruction of sparsity bitmaps within an individual PE. The embodiments of FIGS. 8A-8B may be implemented within the PE 452 (FIG. 5) to be part of the PE 452; para. [0054]) configured to store activation map words (i.e. The reason for the above is that when the sparsity decoder decodes the byte stream, the sparsity decoder of a PE may first identify the sparsity bitmap (e.g., Bit 0-Bit 15) to determine which byte positions are non-zero. The bytes may be broadcast to a group of PEs, so the decoder must step through the relevant portions of sparsity bitmap that are associated with the PE, one byte at a time; para. [0049]) that indicate positions of non-zero activation values in corresponding activation tensors (i.e. each byte of activation data (e.g., intermediate feature maps generated as the outputs from intermediate hidden layers in a DNN) and a corresponding bit in the sparsity bitmap; para. [0040]) and a multiplier-accumulator (MAC) unit such that the sDNA includes a plurality of activation map memories and a plurality of MAC units (i.e. a multiply and accumulate or a MAC may be a computation element of the PE 452; para. [0019, 0041, 0054]), the method comprising performing, by the sDNA (i.e. any aspect of the embodiments described herein may be implemented in the processors and/or accelerators dedicated to AI and/or NN processing such as AI accelerator 148; para. [0057]): implementing a non-zero Activation jump algorithm (i.e. In FIG. 7A, the lookahead example 702 with a look ahead window of 1 may identify the immediate byte as well as the following byte in the sparsity bitmap to check if the following byte is 0. If a 0 is detected, then a skip signal may be triggered to skip the load; para. [0050]) that uses the activation map words (i.e. Illustrated processing block 482 identifies a decode operation 482. Illustrated processing block 484 identifies a lookahead window for a sparsity bitmap decode operation based on a current position in the bitmap. Illustrated processing block 486 determines if any of the sparsity bitmap values from the sparsity bitmap in the lookahead window are associated with a non-zero number; para. [0043, 0049]) to generate of non-zero activation values (i.e. Based on the activation skip signal, the sparsity activation pointer, which is the activation sparsity bitmap write pointer, may be incremented. When the activate skip signal is equal to 0 (a non-zero value detected), the MUX 816 may increment the value of the sparsity activation pointer by 1; para. [0054]) in a compressed activation memory (i.e. The activation data and the write enable may be together used to write the sparsity bitmap and the compressed data in the activation register file; para. [0041, 0054]) of an NPU of the plurality of NPUs (i.e. FIGS. 8A and 8B illustrate a layout of compressed data and the reconstruction of sparsity bitmaps within an individual PE; para. [0054]), and the compressed activation memory and the plurality of MAC units, the routing multiplexer network selectively routing the non-zero activation values to the plurality of MAC units (i.e. FIGS. 7A-7C illustrate the above. For example, in FIGS. 7A-7C, 16B activations may be being broadcast into a group of PEs … a MAC may be a computation element of the PE 452; para. [0041, 0048]). Chinya does not explicitly teach to generate addresses of non-zero activation values, and to control a routing multiplexer network coupled between the compressed activation memory and the plurality of MAC units. However, Chinya ‘137 teaches implementing a non-zero Activation jump algorithm (i.e. the first sparse decoder 114 translates the schedule-dependent byte select signals (e.g., Byte_Sel[0]-Byte_Sel[N]) to sparsity-dependent byte select signals (e.g., Sparse_Byte_Sel[0]-Sparse_Byte_Sel[N]) based on the sparsity bitmap (SB). The example first sparse decoder 114 can then apply the sparsity-dependent byte select signals (e.g., Sparse_Byte_Sel[0]-Sparse_Byte_Sel[N]) to one or more ZVC data vectors. Based on the sparsity bitmap, the example first sparse decoder 114 generates write enable signals (e.g., write_en[0]-write_en[N]) to enable each PE with selected data from the ZVC data vector; para. [0043, 0058]) that uses the activation map words (i.e. FIG. 2, the data controller 210 generates the sparse byte select signals (e.g., Sparse_Byte_Sel[0]-Sparse_Byte_Sel[7]) based on the byte select signals (Byte_Sel[0]-Byte_Sel[7]) and/or the sparsity bitmap (e.g., the sparsity bitmap 204) … the sparse byte signals (e.g. Sparse_Byte_Sel[0]-Sparse_Byte_Sel[7]) are sparsity-aware byte select signals to control the first multiplexer array 116; para. [0058, 0059]) to generate addresses of non-zero activation values (i.e. the data controller 210 also sums the values of the bits between (a) the position in the sparsity bitmap (e.g., the sparsity bitmap 204) corresponding to the value of the byte select signal for a given PE (e.g., Byte_Sel[0] for the first PE 126) and (b) the LSB of the sparsity bitmap (e.g., the sparsity bitmap 204) and sets the sparse byte select signal for the given PE (e.g., Sparse_Byte_Sel[0] for the first PE 126) equal to the summed value minus one; para. [0070, 0071]) in a compressed activation memory (i.e. the first input buffer 112 includes an example header 202, an example sparsity bitmap 204, and an example ZVC data vector 206 … compressed data includes a sparsity bitmap (e.g., the sparsity bitmap 204) and a ZVC data vector (e.g., the ZVC data vector 206); para. [0051, 0053]) of an NPU of the plurality of NPUs (i.e. FIG. 3 is a block diagram of an example processing element (PE) 300 constructed in accordance with the teachings of this disclosure. For example, the PE 300 is an example implementation of the first PE 126; para. [0076, 0078, 0081]), and to control a routing multiplexer network (i.e. the first multiplexer array 116 is driven by n sparsity-dependent byte-select signals … sparsity-aware byte select signals to control the first multiplexer array 116, which apply to a first portion of data from the first input buffer 112 and route the data to the designated PEs; para. [0045, 0059]) coupled between the compressed activation memory and the plurality of MAC units (i.e. FIG. 2 is a block diagram showing an example implementation of the first schedule-aware sparse distribution controller 102 a of FIG. 1. The example first schedule-aware sparse distribution controller 102 a includes the example first input buffer 112, the example first sparse decoder 114, the example first multiplexer array 116, and the example first PE column 118 … FIG. 3, the activation transmission gate 302 is coupled to the output of a multiplexer (e.g., the first multiplexer 120) in a multiplexer array (e.g., the first multiplexer array 116) and a write controller (e.g., the write controller 212). The activation register 304 is coupled to the output of the activation transmission gate 302 and to the multiplier 318. The activation sparsity bitmap register 306 is coupled to the write controller (e.g., the write controller 212); para. [0051, 0079]), the routing multiplexer network selectively routing (i.e. FIG. 7, the first multiplexer array 116 routes 16 bytes of compressed data output from the first input buffer 112 to the designated PEs; para. [0059, 0096]) the non-zero activation values to the plurality of MAC units (i.e. Based on the sparsity bitmap, the example first sparse decoder 114 generates write enable signals (e.g., write_en[0]-write_en[N]) to enable each PE with selected data from the ZVC data vector; para. [0043, 0063, 0071]). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to modify the invention of Chinya to include the feature of Chinya ‘137. One would have been motivated to make this modification because it improves accelerator utilization and throughput. Claim 2: Chinya and Chinya ‘137 teach the method of claim 1. Chinya does not explicitly teach generating different combinations of vector multiplication tensors for machine learning models or algorithms. However, Chinya ‘137 further teaches generating different combinations (i.e. the shape of the tensor (e.g., two by two by three, etc.) to be processed and the volume processed by each PE according to a schedule … The example InSAD system 100 combines both flexible tensor distribution and sparse data compression by (1) decoding ZVC data vectors with software programed byte select signals (e.g., Byte_Sel[0]-Byte_Sel[N]) to distribute non-zero data to respective PE arrays, (2) reconstructing the sparsity bitmap at each PE on the fly for different tensor shapes, (3) eliminating one or more storage requirements for uncompressed data across on-chip memory hierarchy, and (4) serving different tensor shapes (e.g., one or more multi-dimension array dimensions) for each PE; para. [0041, 0050, 30, 109, 110]) of vector multiplication tensors (i.e. Input images, sometimes referred to as input activations or input feature maps, are also loaded into PE arrays, where PEs execute multiply accumulate (MAC) operations via one or more input channels (Ic) and generate output activations. One or more sets of weight tensors (Oc) are often used for a given set of input activations to produce an output tensor volume; para. [0026]) for machine learning models or algorithms (i.e. Machine learning accelerators (e.g., those utilizing DNN engines, CNN engines, etc.) handle a large amount of tensor data (e.g., data stored in multi-dimensional data structures) for performing inference tasks; para. [0025]). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to modify the invention of Chinya to include the feature of Chinya ‘137. One would have been motivated to make this modification because it improves accelerator utilization and throughput. Claim 3: Chinya and Chinya ‘137 teach the method of claim 1. Chinya does not explicitly teach supporting at least one of multiple different parallel modes including at least one of: a multiple points (pixels) parallel scheme, a lines parallel scheme, a multiple input channels parallel scheme, or a multiple output channels parallel scheme. However, Chinya ‘137 further teaches supporting at least one of multiple different parallel modes (i.e. Common DNN accelerators are built from a spatial array of PEs … FIG. 9 illustrates a use case of the InSAD system 100 for five different schedules with different byte select signals … FIGS. 9 and 10 illustrate five example distribution schemes including (1) unicast data of different tensor shapes (scheme 1-3), (2) broadcast data (scheme 4), and (3) multicast data (scheme 5); para. [0026, 0107, 0110]) including at least one of: a multiple points (pixels) parallel scheme, a lines parallel scheme, a multiple input channels parallel scheme, or a multiple output channels parallel scheme (i.e. The example of FIG. 9 illustrates the dense data case (no zeros in the tensor volume), where different shading patterns show the different points in this tensor. For example, communication schemes 1-3 illustrate unicast cases in which each PE has different data points; para. [0026, 0107, 0108]). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to modify the invention of Chinya to include the feature of Chinya ‘137. One would have been motivated to make this modification because it improves accelerator utilization and throughput. Claim 4: Chinya and Chinya ‘137 teach the method of claim 1. Chinya does not explicitly teach implementing a sequential execution NPU, a concurrent execution NPU, or a combination of a sequential execution NPU and a concurrent execution NPU to implement the vector multiplication. However, Chinya ‘137 further teaches implementing a sequential execution NPU, a concurrent execution NPU, or a combination of a sequential execution NPU and a concurrent execution NPU to implement the vector multiplication (i.e. Common DNN accelerators are built from a spatial array of PEs and local storage such as register files (RF) and static random access memory (SRAM) banks. For inference tasks, the weights or filters are pre-trained and layer-specific. As such, the weights and/or filters need to be loaded to PE arrays from the storage (e.g. dynamic random access memory (DRAM) and/or SRAM buffers). Input images, sometimes referred to as input activations or input feature maps, are also loaded into PE arrays, where PEs execute multiply accumulate (MAC) operations via one or more input channels (Ic) and generate output activation; para. [0026]). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to modify the invention of Chinya to include the feature of Chinya ‘137. One would have been motivated to make this modification because it improves accelerator utilization and throughput. Claims 10-12 are similar in scope to Claims 2-4 and are rejected under a similar rationale. 11. Claims 5, 6, 13, and 14 are rejected under 35 U.S.C. 103 as being unpatentable over Chinya in view of Chinya ‘137, and further in view of Labbe et al. (U.S. Patent Application Pub. No. US 20200051309 A1). Claim 5: Chinya and Chinya ‘137 teach the method of claim 4. Chinya does not explicitly teach wherein: implementing a sequential execution comprises storing back (feedback) an output of each neural network layer to a current layer of an Activation Memory Matrix (AMM); and implementing a concurrent execution comprises allocating different hardware resources to different DNN layers to process the DNN layers in parallel (concurrently). However, Labbe teaches wherein: implementing a sequential execution comprises storing back (feedback) an output of each neural network layer to a current layer of an Activation Memory Matrix (AMM); and implementing a concurrent execution comprises allocating different hardware resources to different DNN layers to process the DNN layers in parallel (concurrently) (i.e. the single layer hardware neural network block 2500 is programmed to wrap back upon itself as the block processes each layer … Although the multiple layer hardware neural network block 2600 consumes a greater die area than the single layer hardware neural network block 2500, the multiple layer block can pipeline layer processing, enabling improved performance; para. [0231, 0233]). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to modify the combination of Chinya and Chinya ‘137 to include the feature of Labbe. One would have been motivated to make this modification because it enables sequential reuse of compute and pipeline design. Claim 6: Chinya, Chinya ‘137, and Labbe teach the method of claim 5. Chinya does not explicitly teach implementing a sequential execution comprises reusing hardware resources to calculate different layers of a same neural network; and implementing a concurrent execution comprises providing results of each DNN layer to another hardware logic that executes a next DNN layer. However, Labbe further teaches implementing a sequential execution comprises reusing hardware resources to calculate different layers of a same neural network (i.e. the single layer hardware neural network block 2500 is programmed to wrap back upon itself as the block processes each layer. Each neuron is programmed with a specific set of weights and feeds inputs either from the initial input or the prior layer. As each layer is processed, the neural network block consumes consume a portion of a weight buffer, which can be stored in the weights cache 2510, and applies the neuron calculation and activation function via the neural network & activation operations unit 2504. When a layer is completed, the output buffer becomes the input buffer for the next layer; para. [0231]); and implementing a concurrent execution comprises providing results of each DNN layer to another hardware logic that executes a next DNN layer (i.e. FIG. 26 illustrates a multiple layer hardware neural network block 2600, according to an embodiment. In one embodiment, the multiple layer hardware neural network block 2600 includes multiple neural network blocks 2520 illustrated in FIG. 25. For example, the illustrated multiple layer hardware neural network block 2600 includes three neural network blocks 2520A-2520C, although a different number of blocks can be included; para. [0232, 0233]). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to modify the combination of Chinya and Chinya ‘137 to include the feature of Labbe. One would have been motivated to make this modification because it enables sequential reuse of compute and pipeline design. Claims 13-14 are similar in scope to Claims 5-6 and are rejected under a similar rationale. 12. Claims 7 and 8 are rejected under 35 U.S.C. 103 as being unpatentable over Chinya in view of Chinya ‘137, and further in view of Mills (U.S. Patent Application Pub. No. US 20190340498 A1). Claim 7: Chinya and Chinya ‘137 teach the method of claim 1. Chinya does not explicitly teach different size convolution operations. However, Mills teaches comprising supporting different size convolution operations (i.e. per-sub-channel kernels of different sizes can be generated at each neural engine 314. For example, 5×5 shaped kernel data 326 may be sub-sampled (e.g., at kernel extract circuit 432) into sub-kernels of sizes 3×3, 2×3, 3×2 and 2×2, and provided as corresponding kernel coefficients 422 to MAC 404 for sub-channel convolutions with corresponding sub-channels of portion 408 of input data; para. [0095]). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to modify the combination of Chinya and Chinya ‘137 to include the feature of Mills. One would have been motivated to make this modification because it improves accelerator applicability and performance by avoiding software fallbacks and reducing memory overhead. Claim 8: Chinya, Chinya ‘137, and Mills teach the method of claim 7. Chinya does not explicitly teach wherein supporting different size convolution operations comprises supporting two different n*n convolution operations, and wherein n in a first of the convolution operations has a first value that is different than a second value of n in a second of the convolution operations. However, Mills further teaches wherein supporting different size convolution operations comprises supporting two different n*n convolution operations, and wherein n in a first of the convolution operations has a first value that is different than a second value of n in a second of the convolution operations (i.e. per-sub-channel kernels of different sizes can be generated at each neural engine 314. For example, 5×5 shaped kernel data 326 may be sub-sampled (e.g., at kernel extract circuit 432) into sub-kernels of sizes 3×3, 2×3, 3×2 and 2×2, and provided as corresponding kernel coefficients 422 to MAC 404 for sub-channel convolutions with corresponding sub-channels of portion 408 of input data; para. [0095, 0104]). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to modify the combination of Chinya and Chinya ‘137 to include the feature of Mills. One would have been motivated to make this modification because it improves accelerator applicability and performance by avoiding software fallbacks and reducing memory overhead. 13. Claims 15 and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Chinya in view of Chinya ‘137, and further in view of Mills (U.S. Patent Application Pub. No. US 20190340498 A1). Claim 15: Chinya teaches the method of claim 9. Chinya does not explicitly teach different size convolution operations. However, Mills teaches comprising supporting different size convolution operations (i.e. per-sub-channel kernels of different sizes can be generated at each neural engine 314. For example, 5×5 shaped kernel data 326 may be sub-sampled (e.g., at kernel extract circuit 432) into sub-kernels of sizes 3×3, 2×3, 3×2 and 2×2, and provided as corresponding kernel coefficients 422 to MAC 404 for sub-channel convolutions with corresponding sub-channels of portion 408 of input data; para. [0095]). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to modify the invention of Chinya to include the feature of Mills. One would have been motivated to make this modification because it improves accelerator applicability and performance by avoiding software fallbacks and reducing memory overhead. Claim 16: Chinya and Mills teach the method of claim 15. Chinya does not explicitly teach wherein supporting different size convolution operations comprises supporting two different n*n convolution operations, and wherein n in a first of the convolution operations has a first value that is different than a second value of n in a second of the convolution operations. However, Mills further teaches wherein supporting different size convolution operations comprises supporting two different n*n convolution operations, and wherein n in a first of the convolution operations has a first value that is different than a second value of n in a second of the convolution operations (i.e. per-sub-channel kernels of different sizes can be generated at each neural engine 314. For example, 5×5 shaped kernel data 326 may be sub-sampled (e.g., at kernel extract circuit 432) into sub-kernels of sizes 3×3, 2×3, 3×2 and 2×2, and provided as corresponding kernel coefficients 422 to MAC 404 for sub-channel convolutions with corresponding sub-channels of portion 408 of input data; para. [0095, 0104]). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to modify the invention of Chinya to include the feature of Mills. One would have been motivated to make this modification because it improves accelerator applicability and performance by avoiding software fallbacks and reducing memory overhead. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant’s disclosure. Park et al. (Pub. No. US 20220066960 A1), Although illustrated are an 8×8 input data map, a 3×3 filter, an adder tree-based 16-operand MAC, examples are not limited thereto. The foregoing descriptions are also applicable to convolution operations based on input data maps, filters, and MACs of various structures and sizes. Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any extension fee pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the date of this final action. It is noted that any citation to specific pages, columns, lines, or figures in the prior art references and any interpretation of the references should not be considered to be limiting in any way. A reference is relevant for all it contains and may be relied upon for all that it would have reasonably suggested to one having ordinary skill in the art. In re Heck, 699 F.2d 1331, 1332-33, 216 U.S.P.Q. 1038, 1039 (Fed. Cir. 1983) (quoting In re Lemelson, 397 F.2d 1006, 1009, 158 U.S.P.Q. 275, 277 (C.C.P.A. 1968)). Any inquiry concerning this communication or earlier communications from the examiner should be directed to TAN TRAN whose telephone number is (303)297-4266. The examiner can normally be reached on Monday - Thursday - 8:00 am - 5:00 pm MT. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Matt Ell can be reached on 571-270-3264. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /TAN H TRAN/Primary Examiner, Art Unit 2141
Read full office action

Prosecution Timeline

Jun 02, 2023
Application Filed
Feb 26, 2026
Non-Final Rejection mailed — §101, §102, §103
Apr 28, 2026
Interview Requested
May 27, 2026
Examiner Interview Summary
May 27, 2026
Applicant Interview (Telephonic)
Jun 26, 2026
Response Filed
Aug 13, 2026
Final Rejection mailed — §101, §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12748960
Analog Hardware Realization of Neural Networks
5y 7m to grant Granted Sep 29, 2026
Patent 12718079
Systems and Methods for Generating Libraries for Hardware Realization of Neural Networks
5y 5m to grant Granted Aug 25, 2026
Patent 12718088
DESIGNING LADDER AND LAGUERRE ORTHOGONAL RECURRENT NEURAL NETWORK ARCHITECTURES INSPIRED BY DISCRETE-TIME DYNAMICAL SYSTEMS
4y 8m to grant Granted Aug 25, 2026
Patent 12688413
METHODS FOR RELIABLE OVER-THE-AIR COMPUTATION WITH PULSES FOR DISTRIBUTED LEARNING AND WITH FEDERATED EDGE LEARNING WITHOUT CHANNEL STATE INFORMATION
4y 1m to grant Granted Jul 21, 2026
Patent 12682274
MODEL INTEGRATION APPARATUS, MODEL INTEGRATION METHOD, COMPUTER-READABLE STORAGE MEDIUM STORING A MODEL INTEGRATION PROGRAM, INFERENCE SYSTEM, INSPECTION SYSTEM, AND CONTROL SYSTEM
5y 0m to grant Granted Jul 14, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
61%
Grant Probability
94%
With Interview (+32.6%)
3y 6m (~1m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 320 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month