CTNF 17/653,095 CTNF 100392 DETAILED ACTION 07-03-aia AIA 15-10-aia The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA. This Action is non-final and is in response to the claims filed 04/10/2026. Claims 1-16 are currently pending, of which claims 1-14 and 16 are currently rejected. Claim 15 is allowed. Response to Arguments Applicant’s arguments filed on 04/10/2026 have been fully considered. Objection to Title : Objection to title has been withdrawn necessitated by amendment to the title. 35 U.S.C. 112(b) : 35 U.S.C. 112(b) rejections have been withdrawn necessitated by amendments. 35 U.S.C. 103 : Applicant’s arguments filed on 04/10/2026 have been fully considered, but they are not persuasive. Applicant argues in pages 10-14 that Ovtcharov in view of Xilinx in view of Bokhari in view of Zhang do not teach performing two distinct convolution results in parallel. Applicant specifically argues “Applicant respectfully asserts that such a piecemeal combination cannot reasonably be interpreted as teaching both selective bifurcation of parallel convolution results, where one result is reused on-chip through a feedback operation to the input buffer unit, while the other result is written to external memory.” 07-37-08 Examiner respectfully disagrees. In response to applicant's argument that the references fail to show certain features of the invention, it is noted that the features upon which applicant relies (i.e., selective bifurcation of parallel convolution results) are not recited in the rejected claim(s). Although the claims are interpreted in light of the specification, limitations from the specification are not read into the claims. See In re Van Geuns , 988 F.2d 1181, 26 USPQ2d 1057 (Fed. Cir. 1993). Applicant further argues in pages 14 and 15 that the rejection of claim 1 does not articulate reasoning perform the combination. Specifically, Applicant argues “The rejection of Claim 1 does not articulate how Bokhari's router-level flit buffers would be repurposed to function as the claimed output buffer units with the above result-management semantics, nor does it explain why a person of ordinary skill would have had a reasonable expectation of success in doing so. Instead, the rejection effectively re-labels router input buffers as "output buffer units" and then attributes to those structures claim-required behaviors (result storage, selective feedback, and selective writeback) that are not taught by Bokhari or Zhang.” Examiner respectfully disagrees. Examiner explains in the non-final rejection of claim 1 mailed on 12/22/2025 the motivation to perform each motivation, and how this combination would lead to the input buffers of Bokhari to receive convolution results of Ovtcharov. Regarding the combination using the Network-on-Chip architecture of Bokhari, Examiner explains “Bokhari enhances the model of Ovtcharov in view of Xilinx because "The wormhole flow control scheme results in low-area routers, and it is therefore widely used in most on-chip networks" (Bokhari: Page 471, Fourth paragraph).” See non-final rejection on 12/22/2025, Page 8. Claim Rejections - 35 USC § 101 07-04-01 AIA 07-04 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 8-14, and 16 is rejected under 35 U.S.C. 101 because a computer-readable storage media having instructions stored would normally be considered statutory unless the specification defines "recording medium being a computer-readable recording medium" as including transient media such as signals, carries waves, transmissions, optical waves, transmission media or other media incapable of being touched or perceived absent the non-transitory medium through which they are conveyed. Claims 8-14 and 16 are not limited to non-transitory embodiments. Specifically, in view of the specification (Page 5: First paragraph), the computer-readable storage media is not limited to non-transitory embodiments. Instead, the specification merely repeats claim language, and does not explain how computer-readable storage media is non-transitory. Therefore, the claims are not limited to statutory subject matter, hence Claims 8-14 and 16 is non-statutory. Claim Rejections - 35 USC § 103 07-20-aia AIA The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. 07-21-aia AIA Claim s 1-4, 6-11, and 13-14 are rejected under 35 U.S.C. 103 as being unpatentable over Kalin Ovtcharov in NPL: “Accelerating Deep Convolutional Neural Networks Using Specialized Hardware” (https://www.microsoft.com/en-us/research/wp-content/uploads/2016/02/CNN20Whitepaper.pdf), hereinafter “Ovtcharov”, in view of Xilinx’s NPL: “Accelerating DNNs with Xilinx Alveo Accelerator Cards” (https://docs.amd.com/v/u/en-US/wp504-accel-dnns), hereinafter “Xilinx”, in view of Haseeb Bokhari in NPL: “Network-on-Chip Design” (https://liacs.leidenuniv.nl/~stefanovtp/courses/ES/papers/Ch15_NoC_Design.pdf), hereinafter “Bokhari”, further in view of Chen Zhang in NPL: “Optimizing FPGA-based Accelerator Design for Deep Convolutional Neural Networks” (https://dl.acm.org/doi/pdf/10.1145/2684746.2689060), hereinafter “Zhang” . Regarding Claim 1 , Ovtcharov teaches: A data processing apparatus comprising: an external memory that stores processing target data (Fig. 3, e.g., DRAM Channels (external memory); Page 2, Last paragraph, e.g., input image pixels are stored in a multi-banked input buffer from the DRAM) ; an input buffer unit that stores at least part of the processing target data stored in the external memory (Fig. 3, e.g., shows Multi-Banked input Buffer (input buffer unit) Page 2, Last paragraph, e.g., input image pixels are stored in a multi-banked input buffer from the DRAM) ; [a first PE array] that performs … convolution processing using the processing target data stored in the input buffer unit (Fig. 3, e.g., shows PE arrays receiving data from Multi-Banked Input Buffer (input buffer unit); Page 2, Last paragraph, e.g., inputs from Multi-Banked Input Buffer (input buffer unit) are streamed into multiple PE arrays to perform convolutional operations) ; [a second PE array] that performs … convolution processing using the processing target data stored in the input buffer unit (Fig. 3, e.g., shows PE arrays receiving data from Multi-Banked Input Buffer (input buffer unit); Page 2, Last paragraph, e.g., inputs from Multi-Banked Input Buffer (input buffer unit) are streamed into multiple PE arrays to perform convolutional operations) ; a [Network-on-Chip] that stores, as a first stored result, one of a first result of processing by the [first PE array] or a second result of processing by the [second PE array] (Fig. 3, e.g., shows PE arrays outputting data (first stored result) to Network-on-Chip) ; and [the Network-on-Chip] that stores, as a second stored result, the other of the first result or the second result … (Page 3, Paragraph above Fig. 3, e.g., accumulated results are sent back to input buffer for next round of layer computation; Fig. 3, e.g., shows PE arrays outputting data (second stored result) to Network-on-Chip) , wherein the first stored result stored in the [Network-on-Chip] is stored in the input buffer unit (Fig. 3, shows Network-on-Chip outputting data to Multi-Banked Input Buffer (input buffer unit)) , Ovtcharov does not teach: an M x M data processing unit that performs M x M convolution processing using the processing target data stored in the input buffer unit; an N x N data processing unit that performs N x N convolution processing using the processing target data stored in the input buffer unit; a first output buffer unit that stores, as a first stored result, one of a first result of processing by the M x M data processing unit or a second result of processing by the N x N data processing unit; a second output buffer unit that stores, as a second stored result, the other of the first result or the second result not stored in the first output buffer unit , wherein the first stored result stored in the first output buffer unit is stored in the input buffer unit, the second stored result stored in the second output buffer unit is transferred to the external memory. However, Xilinx teaches: an M x M data processing unit that performs M x M convolution processing (Fig. 2, e.g., 3x3 convolution is performed; Page 4, Top paragraph, e.g., xDNN processing engine has dedicated execution paths for different size convolutions) … an N x N data processing unit that performs N x N convolution processing (Fig. 2, e.g., 1x1 convolution is performed; Page 4, Top paragraph, e.g., xDNN processing engine has dedicated execution paths for different size convolutions) … Therefore, it would have been obvious before the effective filing date of the claimed invention to one of ordinary skill in the art to which said subject matter pertains to combine the different convolution sizes performed by xDNN processing engine as taught by Xilinx with the PE arrays performing convolution operations as taught by Ovtcharov. One would have been motivated to combine these references because both references disclose PE arrays performing convolutional operations for neural networks, and Xilinx enhances the model of Ovtcharov because “The xDNN processing engine has dedicated execution paths for each type of command (download, conv, pooling, element-wise, and upload). This allows for convolution commands to be run in parallel with other commands if the network graph allows it.” (Xilinx: Page 4, First paragraph) Ovtcharov in view of Xilinx do not teach: a first output buffer unit that stores, as a first stored result, one of a first result of processing by the M x M data processing unit or a second result of processing by the N x N data processing unit; a second output buffer unit that stores, as a second stored result, the other of the first result or the second result not stored in the first output buffer unit , wherein the first stored result stored in the first output buffer unit is stored in the input buffer unit, the second stored result stored in the second output buffer unit is transferred to the external memory. However, Bokhari teaches the structure of a Network-on-Chip (NoC) architecture. Specifically, Bokhari teaches how the Wormhole routing architecture is widely used on NoC. Bokhari explains “The wormhole flow control scheme results in low-area routers, and it is therefore widely used in most on-chip networks” (Bokhari: Page 471, Fourth paragraph). Bokhari shows in Fig. 15.5 how the Wormhole router architecture receives inputs through input buffers to then be selected by the XBAR and outputted. Additionally, Ovtcharov teaches using a Network-on-chip, but does not teach the specific architecture of the Network-on-Chip. Therefore, it would have been obvious before the effective filing date of the claimed invention to one of ordinary skill in the art to which said subject matter pertains to combine the Wormhole architecture in a Network-on-Chip including input buffers as taught by Bokhari with the Network-on- Chip as taught by Ovtcharov in view of Xilinx. One would have been motivated to combine these references because both references disclose data transmitting using Network-on-Chip architecture, and Bokhari enhances the model of Ovtcharov in view of Xilinx because “The wormhole flow control scheme results in low-area routers, and it is therefore widely used in most on-chip networks” (Bokhari: Page 471, Fourth paragraph). Hence, each output from PE arrays would arrive to a respective input buffer (first and second output buffers). Ovtcharov in view of Xilinx in view of Bokhari do not teach: the second stored result stored in the second output buffer unit is transferred to the external memory. However, in the same field of endeavor, Zhang teaches writing data down to DRAM when N/T n phases (all phases) of computation are complete. Zhang explains “When N/T n phases of computation and data copying are done, the resulting output feature maps are written down to DRAM” (Zhang: Page 168). Additionally, Ovtcharov shows how the Convolutional Neural Network Accelerator can also Writeback data to DRAM Channels (external memory). Therefore, it would have been obvious before the effective filing date of the claimed invention to one of ordinary skill in the art to which said subject matter pertains to combine the storing of the final output in DRAM (external memory) as taught by Zhang with the Convolutional Neural Network Accelerator as taught by Ovtcharov in view of Xilinx in view of Bokhari. One would have been motivated to combine these references because both references disclose PE arrays used for Convolutional Neural Networks, and Zhang enhances the model of Ovtcharov in view of Xilinx in view of Bokhari by allowing for final results of the convolution operations to be stored in an external memory. Regarding Claim 2 , Ovtcharov in view of Xilinx in view of Bokhari in view of Zhang teach: The data processing apparatus according to claim 1, wherein M and N are integers greater than or equal to 1, and M > N (Xilinx: Fig. 2, e.g., 3x3 (MxM) and 1x1 (NxN) convolution is performed; Page 4, Top paragraph, e.g., xDNN processing engine has dedicated execution paths for different size convolutions) . The motivation to combine provided with respect to claim 1 applies equally to claim 2. Regarding Claim 3 , Ovtcharov in view of Xilinx in view of Bokhari in view of Zhang teach: The data processing apparatus according to claim 2, wherein N = 1 (Fig. 2, e.g., 1x1 convolution is performed; Page 4, Top paragraph, e.g., xDNN processing engine has dedicated execution paths for different size convolutions) . The motivation to combine provided with respect to claim 1 applies equally to claim 3. Regarding Claim 4 , Ovtcharov in view of Xilinx in view of Bokhari in view of Zhang teach: The data processing apparatus according to claim 1, wherein the processing target data is data defined by three or more orthogonal axes (Ovtcharov: Page 2, Last paragraph, e.g., PE arrays perform 3D convolution step (three orthogonal axes)) , and the M x M convolution processing or the N x N convolution processing is performed on a first axis and a second axis in the processing target data (Xilinx: Xilinx: Fig. 2, e.g., 3x3 (MxM) and 1x1 (NxN) convolution is performed (first and second axis)) . The motivation to combine provided with respect to claim 1 applies equally to claim 4. Regarding Claim 6 , Ovtcharov in view of Xilinx in view of Bokhari in view of Zhang teach: The data processing apparatus according to claim 4, wherein when the number of data items belonging to a third axis in the data of the second result of the N x N convolution processing is smaller than the number of data items belonging to the third axis in the data of the first result of the M x M convolution processing, the second result of processing by the N x N data processing unit is stored in the second output buffer unit, and the first result of processing by the M x M data processing unit is stored in the first output buffer unit (Ovtcharov: Fig. 3, e.g., results of convolution operations performed in PE arrays is inputted to Network-on-Chip; Bokhari: Fig. 15.5, e.g., Wormhole router architecture in NoC includes input buffers for each input; Combination would cause for each output from PE arrays to arrive to a respective input buffer (first and second output buffers)) . The motivation to combine provided with respect to claim 1 applies equally to claim 6. Regarding Claim 7 , Ovtcharov in view of Xilinx in view of Bokhari in view of Zhang teach: The data processing apparatus according to claim 1, wherein the N x N convolution processing and the M x M convolution processing are performed as part of image processing using a neural network (Ovtcharov: Page 2, Last paragraph, e.g., CNN accelerator receives input image pixels; Xilinx: Fig. 2, e.g., 3x3 (MxM) and 1x1 (NxN) convolution is performed) . The motivation to combine provided with respect to claim 1 applies equally to claim 7. Regarding Claims 8-11 and 13-14 , they are media claims practiced by the apparatus of claims 1-4 and 6-7. They are rejected for the same reasons as claims 1-4 and 6-7 . Allowable Subject Matter 12-151-08 AIA 07-43 12-51-08 Claim s 5 and 12 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. 12-151-07 AIA 07-97 12-51-07 Claim 15 is allowed. 07-43-02 Claim 16 would be allowable if rewritten to overcome the rejection(s) under 35 U.S.C. 101 rejection set forth in this Office action and to include all of the limitations of the base claim and any intervening claims. 13-03-01 AIA The following is a statement of reasons for the indication of allowable subject matter: Ovtcharov teaches a Top-Level Architecture of the Convolutional Neural Network Accelerator that uses a Multi-Banked Input buffer to store input data of different PE Arrays. See Page 3, Figure 3. Ovtcharov does not teach or suggest storing convolution results on a different output buffer unit based on the size of the third axis of the result data. Instead, Ovtcharov teaches storing convolution results from PE arrays in a Network-on-Chip, and does not show where each individual result is stored inside the Network-on-Chip. Therefore, Ovtcharov does not teach the combination of Claim 15, including the limitations of “when the number of data items belonging to a third axis in the data of the first result of the Mx M convolution processing is smaller than the number of data items belonging to the third axis in the data of the second result of the N x N convolution processing, the first result of processing by the M x M data processing unit is stored in the second output buffer unit, and the second result of processing by the N x N data processing unit is stored in the first output buffer unit.” Kfir et al. (U.S. Patent Application No.: US 20190095776 A1), hereinafter “Kfir” – teaches a block diagram that schematically illustrates a convolutional neural network that uses an input buffer to store input data for convolution operations using processing elements. See Fig. 1 and ¶0024-0027. Kfir does not teach or suggest storing convolution results on a different output buffer unit based on the size of the third axis of the result data. Instead, Kfir teaches storing each convolution result in respective output buffers, and does not disclose how it would store convolution results depending on the size. Therefore, Kfir does not teach the combination of Claim 15, including the limitations of “when the number of data items belonging to a third axis in the data of the first result of the Mx M convolution processing is smaller than the number of data items belonging to the third axis in the data of the second result of the N x N convolution processing, the first result of processing by the M x M data processing unit is stored in the second output buffer unit, and the second result of processing by the N x N data processing unit is stored in the first output buffer unit.” Xu et al. (Patent Application Publication No.: US 12602743 B1), hereinafter “Xu” – teaches a streaming multi-processor that performs convolution operations for neural network training using tensor cores. See Fig. 35 and Column 63 Line 1- Column 64 Line 65. Xu does not teach or suggest storing convolution results on a different output buffer unit based on the size of the third axis of the result data. Instead, Xu teaches using an interconnect network to transfer data from processing cores to a shared memory/L1 cache. Therefore, Xu does not teach the combination of Claim 15, including the limitations of “when the number of data items belonging to a third axis in the data of the first result of the Mx M convolution processing is smaller than the number of data items belonging to the third axis in the data of the second result of the N x N convolution processing, the first result of processing by the M x M data processing unit is stored in the second output buffer unit, and the second result of processing by the N x N data processing unit is stored in the first output buffer unit.” Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to CARLOS H DE LA GARZA whose telephone number is (571)272-0474. The examiner can normally be reached Monday-Friday 9:30AM-6PM. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Andrew Caldwell can be reached at (571) 272-3702. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /C.H.D./ Carlos H. De La GarzaExaminer, Art Unit 2182 /EMILY E LAROCQUE/ Primary Examiner, Art Unit 2182 Application/Control Number: 17/653,095 Page 2 Art Unit: 2182 Application/Control Number: 17/653,095 Page 3 Art Unit: 2182 Application/Control Number: 17/653,095 Page 4 Art Unit: 2182 Application/Control Number: 17/653,095 Page 5 Art Unit: 2182 Application/Control Number: 17/653,095 Page 6 Art Unit: 2182 Application/Control Number: 17/653,095 Page 7 Art Unit: 2182 Application/Control Number: 17/653,095 Page 8 Art Unit: 2182 Application/Control Number: 17/653,095 Page 9 Art Unit: 2182 Application/Control Number: 17/653,095 Page 10 Art Unit: 2182 Application/Control Number: 17/653,095 Page 11 Art Unit: 2182 Application/Control Number: 17/653,095 Page 12 Art Unit: 2182