Prosecution Insights
Last updated: August 17, 2026
Application No. 18/050,699

SEMICONDUCTOR DEVICE, METHOD OF OPERATING SEMICONDUCTOR DEVICE, AND SEMICONDUCTOR SYSTEM

Non-Final OA §103
Filed
Oct 28, 2022
Priority
Nov 15, 2021 — RE 10-2021-0157109 +1 more
Examiner
MARU, MATIYAS T
Art Unit
2148
Tech Center
2100 — Computer Architecture & Software
Assignee
Samsung Electronics Co., Ltd.
OA Round
3 (Non-Final)
65%
Grant Probability
Moderate
3-4
OA Rounds
5m
Est. Remaining
70%
With Interview

Examiner Intelligence

Grants 65% of resolved cases
65%
Career Allowance Rate
33 granted / 51 resolved
+9.7% vs TC avg
Minimal +5% lift
Without
With
+4.8%
Interview Lift
resolved cases with interview
Typical timeline
4y 2m
Avg Prosecution
28 currently pending
Career history
81
Total Applications
across all art units

Statute-Specific Performance

§101
35.3%
-4.7% vs TC avg
§103
52.6%
+12.6% vs TC avg
§102
1.9%
-38.1% vs TC avg
§112
10.2%
-29.8% vs TC avg
Black line = Tech Center average estimate • Based on career data from 51 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Continued Examination Under 37 CFR 1.114 A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 04/23/2026 has been entered. Response to Argument Applicant's arguments filed 03/20/2026 ("Arguments/Remarks") have been fully considered but they are not persuasive. Applicant’s arguments with respect to amended claim have been considered but are moot, because arguments/remarks are directed to amended claim limitations that were not previously examined by the examiner. The rejections are noted in the current office action to address amended claim limitations. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1 and 8 are rejected under 35 U.S.C. 103 as being unpatentable over Vemuri et al., Pub. No.: US11704535B1, in view of Temam et al., Pub. No.: US10504022B2, Huang et al., Pub. No.: US20200117989A1 and Whatmough et al., Pub. No.: US11501151B2. Regarding claim 1, Vemuri teaches: A semiconductor device, comprising: an operator including a multiplier and an accumulator and configured to perform an artificial intelligence operation; (Vemuri, col. 5 line [24 – 34], “ In some embodiments, NNUs 134 1-134 N comprise non-programmable logic i.e., are hardened specialized processing elements. In such embodiments, the NNUs comprise hardware elements [configured to perform an artificial intelligence operation] including, but not limited to, program memories, an instruction fetch/decode unit, fixed-point vector units, floating-point vector units, arithmetic logic units (ALUs), and multiply accumulators (MAC) [A semiconductor device comprising: an operator including a multiplier and an accumulator and]. Although the NNUs 134 1-134 N may be hardened, this does not mean the NNUs are not programmable. That is, the NNUs 134 1-134 N can be configured to perform different operations based on the configuration data.”) wherein the operator (i) reads the feature map data stored in the first memory for a first neural network layer to perform the artificial intelligence operation and stores an operation result of the artificial intelligence operation in the second memory, (Vemuri, col. 1 line [46-67] – col.2 line [1-13], “An integrated circuit (IC) for processing and accelerating data passing through a neural network is disclosed. One example is a reconfigurable IC that includes a digital processing engine (DPE) array having a plurality of DPEs configured to execute one or more layers of a neural network. The reconfigurable IC also includes programmable logic, which includes: an IO controller coupled to input ping-pong buffers the IO controller receiving input data from an interconnect coupled to the programmable logic, wherein the input data fills a first input buffer of the input ping-pong buffers while data stored in a second input buffer of the input ping-pong buffers is processed; a feeding controller coupled to feeding ping-pong buffers the feeding controller receiving input data from the IO controller via the input ping-pong buffers and transmitting the input data through the feeding ping-pong buffers to the DPE array wherein the input data fills a first feeding buffer of the feeding ping-pong buffers [wherein the operator (i) reads the feature map data stored in the first memory for a first neural network layer] while data stored in a second feeding buffer of the feeding ping-pong buffers is processed by the one or more layers executing in the DPE array; [to perform the artificial intelligence operation] a weight controller coupled to weight ping-pong buffers the weight controller receiving weight data from the interconnect and transmitting the weight data through the weight ping-pong buffers to the DPE array wherein the weight data fills a first weight buffer of the weight ping-pong buffers while data stored in a second weight buffer of the weight ping-pong buffers is processed by the one or more layers executing in the DPE array; and an output controller coupled to the plurality of output ping-pong buffers [and stores an operation result of the artificial intelligence operation in the second memory] (i.e.: the output ping-pong buffer corresponds to the “second memory” holding output data after processing) the output controller receiving output data from the one or more layers executing in the DPE array via the output ping-pong buffers, wherein the output data fills a first output buffer of the output ping-pong buffers while data stored in a second output buffer of the output ping-pong buffers is outputted to a host computing system communicatively coupled to the IC.”) Vemuri does not teach: a first memory and a second memory each configured to store feature map data used in the artificial intelligence operation; and a third memory configured to store a training parameter used in the artificial intelligence operation, (ii) reads the feature map data stored in the second memory for a second neural network layer following the first neural network layer to perform the artificial intelligence operation and stores an operation result of the artificial intelligence operation in the first memory. wherein the operator: divides the artificial intelligence operation into a data fetch step, a multiplication step, an accumulation step, and a write memory step; and repeatedly performs the data fetch step, the multiplication step, and the accumulation step, and only after repeatedly performing the data fetch step, the multiplication step, and the accumulation step, perform the write memory step. Temam teaches: a first memory and a second memory each configured to store feature map data used in the artificial intelligence operation; and a third memory configured to store a training parameter used in the artificial intelligence operation, (Temam, abstract, “One embodiment of an accelerator includes a computing unit; a first memory bank for storing input activations [a first memory and a second memory each configured to store feature map data used in the artificial intelligence operation;] and a second memory bank for storing parameters used in performing computations, the second memory bank configured to store a sufficient amount of the neural network parameters [a third memory configured to store a training parameter used in the artificial intelligence operation,] (i.e.: the second memory bank is a third memory) on the computing unit to allow for latency below a specified level with throughput above a specified level. The computing unit includes at least one cell comprising at least one multiply accumulate (“MAC”) operator that receives parameters from the second memory bank and performs computations. The computing unit further includes a first traversal unit that provides a control signal to the first memory bank to cause an input activation to be provided to a data bus accessible by the MAC operator. The computing unit performs computations associated with at least one element of a data array, the one or more computations performed by the MAC operator.”) Temam and Vemuri are related to the same field of endeavor (i.e.: memory management and execution flow design). It would have been obvious to a person having ordinary skill in the art before the effective filing date of the invention to combine the teaching of Temam with teachings of Vemuri to enhance a memory management and execution workflow system by enabling low-latency, high-throughput neural network processing through local storage of sufficient parameters on the accelerator (Temam, Abstract). Vemuri in view of Temam do not teach: (ii) reads the feature map data stored in the second memory for a second neural network layer following the first neural network layer to perform the artificial intelligence operation and stores an operation result of the artificial intelligence operation in the first memory. wherein the operator: divides the artificial intelligence operation into a data fetch step, a multiplication step, an accumulation step, and a write memory step; and repeatedly performs the data fetch step, the multiplication step, and the accumulation step, and only after repeatedly performing the data fetch step, the multiplication step, and the accumulation step, perform the write memory step. Huang teaches: (ii) reads the feature map data stored in the second memory for a second neural network layer following the first neural network layer to perform the artificial intelligence operation and stores an operation result of the artificial intelligence operation in the first memory. (Huang, “[0053] In this embodiment, an artificial intelligence engine 410 may perform a convolutional neural network operation, for example. The artificial intelligence engine 410 accesses the data buffer areas 431, 432, the weight data area 433, and the feature map data areas 434, 435 via respective dedicated memory controllers and repsective dedicated buses. Here, the artificial intelligence engine 410 alternately accesses the feature map data areas 434, 435. For instance, at the very first beginning, after the artificial intelligence engine 410 reads the digitized input data D1 in the data buffer area 431 to perform the convolutional neural network operation, the artificial intelligence engine 410 generates first feature map data F1. The artificial intelligence engine 410 stores the first feature map data F1 in the feature map data area 434. Then, when the artificial intelligence engine 410 performs the next convolution neural network operation, the artificial intelligence engine 410 reads the first feature map data F1 of the feature map data area 434 for the operation, and generates second feature map data F2. The artificial intelligence engine 410 stores the second feature map data F2 in the feature map data area 435. By analogy, the artificial intelligence engine 410 alternately reads the feature map data generated by the previous operation from the memory banks of the feature map data areas 434 or 435 [(ii) reads the feature map data stored in the second memory for a second neural network layer following the first neural network layer to perform the artificial intelligence operation], and then stores current feature map data generated during the current neural network operation in the memory banks of the corresponding feature map data areas 435 or 434 [and stores an operation result of the artificial intelligence operation in the first memory]. Further, in this embodiment, the digitized input data D2 may be simultaneously stored in or read from the data buffer area 432 by the external processor. This implementation is not limited to the convolutional neural network, but is also applicable to other types of networks.”) Huang, Vemuri and Temam are related to the same field of endeavor (i.e.: memory management and execution flow design). It would have been obvious to a person having ordinary skill in the art before the effective filing date of the invention to combine the teaching of Huang with teachings of Vemuri and Temam by integrating the AI engine with the memory array and using a dedicated bus, can achieve faster neural network operation and lower latency, improving overall efficiency and throughput, (Huang, Abstract). Vemuri in view of Temam and Huang do not teach: wherein the operator: divides the artificial intelligence operation into a data fetch step, a multiplication step, an accumulation step, and a write memory step; and repeatedly performs the data fetch step, the multiplication step, and the accumulation step, and only after repeatedly performing the data fetch step, the multiplication step, and the accumulation step, perform the write memory step. Whatmough teaches: wherein the operator: divides the artificial intelligence operation into a data fetch step, a multiplication step, an accumulation step, and a write memory step; (Whatmough, (col. 9 line 64 – col. 10 line 6), “Controller 210 is coupled to control port 206, multiplexor 220, input register 230 [wherein the operator: divides the artificial intelligence operation into a data fetch step], multi-stage add module 240 and output register 250. Generally, controller 210 manages the operation of pipelined accumulator 200 [a multiplication step, an accumulation step], including, for example, reset, input selection for multiplexor 220, etc. Controller 210 may be a simple microcontroller, a programmable circuit, etc. Input port 202 receives a sequence of operands to be accumulated, and output port 204 outputs the final sum of the accumulation calculation [and a write memory step].”) repeatedly performs the data fetch step, the multiplication step, and the accumulation step, (Whatmough, (col. 10 line 65 – 67]), “Generally, pipelined accumulator 200 executes a number of processing cycles that depends upon the number of operands that are to be accumulated [repeatedly performs the data fetch step, the multiplication step, and the accumulation step].”) and only after repeatedly performing the data fetch step, the multiplication step, and the accumulation step, perform the write memory step. (Whatmough, (col. 13 line 52 – 58]), “During the 10th cycle (output) [and only after repeatedly performing the data fetch step, the multiplication step, and the accumulation step], the partial result stored in pipeline register 242 (i.e., [6+9]) is presented to the input of 2nd stage adder 243, which latches the input data, converts the partial result into a complete result, and outputs the complete result (i.e., 15) to output register 250 for storage. Output register 250 then outputs this value to output port 204 [perform the write memory step].”) Whatmough, Vemuri, Temam and Huang are related to the same field of endeavor (i.e.: memory management and execution flow design). It would have been obvious to a person having ordinary skill in the art before the effective filing date of the invention to combine the teaching of Whatmough with teachings of Vemuri, Temam and Huang to add a pipelined accumulation techniques that enables efficient processing of sequential summation operations to reduce processing delays and improved utilization of resources, (Whatmough, Abstract). Regarding claim 8, Vemuri in view of Temam, Huang and Whatmough teach the method of claim 1. Vemuri further teaches: further comprising a post-processor configured to perform post-processing on output data from the artificial intelligence operation and provide the post-processed data to one or more domains. (Vemuri, col. 1 line [46-67] – col.2 line [1-13], “An integrated circuit (IC) for processing and accelerating data passing through a neural network is disclosed. One example is a reconfigurable IC that includes a digital processing engine (DPE) array having a plurality of DPEs configured to execute one or more layers of a neural network. The reconfigurable IC also includes programmable logic, which includes: an IO controller coupled to input ping-pong buffers the IO controller receiving input data from an interconnect coupled to the programmable logic, wherein the input data fills a first input buffer of the input ping-pong buffers while data stored in a second input buffer of the input ping-pong buffers is processed; a feeding controller coupled to feeding ping-pong buffers the feeding controller receiving input data from the IO controller via the input ping-pong buffers and transmitting the input data through the feeding ping-pong buffers to the DPE array wherein the input data fills a first feeding buffer of the feeding ping-pong buffers while data stored in a second feeding buffer of the feeding ping-pong buffers is processed by the one or more layers executing in the DPE array; a weight controller coupled to weight ping-pong buffers the weight controller receiving weight data from the interconnect and transmitting the weight data through the weight ping-pong buffers to the DPE array wherein the weight data fills a first weight buffer of the weight ping-pong buffers while data stored in a second weight buffer of the weight ping-pong buffers is processed by the one or more layers executing in the DPE array; and an output controller coupled to the plurality of output ping-pong buffers the output controller receiving output data from the one or more layers executing in the DPE array via the output ping-pong buffers [further comprising a post-processor configured to perform post-processing on output data from the artificial intelligence operation], wherein the output data fills a first output buffer of the output ping-pong buffers while data stored in a second output buffer of the output ping-pong buffers is outputted to a host computing system communicatively coupled to the IC [provide the post-processed data to one or more domains].”) Claim(s) 5 – 6, 19 – 20 and 26 – 27 are rejected under 35 U.S.C. 103 as being unpatentable over Vemuri in view of Temam, Huang, Whatmough and in further view of Sachs et al., Pub. No.: US20190294975A1. Regarding claim 5, Vemuri in view of Temam, Huang and Whatmough teach the method of claim 1. Whatmough further teaches: the operator performs the write memory step only once whenever performing the data fetch step, the multiplication step, and the accumulation step (Whatmough, (col. 13 line 52 – 58]), “During the 10th cycle (output) [whenever performing the data fetch step, the multiplication step, and the accumulation step], the partial result stored in pipeline register 242 (i.e., [6+9]) is presented to the input of 2nd stage adder 243, which latches the input data, converts the partial result into a complete result, and outputs the complete result (i.e., 15) to output register 250 for storage. Output register 250 then outputs this value to output port 204 [the operator performs the write memory step only once].”) It would have been obvious to one of ordinary skill in the art before the effective filling date of the present application to combine the teachings of Whatmough with teachings of Vemuri, Temam and Huang for the same reasons disclosed for claim 1. Vemuri in view of Temam, Huang and Whatmough do not teach: wherein when the feature map data includes N rows and M columns, and when the neural network layer corresponds to a column layer N times, wherein N is an integer greater than or equal to two, and M is an integer greater than or equal to two. Sachs teaches: wherein when the feature map data includes N rows and M columns, (Sachs, “[0060] Each feature map has columns, one column [and M columns] per time step. Each row [wherein when the feature map data includes N rows] of a feature map has results from a different convolutional filter of a convolutional neural network layer 506, 512. In a preferred embodiment, each convolutional neural network layer 506 has a plurality of different convolutional filters which are the same height as a column of the input tensor but which have different widths, where the widths correspond to numbers of time steps. By using convolutional filters which are the same height as one another efficiencies are gained without significantly sacrificing accuracy. In other example, both the width and height of the convolutional filters varies.”) and when the neural network layer corresponds to a column layer (Sachs, “[0061] The effect of a convolutional neural network layer [and when the neural network layer] can be thought of as sliding each convolutional filter over the input tensor, from column to column [corresponds to a column layer], and computing a convolution, which is an aggregation of the neural network node signals falling within the footprint of the filter, at each position of the filter as it is slid from column to column. This gives a convolution result, which aggregates each schema field over the time steps that fall within the footprint of the filter. For a given column, there is a convolution result from each convolutional filter. One of the convolution results is selected and stored in the corresponding feature map column. In an example, the selection is done by selecting the maximum convolution result.”) N times, wherein N is an integer greater than or equal to two, and M is an integer greater than or equal to two. (Sachs, “[0058] FIG. 5A shows an input structure comprising a tensor 500 of columns and rows. The tensor is equivalent to an image in many respects and so is operable with machine learning technology typically used for image processing. The columns 502 [and M is an integer greater than or equal to two] each contain state data in a different time step and the rows 504 contain schema fields [N times, wherein N is an integer greater than or equal to two]. The input structure is input to a convolutional neural network layer 506 which computes a feature map 508 as output. The feature map is input to a second convolutional neural network layer 512 which computes a second feature map 514 as output.”) Sachs, Vemuri, Temam, Huang and Whatmough are related to the same field of endeavor (i.e.: memory management and execution flow design). It would have been obvious to a person having ordinary skill in the art before the effective filing date of the invention to combine the teaching of Sachs with teachings of Vemuri, Temam, Huang and Whatmough by comparing and aggregating information from multiple digital twins, it enables distribution inference, pattern recognition and relationship mapping to improve decision-making and system coordination, (Sachs, Abstract). Claim(s) 19 and 26, recite limitations analogous to claim 5, so are rejected under the same rationale. Regarding claim 6, Vemuri in view of Temam, Huang and Whatmough teach the method of claim 1. Whatmough further teaches: the operator performs the write memory step only once whenever performing the data fetch step, the multiplication step, and the accumulation step (Whatmough, (col. 13 line 52 – 58]), “During the 10th cycle (output) [whenever performing the data fetch step, the multiplication step, and the accumulation step], the partial result stored in pipeline register 242 (i.e., [6+9]) is presented to the input of 2nd stage adder 243, which latches the input data, converts the partial result into a complete result, and outputs the complete result (i.e., 15) to output register 250 for storage. Output register 250 then outputs this value to output port 204 [the operator performs the write memory step only once].”) It would have been obvious to one of ordinary skill in the art before the effective filling date of the present application to combine the teachings of Whatmough with teachings of Vemuri, Temam and Huang for the same reasons disclosed for claim 1. Vemuri in view of Temam, Huang and Whatmough do not teach: wherein when the feature map data includes N rows and M columns, and when the neural network layer corresponds to a row layer, M times, wherein N is an integer greater than or equal to two, and M is an integer greater than or equal to two. Sachs teaches: wherein when the feature map data includes N rows and M columns, and when the neural network layer corresponds to a row layer, (Sachs, “[0059] Using a convolutional neural network in the context of the present technology gives unexpected benefits. Typically convolutional neural networks [when the neural network layer] are used for image processing where spatial information is contained in the image so that there are relationships expected between rows [corresponds to a row layer] and columns of the image. In contrast, the present technology does not use images as inputs but rather has matrices formed from time steps of data from schema fields of event streams. Relationships are not expected between the schema field data. However, it is unexpectedly found that using convolution where the convolutional filters span both one or more time steps and one or more schema fields gives good quality prediction results.”) M times, wherein N is an integer greater than or equal to two, and M is an integer greater than or equal to two. (Sachs, “[0058] FIG. 5A shows an input structure comprising a tensor 500 of columns and rows [N times]. The tensor is equivalent to an image in many respects and so is operable with machine learning technology typically used for image processing. The columns 502 [and M is an integer greater than or equal to two] each contain state data in a different time step and the rows 504 contain schema fields [wherein N is an integer greater than or equal to two]. The input structure is input to a convolutional neural network layer 506 which computes a feature map 508 as output. The feature map is input to a second convolutional neural network layer 512 which computes a second feature map 514 as output.”) Sachs, Vemuri, Temam, Huang and Whatmough are related to the same field of endeavor (i.e.: memory management and execution flow design). It would have been obvious to a person having ordinary skill in the art before the effective filing date of the invention to combine the teaching of Sachs with teachings of Vemuri, Temam, Huang and Whatmough by comparing and aggregating information from multiple digital twins, it enables distribution inference, pattern recognition and relationship mapping to improve decision-making and system coordination, (Sachs, Abstract). Claim(s) 20 and 27, recite limitations analogous to claim 6 , so are rejected under the same rationale. Claim 7 is rejected under 35 U.S.C. 103 as being unpatentable over Vemuri in view of Temam, Huang, Whatmough and in further view of Motozuka et al., Pub. No.: US20160241360A1. Vemuri in view of Temam, Huang and Whatmough teach the method of claim 1. Vemuri further teaches: provide the pre-processed data to the first memory or the second memory. (Vemuri, col. 1 line [46-67] – col.2 line [1-13], “An integrated circuit (IC) for processing and accelerating data passing through a neural network is disclosed. One example is a reconfigurable IC that includes a digital processing engine (DPE) array having a plurality of DPEs configured to execute one or more layers of a neural network. The reconfigurable IC also includes programmable logic, which includes: an IO controller coupled to input ping-pong buffers the IO controller receiving input data from an interconnect coupled to the programmable logic, wherein the input data fills a first input buffer of the input ping-pong buffers while data stored in a second input buffer of the input ping-pong buffers is processed [provide the pre-processed data to the first memory or the second memory] (i.e.: input ping-pong buffers correspond to the “first memory”); a feeding controller coupled to feeding ping-pong buffers the feeding controller receiving input data from the IO controller via the input ping-pong buffers and transmitting the input data through the feeding ping-pong buffers to the DPE array wherein the input data fills a first feeding buffer of the feeding ping-pong buffers while data stored in a second feeding buffer of the feeding ping-pong buffers is processed by the one or more layers executing in the DPE array;”) Vemuri in view of Temam, Huang and Whatmough do not teach: further comprising a pre-processor configured to perform pre-processing on input data from one or more domains for the artificial intelligence operation and Motozuka teaches: further comprising a pre-processor configured to perform pre-processing on input data from one or more domains for the artificial intelligence operation and (Motozuka, “[0190] Pieces D1 and D2 of input data are input to the preprocessing section 301 [further comprising a pre-processor configured to perform pre-processing on input data from one or more domains for the artificial intelligence operation] while pieces D3 and D4 of input data are input to the minimum value selection section 302. That is, the four pieces D1 to D4 of data are input in parallel to the minimum value selection circuit 300, as in the second embodiment.”) Motozuka, Vemuri, Temam, Huang and Whatmough are related to the same field of endeavor (i.e.: memory management and execution flow design). It would have been obvious to a person having ordinary skill in the art before the effective filing date of the invention to combine the teaching of Motozuka with teachings of Vemuri, Temam, Huang and Whatmough to reduce computational load on general-purpose processors and improves performance for tasks like sorting, selection or optimization, (Motozuka, Abstract). Claim(s) 9 – 10 are rejected under 35 U.S.C. 103 as being unpatentable over Vemuri in view of Temam, Huang, Whatmough and in further view of Tomida et al., Pub. No.: US20230316071A1. Regarding claim 9, Vemuri in view of Temam, Huang and Whatmough teach the method of claim 1. Vemuri in view of Temam, Huang and Whatmough do not teach: wherein the operator includes a first operator that performs a first artificial intelligence operation and a second operator that performs a second artificial intelligence operation different from the first artificial intelligence operation, and the first operator uses, for the neural network layer, a partial area of the first memory and a partial area of the second memory as a third space for storing data before the first artificial intelligence operation and a fourth space for storing data after the first artificial intelligence operation, respectively. Tomida teaches: wherein the operator includes a first operator that performs a first artificial intelligence operation and a second operator that performs a second artificial intelligence operation different from the first artificial intelligence operation, and (Tomida, “[0131] The convolution operation circuit 4 reads a portion of the input data at from the first memory 1 and performs a layer-(2M−1) convolution operation [wherein the operator includes a first operator that performs a first artificial intelligence operation] (i.e.: first operation = convolution operation) that outputs the partial tensor ft (referred to as the first partial tensor ft1). The first partial tensor ft1 is written into the second memory 2. Before implementing a convolution operation with respect to the remainder of the input data a stored in the first memory 1, the quantization operation circuit 5 performs a layer-2M quantization operation [and a second operator that performs a second artificial intelligence operation different from the first artificial intelligence operation] (i.e.: second operation = quantization operation) corresponding to the first partial tensor ft1 stored in the second memory 2. The output data from the layer-2M quantization operation is written into the first memory 1. As a result thereof, the first partial tensor ft1 written into the second memory 2 becomes unnecessary.”) the first operator uses, for the neural network layer, a partial area of the first memory and a partial area of the second memory as a third space for storing data before the first artificial intelligence operation (Tomida, “[0131] The convolution operation [the first operator uses, for the neural network layer] circuit 4 reads a portion of the input data at from the first memory 1 [a partial area of the first memory] and performs a layer-(2M−1) convolution operation that outputs the partial tensor ft (referred to as the first partial tensor ft1). The first partial tensor ft1 is written into the second memory 2 [and a partial area of the second memory as a third space for storing data]. Before implementing a convolution operation [before the first artificial intelligence operation] (i.e.: a partial tensor, of the input data is stored/used from the first/second memory before convolution) with respect to the remainder of the input data a stored in the first memory 1, the quantization operation circuit 5 performs a layer-2M quantization operation corresponding to the first partial tensor ft1 stored in the second memory 2. The output data from the layer-2M quantization operation is written into the first memory 1. As a result thereof, the first partial tensor ft1 written into the second memory 2 becomes unnecessary.”) and a fourth space for storing data after the first artificial intelligence operation, respectively. (Tomida, “[0104] In the case in which an operation that cannot be implemented by the NN execution model 100 is included among the operations of the CNN 200, the NN execution model 100 transfers intermediate data to an external operation device such as an external host CPU. After the external operation device has performed operations [after the first artificial intelligence operation, respectively] on the intermediate data, operation results by the external operation device are input to the first memory 1 or the second memory 2 [and a fourth space for storing data]. The NN execution model 100 resumes operations on the operation results by the external operation device.”) Tomida, Vemuri, Temam, Huang and Whatmough are related to the same field of endeavor (i.e.: memory management and execution flow design). It would have been obvious to a person having ordinary skill in the art before the effective filing date of the invention to combine the teaching of Tomida with teachings of Vemuri, Temam, Huang and Whatmough to add automated, hardware-aware neural network generation to the system by creating execution models tailored to the specific hardware on which it will run, enables optimized neural network deployment with improved performance, (Tomida, ¶[0008] – [0010]). Regarding claim 10, Vemuri in view of Temam, Huang, Whatmough and Tomida teach the method of claim 9. Tomida further teaches: wherein the second operator uses, for the neural network layer, another partial area of the first memory and another partial area of the second memory as a fifth space for storing data before the second artificial intelligence operation (Tomida, “[0136] As illustrated in FIG. 11 , the convolution operation circuit 4 reads a portion of the partial tensor as (first partial tensor as1) from the first memory 1 and performs a layer-(2M−1) convolution operation that outputs the partial tensor ft (referred to as the first partial tensor ft1). The first partial tensor ft1 is written into the second memory 2. Before implementing a convolution operation with respect to the remainder of the first partial tensor as1 stored in the first memory 1, the quantization operation circuit 5 [wherein the second operator uses, for the neural network layer] performs a layer-2M quantization operation corresponding to the first partial tensor ft1 stored in the second memory 2 [and another partial area of the second memory as a fifth space for storing data before the second artificial intelligence operation]. The output data from the layer-2M quantization operation is written into the first memory 1 [another partial area of the first memory]. As a result thereof, the first partial tensor ft1 written into the second memory 2 becomes unnecessary.”) and a sixth space for storing data after the second artificial intelligence operation. (Tomida, “[0095] The DMAC 3 stores the input data a input to layer-1 (see FIG. 3 ) in the first memory 1 [and a sixth space for storing data]. The DMAC 3 may transfer the input data a input to layer-1 after partitioning the data in accordance with the order of convolution operations performed by the convolution operation circuit 4 [after the second artificial intelligence operation].”) It would have been obvious to one of ordinary skill in the art before the effective filling date of the present application to combine the teachings of Tomida with teachings of Vemuri, Temam, Huang and Whatmough for the same reasons disclosed for claim 9. Claim(s) 16 – 17 are rejected under 35 U.S.C. 103 as being unpatentable over Tanaka et al. Pub. No.: US20180329536A1, in view of Tolias et al., Pub. No.: US20220284288, Vemuri, Temam, Huang and Whatmough. Regarding claim 16, Tanaka teaches: A semiconductor system, comprising: a display driver configured to drive a display panel based on input image data; (Tanaka, “[0023] FIG. 1 is a block diagram schematically illustrating a display device 100 in accordance with some embodiments. In one embodiment, the display device 100 includes an LCD panel 1 and a touch controller-embedded display driver 2. The display device 100 may be configured to receive image data from an application processor 3 and display an image [a display driver configured to drive a display panel] corresponding to the received image data on the LCD panel 1 [based on input image data]. The display device 100 may be configured to perform touch sensing of user input such as a position at which a conductor, such as a human finger and/or a stylus, is in contact with the LCD panel 1.”) a touch controller configured to convert a touch sensing signal received from a touch sensor into touch sensing data; (Tanaka, “[0010] In another embodiment, a display device comprises a liquid crystal display panel including a plurality of pixel circuits and a plurality of sense electrodes, and display driver circuitry configured to drive the plurality of pixel circuits. The display device further comprises touch controller circuitry configured to obtain capacitance detection data depending on capacitances of the plurality of sense electrodes. The touch controller circuitry [a touch controller] is further configured to generate touch sensing data [into touch sensing data] associated with a current touch sensing frame based on capacitance detection data associated with the current touch sensing frame [configured to convert a touch sensing signal received from a touch sensor] and capacitance detection data associated with a former touch sensing frame selected in response to a state in which the liquid crystal display panel is placed in the current touch sensing frame.”) a host processor configured to provide the input image data to the display driver and receives the touch sensing data from the touch controller; and (Tanaka, “[0023] FIG. 1 is a block diagram schematically illustrating a display device 100 in accordance with some embodiments. In one embodiment, the display device 100 includes an LCD panel 1 and a touch controller-embedded display driver 2. The display device 100 may be configured to receive image data from an application processor 3 and display an image corresponding to the received image data on the LCD panel 1. The display device 100 [a host processor configured to provide the input image data to the display driver] may be configured to perform touch sensing [and receives the touch sensing data from the touch controller] of user input such as a position at which a conductor, such as a human finger and/or a stylus, is in contact with the LCD panel 1.”) Tanaka does not teach: an artificial intelligence unit configured to perform an artificial intelligence operation generating predictive noise data corresponding to the input image data, wherein the artificial intelligence unit includes: an operator including a multiplier and an accumulator and configured to perform the artificial intelligence operation; a first memory and a second memory each configured to store feature map data used in the artificial intelligence operation; and a third memory configured to store a training parameter used in the artificial intelligence operation, wherein the operator (i) reads the feature map data stored in the first memory for a first neural network layer to perform the artificial intelligence operation and stores an operation result of the artificial intelligence operation in the second memory, and (ii) reads the feature map data stored in the second memory for a second neural network layer following the first neural network layer to perform the artificial intelligence operation and stores an operation result of the artificial intelligence operation in the first memory. wherein the operator: divides the artificial intelligence operation into a data fetch step, a multiplication step, an accumulation step, and a write memory step; and accumulation step, and only after repeatedly performing the data fetch step, the multiplication step, and the accumulation step, perform the write memory step. Tolias teaches: an artificial intelligence unit configured to perform an artificial intelligence operation generating predictive noise data corresponding to the input image data, wherein the artificial intelligence unit includes: (Tolias, “[0055] As discussed herein in detail, most neural response data such as cortex scans from in vivo experimental conditions are too noisy for regularizing task models. Accordingly, various embodiments are directed to in silico models that can take neural stimuli such as images [corresponding to the input image data] and predict the neural response of a biological system to the neural stimuli. The use of in silico model neuron responses as a proxy for the real in vivo neurons enables isolation of the relevant features from the biological system (e.g., brain) for use to regularize the artificial intelligence system [an artificial intelligence unit configured to perform an artificial intelligence operation]. For example, the in silico predictive model eliminates random noise [generating predictive noise data] and the model's shifter and modulator circuits can be configured to account for the irrelevant non-stimuli data such as eye and body movements, and thereby extract the purely visual stimuli-driven responses. Extensions of the in silico predictive model can be used to extract all kinds of other features from the brain including the structure of the noise which can also be used as an additional regularizer. Mimicking neural noise can be used to bias AI models towards more probabilistic representations of sensory information.”) Tolias and Tanaka are related to the same field of endeavor (i.e.: memory management and execution flow design). It would have been obvious to a person having ordinary skill in the art before the effective filing date of the invention to combine the teaching of Tolias with teachings of Tanaka to enhance learning efficiency, robustness and adaptability across diverse AI domain such as perception, memory and decision making (Tolias, Abstract). Tanaka and Tolias do not teach: an operator including a multiplier and an accumulator and configured to perform the artificial intelligence operation; a first memory and a second memory each configured to store feature map data used in the artificial intelligence operation; and a third memory configured to store a training parameter used in the artificial intelligence operation, wherein the operator (i) reads the feature map data stored in the first memory for a first neural network layer to perform the artificial intelligence operation and stores an operation result of the artificial intelligence operation in the second memory, and (ii) reads the feature map data stored in the second memory for a second neural network layer following the first neural network layer to perform the artificial intelligence operation and stores an operation result of the artificial intelligence operation in the first memory. wherein the operator: divides the artificial intelligence operation into a data fetch step, a multiplication step, an accumulation step, and a write memory step; and accumulation step, and only after repeatedly performing the data fetch step, the multiplication step, and the accumulation step, perform the write memory step. Temam teaches: an operator including a multiplier and an accumulator and configured to perform the artificial intelligence operation; (Temam, col. 16, line [48 – 64], “Accordingly, the activations feed one of the inputs of each MAC operator 215 and each MAC operator 215 in the cells of MAC array 214 get their second multiplier [an operator including a multiplier] input from wide memory 212. At block 608 of process 600, MAC array 214 of compute tile 200 performs tensor computations comprising dot product computations based on elements of a data array structure accessed from memory. Wide memory 212 can have a width in bits that is equal to the width of the linear unit (e.g., 32-bits). The linear unit (LU) is thus a SIMD vector arithmetic logic unit (ALU) unit that receives data from a vector memory (i.e., wide memory 212). In some implementations, MAC operators 215 may also get the accumulator [and an accumulator and configured to perform the artificial intelligence operation] inputs (partial sums) from wide memory 212 as well. In some implementations, there is time sharing relative to the wide memory 212 port for reads and/or writes relating to the two different operands (parameters and partial sum).”) a first memory and a second memory each configured to store feature map data used in the artificial intelligence operation; and a third memory configured to store a training parameter used in the artificial intelligence operation, (Temam, abstract, “One embodiment of an accelerator includes a computing unit; a first memory bank for storing input activations [a first memory and a second memory each configured to store feature map data used in the artificial intelligence operation;] and a second memory bank for storing parameters used in performing computations, the second memory bank configured to store a sufficient amount of the neural network parameters [a third memory configured to store a training parameter used in the artificial intelligence operation,] (i.e.: the second memory bank is a third memory) on the computing unit to allow for latency below a specified level with throughput above a specified level. The computing unit includes at least one cell comprising at least one multiply accumulate (“MAC”) operator that receives parameters from the second memory bank and performs computations. The computing unit further includes a first traversal unit that provides a control signal to the first memory bank to cause an input activation to be provided to a data bus accessible by the MAC operator. The computing unit performs computations associated with at least one element of a data array, the one or more computations performed by the MAC operator.”) Temam, Tanaka and Tolias are related to the same field of endeavor (i.e.: memory management and execution flow design). It would have been obvious to a person having ordinary skill in the art before the effective filing date of the invention to combine the teaching of Temam with teachings of Tanaka and Tolias to add low-latency, high-throughput neural network acceleration to the system by storing sufficient parameters locally on the accelerator and optimizing data flow between memory banks and computation units. (Temam, Abstract). Tanaka, Tolias and Temam do not teach: wherein the operator (i) reads the feature map data stored in the first memory for a first neural network layer to perform the artificial intelligence operation and stores an operation result of the artificial intelligence operation in the second memory, (ii) reads the feature map data stored in the second memory for a second neural network layer following the first neural network layer to perform the artificial intelligence operation and stores an operation result of the artificial intelligence operation in the first memory. wherein the operator: divides the artificial intelligence operation into a data fetch step, a multiplication step, an accumulation step, and a write memory step; and accumulation step, and only after repeatedly performing the data fetch step, the multiplication step, and the accumulation step, perform the write memory step. Vemuri teaches: wherein the operator (i) reads the feature map data stored in the first memory for a first neural network layer to perform the artificial intelligence operation and stores an operation result of the artificial intelligence operation in the second memory, (Vemuri, col. 1 line [46-67] – col.2 line [1-13], “An integrated circuit (IC) for processing and accelerating data passing through a neural network is disclosed. One example is a reconfigurable IC that includes a digital processing engine (DPE) array having a plurality of DPEs configured to execute one or more layers of a neural network. The reconfigurable IC also includes programmable logic, which includes: an IO controller coupled to input ping-pong buffers the IO controller receiving input data from an interconnect coupled to the programmable logic, wherein the input data fills a first input buffer of the input ping-pong buffers while data stored in a second input buffer of the input ping-pong buffers is processed; a feeding controller coupled to feeding ping-pong buffers the feeding controller receiving input data from the IO controller via the input ping-pong buffers and transmitting the input data through the feeding ping-pong buffers to the DPE array wherein the input data fills a first feeding buffer of the feeding ping-pong buffers [wherein the operator (i) reads the feature map data stored in the first memory for a first neural network layer] while data stored in a second feeding buffer of the feeding ping-pong buffers is processed by the one or more layers executing in the DPE array; [to perform the artificial intelligence operation] a weight controller coupled to weight ping-pong buffers the weight controller receiving weight data from the interconnect and transmitting the weight data through the weight ping-pong buffers to the DPE array wherein the weight data fills a first weight buffer of the weight ping-pong buffers while data stored in a second weight buffer of the weight ping-pong buffers is processed by the one or more layers executing in the DPE array; and an output controller coupled to the plurality of output ping-pong buffers [and stores an operation result of the artificial intelligence operation in the second memory] (i.e.: the output ping-pong buffer corresponds to the “second memory” holding output data after processing) the output controller receiving output data from the one or more layers executing in the DPE array via the output ping-pong buffers, wherein the output data fills a first output buffer of the output ping-pong buffers while data stored in a second output buffer of the output ping-pong buffers is outputted to a host computing system communicatively coupled to the IC.”) Vemuri, Tanaka, Tolias and Temam are related to the same field of endeavor (i.e.: memory management and execution flow design). It would have been obvious to a person having ordinary skill in the art before the effective filing date of the invention to combine the teaching of Vemuri with teachings of Tanaka, Tolias and Temam to improve optimization on neural network processing to the system by locally storing key parameters to meet strict latency and throughput requirements. (Vemuri, Abstract). Tanaka in view of Tolias, Temam and Vemuri do not teach: (ii) reads the feature map data stored in the second memory for a second neural network layer following the first neural network layer to perform the artificial intelligence operation and stores an operation result of the artificial intelligence operation in the first memory. wherein the operator: divides the artificial intelligence operation into a data fetch step, a multiplication step, an accumulation step, and a write memory step; and accumulation step, and only after repeatedly performing the data fetch step, the multiplication step, and the accumulation step, perform the write memory step. Huang teaches: (ii) reads the feature map data stored in the second memory for a second neural network layer following the first neural network layer to perform the artificial intelligence operation and stores an operation result of the artificial intelligence operation in the first memory. (Huang, “[0053] In this embodiment, an artificial intelligence engine 410 may perform a convolutional neural network operation, for example. The artificial intelligence engine 410 accesses the data buffer areas 431, 432, the weight data area 433, and the feature map data areas 434, 435 via respective dedicated memory controllers and repsective dedicated buses. Here, the artificial intelligence engine 410 alternately accesses the feature map data areas 434, 435. For instance, at the very first beginning, after the artificial intelligence engine 410 reads the digitized input data D1 in the data buffer area 431 to perform the convolutional neural network operation, the artificial intelligence engine 410 generates first feature map data F1. The artificial intelligence engine 410 stores the first feature map data F1 in the feature map data area 434. Then, when the artificial intelligence engine 410 performs the next convolution neural network operation, the artificial intelligence engine 410 reads the first feature map data F1 of the feature map data area 434 for the operation, and generates second feature map data F2. The artificial intelligence engine 410 stores the second feature map data F2 in the feature map data area 435. By analogy, the artificial intelligence engine 410 alternately reads the feature map data generated by the previous operation from the memory banks of the feature map data areas 434 or 435 [(ii) reads the feature map data stored in the second memory for a second neural network layer following the first neural network layer to perform the artificial intelligence operation], and then stores current feature map data generated during the current neural network operation in the memory banks of the corresponding feature map data areas 435 or 434 [and stores an operation result of the artificial intelligence operation in the first memory]. Further, in this embodiment, the digitized input data D2 may be simultaneously stored in or read from the data buffer area 432 by the external processor. This implementation is not limited to the convolutional neural network, but is also applicable to other types of networks.”) Huang, Tanaka, Tolias, Temam and Vemuri are related to the same field of endeavor (i.e.: memory management and execution flow design). It would have been obvious to a person having ordinary skill in the art before the effective filing date of the invention to combine the teaching of Huang with teachings of Tanaka, Tolias, Temam and Vemuri by integrating the AI engine with the memory array and using a dedicated bus, can achieve faster neural network operation and lower latency, improving overall efficiency and throughput, (Huang, Abstract). Tanaka in view of Tolias, Temam, Vemuri and Huang do not teach: wherein the operator: divides the artificial intelligence operation into a data fetch step, a multiplication step, an accumulation step, and a write memory step; and accumulation step, and only after repeatedly performing the data fetch step, the multiplication step, and the accumulation step, perform the write memory step. Whatmough teaches: wherein the operator: divides the artificial intelligence operation into a data fetch step, a multiplication step, an accumulation step, and a write memory step; (Whatmough, (col. 9 line 64 – col. 10 line 6), “Controller 210 is coupled to control port 206, multiplexor 220, input register 230 [wherein the operator: divides the artificial intelligence operation into a data fetch step], multi-stage add module 240 and output register 250. Generally, controller 210 manages the operation of pipelined accumulator 200 [a multiplication step, an accumulation step], including, for example, reset, input selection for multiplexor 220, etc. Controller 210 may be a simple microcontroller, a programmable circuit, etc. Input port 202 receives a sequence of operands to be accumulated, and output port 204 outputs the final sum of the accumulation calculation [and a write memory step].”) repeatedly performs the data fetch step, the multiplication step, and the accumulation step, (Whatmough, (col. 10 line 65 – 67]), “Generally, pipelined accumulator 200 executes a number of processing cycles that depends upon the number of operands that are to be accumulated [repeatedly performs the data fetch step, the multiplication step, and the accumulation step].”) and only after repeatedly performing the data fetch step, the multiplication step, and the accumulation step, perform the write memory step. (Whatmough, (col. 13 line 52 – 58]), “During the 10th cycle (output) [and only after repeatedly performing the data fetch step, the multiplication step, and the accumulation step], the partial result stored in pipeline register 242 (i.e., [6+9]) is presented to the input of 2nd stage adder 243, which latches the input data, converts the partial result into a complete result, and outputs the complete result (i.e., 15) to output register 250 for storage. Output register 250 then outputs this value to output port 204 [perform the write memory step].”) Whatmough, Tanaka, Tolias, Temam, Vemuri and Huang are related to the same field of endeavor (i.e.: memory management and execution flow design). It would have been obvious to a person having ordinary skill in the art before the effective filing date of the invention to combine the teaching of Whatmough with teachings of Tanaka, Tolias, Temam, Vemuri and Huang to add a pipelined accumulation techniques that enables efficient processing of sequential summation operations to reduce processing delays and improved utilization of resources, (Whatmough, Abstract). Regarding claim 17, Tanaka in view of Tolias, Temam, Vemuri, Huang and Whatmough teach the method of claim 16. Tanaka further teaches: wherein the artificial intelligence unit is installed in one of the display driver, the touch controller, and the host processor. (Tanaka, “[0023] FIG. 1 is a block diagram schematically illustrating a display device 100 in accordance with some embodiments. In one embodiment, the display device 100 includes an LCD panel 1 and a touch controller-embedded display driver 2 [wherein the artificial intelligence unit is installed in one of the display driver, the touch controller]. The display device 100 may be configured to receive image data from an application processor 3 [and the host processor] and display an image corresponding to the received image data on the LCD panel 1. The display device 100 may be configured to perform touch sensing of user input such as a position at which a conductor, such as a human finger and/or a stylus, is in contact with the LCD panel 1.”) Claim 23 is rejected under 35 U.S.C. 103 as being unpatentable over Agarwal et al. Pub. No.: US11018926B2, in view of TAO et al., Pub. No.: US20230297386A1, Temam and Whatmough. Agarwal teaches: A semiconductor system, comprising: a first pair of devices comprising a first device and a second device that exchange data in a first domain; a second pair of devices comprising a third device and a fourth device that exchange data in a second domain different from the first domain; (Agarwal, (col. 3, line [7 – 10]), “First domain 105 may comprise a first plurality of devices 115 [a first pair of devices comprising a first device and a second device] in first domain 105 [that exchange data in a first domain] and second domain 110 [that exchange data in a second domain different from the first domain] may comprise a second plurality of devices 120 [a second pair of devices comprising a third device and a fourth device] in second domain 110.”) an artificial intelligence unit that performs an artificial intelligence operation on data in the first domain or data in the second domain; (Agarwal, (col. 3, line [13 – 21]), “Each of first plurality of devices 115 in first domain 105 [an artificial intelligence unit that performs an artificial intelligence operation on data in the first domain or data in the second domain] may comprise, but are not limited to, a host, a router, or a switch. As shown in FIG. 1A, first plurality of devices 115 may be connected by links that carry data traffic between first plurality of devices 115. For example, first device 125 and second device 130 may be connected in first domain 105 with data traffic passing through third device 135 for example.”) Agarwal does not teach: a first pre/post-processor that performs first pre-processing to provide the data the first domain to the artificial intelligence unit or that performs first post-processing to provide an operation result of the artificial intelligence unit to the first domain; and a second pre/post-processor that performs second pre-processing to provide the data in the second domain to the artificial intelligence unit or that performs second post-processing to provide an operation result of the artificial intelligence unit to the second domain. an operator performing the artificial intelligence operation; a first memory and a second memory that store feature map data used in the artificial intelligence operation; and a third memory that stores a training parameter used in the artificial intelligence operation, wherein the operator uses, for a neural network layer, the first memory and the second memory as a first space for storing data before the artificial intelligence operation and a second space for storing data after the artificial intelligence operation, respectively, and wherein the operator: divides the artificial intelligence operation into a data fetch step, a multiplication step, an accumulation step, and a write memory step; and repeatedly performs the data fetch step, the multiplication step, and the accumulation step, and only after repeatedly performing the data fetch step, the multiplication step, and the accumulation step, perform the write memory step. TAO teaches: a first pre/post-processor that performs first pre-processing to provide the data the first domain to the artificial intelligence unit or that performs first post-processing to provide an operation result of the artificial intelligence unit to the first domain; and (TAO, “[0035] As shown in FIG. 3, the master processing circuit 300 may include a data processing unit 302, a first-group pipeline operation circuit 304, a last-group pipeline operation circuit 306, and one or a plurality of groups of pipeline operation circuits (which are replaced by black circles) between the first-group pipeline operation circuit 304 and the last-group pipeline operation circuit 306. In an embodiment, the data processing unit 302 includes a data conversion circuit 3021 and a data concatenation circuit 3022 [a first pre/post-processor that performs first pre-processing]. As described earlier, when a master operation includes a pre-processing operation for a slave operation, such as a data conversion operation or a data concatenation operation [to provide the data the first domain to the artificial intelligence unit] (i.e.: the data handled here originates in the “master” domain and is prepared so the downstream processor (slave operation) can consume it), the data conversion circuit 3021 or the data concatenation circuit 3022 may perform a corresponding conversion operation or concatenation operation according to a corresponding master instruction. The following will explain the conversion operation and the concatenation operation with examples.”) a second pre/post-processor that performs second pre-processing to provide the data in the second domain to the artificial intelligence unit or that performs second post-processing to provide an operation result of the artificial intelligence unit to the second domain. (TAO, “[0032] ... The master memory may be used to store related operation data such as a neuron, a weight, and various constant terms. The master caching unit may be used to store intermediate data temporarily, such as data after the pre-processing operation [a second pre/post-processor that performs second pre-processing] and data before the post-processing operation, and these pieces of intermediate data [to provide the data in the second domain to the artificial intelligence unit] may be invisible to an operator according to settings.”) TAO and Agarwal are related to the same field of endeavor (i.e.: memory management and execution flow design). It would have been obvious to a person having ordinary skill in the art before the effective filing date of the invention to combine the teaching of TAO with teachings of Agarwal to add cooperative, multi-processor workload sharing, where multiple specialized and general-purpose processing units access shared storage through a unified interconnection interface. (TAO, Abstract). Agarwal in view of TAO do not teach: an operator performing the artificial intelligence operation; a first memory and a second memory that store feature map data used in the artificial intelligence operation; and a third memory that stores a training parameter used in the artificial intelligence operation, wherein the operator uses, for a neural network layer, the first memory and the second memory as a first space for storing data before the artificial intelligence operation and a second space for storing data after the artificial intelligence operation, respectively, and wherein the operator: divides the artificial intelligence operation into a data fetch step, a multiplication step, an accumulation step, and a write memory step; and repeatedly performs the data fetch step, the multiplication step, and the accumulation step, and only after repeatedly performing the data fetch step, the multiplication step, and the accumulation step, perform the write memory step. Temam teaches: an operator performing the artificial intelligence operation; a first memory and a second memory that store feature map data used in the artificial intelligence operation; and a third memory that stores a training parameter used in the artificial intelligence operation, (Temam, abstract, “One embodiment of an accelerator includes a computing unit; a first memory bank for storing input activations [an operator performing the artificial intelligence operation; a first memory and a second memory that store feature map data used in the artificial intelligence operation] and a second memory bank for storing parameters used in performing computations, the second memory bank configured to store a sufficient amount of the neural network parameters [and a third memory that stores a training parameter used in the artificial intelligence operation] (i.e.: the second memory bank is a third memory) on the computing unit to allow for latency below a specified level with throughput above a specified level.”) wherein the operator uses, for a neural network layer, the first memory and the second memory as a first space for storing data before the artificial intelligence operation and a second space for storing data after the artificial intelligence operation, respectively, and (Temam, (col. 12 line [45 – 53]), “The first portion is complete when multiply operations produce an output activation [wherein the operator uses, for a neural network layer, the first memory and the second memory as a first space for storing data before the artificial intelligence operation], for example, by completing a multiplication of an input activation and a parameter to generate the output activation. The second portion includes application of a non-linear function to an output activation and the second portion is complete when the output activation is written to narrow memory 210 [and the second memory as a first space for storing data before the artificial intelligence operation and a second space for storing data after the artificial intelligence operation, respectively] after application of the function.”) Temam, Agarwal and TAO are related to the same field of endeavor (i.e.: memory management and execution flow design). It would have been obvious to a person having ordinary skill in the art before the effective filing date of the invention to combine the teaching of Temam with teachings of Agarwal and TAO to enhance a memory management and execution workflow system by enabling low-latency, high-throughput neural network processing through local storage of sufficient parameters on the accelerator (Temam, Abstract). Agarwal in view of TAO and Temam do not teach: wherein the operator: divides the artificial intelligence operation into a data fetch step, a multiplication step, an accumulation step, and a write memory step; and repeatedly performs the data fetch step, the multiplication step, and the accumulation step, and only after repeatedly performing the data fetch step, the multiplication step, and the accumulation step, perform the write memory step. Whatmough teaches: wherein the operator: divides the artificial intelligence operation into a data fetch step, a multiplication step, an accumulation step, and a write memory step; (Whatmough, (col. 9 line 64 – col. 10 line 6), “Controller 210 is coupled to control port 206, multiplexor 220, input register 230 [wherein the operator: divides the artificial intelligence operation into a data fetch step], multi-stage add module 240 and output register 250. Generally, controller 210 manages the operation of pipelined accumulator 200 [a multiplication step, an accumulation step], including, for example, reset, input selection for multiplexor 220, etc. Controller 210 may be a simple microcontroller, a programmable circuit, etc. Input port 202 receives a sequence of operands to be accumulated, and output port 204 outputs the final sum of the accumulation calculation [and a write memory step].”) repeatedly performs the data fetch step, the multiplication step, and the accumulation step, (Whatmough, (col. 10 line 65 – 67]), “Generally, pipelined accumulator 200 executes a number of processing cycles that depends upon the number of operands that are to be accumulated [repeatedly performs the data fetch step, the multiplication step, and the accumulation step].”) and only after repeatedly performing the data fetch step, the multiplication step, and the accumulation step, perform the write memory step. (Whatmough, (col. 13 line 52 – 58]), “During the 10th cycle (output) [and only after repeatedly performing the data fetch step, the multiplication step, and the accumulation step], the partial result stored in pipeline register 242 (i.e., [6+9]) is presented to the input of 2nd stage adder 243, which latches the input data, converts the partial result into a complete result, and outputs the complete result (i.e., 15) to output register 250 for storage. Output register 250 then outputs this value to output port 204 [perform the write memory step].”) Whatmough, Agarwal, TAO and Temam are related to the same field of endeavor (i.e.: memory management and execution flow design). It would have been obvious to a person having ordinary skill in the art before the effective filing date of the invention to combine the teaching of Whatmough with teachings of Agarwal, TAO and Temam to add a pipelined accumulation techniques that enables efficient processing of sequential summation operations to reduce processing delays and improved utilization of resources, (Whatmough, Abstract). Claim 32 is rejected under 35 U.S.C. 103 as being unpatentable over Agarwal in view of Temam, TAO, Whatmough and in further view of Jain et al., Pub. No.: US20200201781A1. Agarwal in view of Temam, TAO and Whatmough teach the method of claim 23. Agarwal in view of Temam, TAO and Whatmough do not teach: wherein the data in the first domain and the data in the second domain differ from each other in one or more of bandwidth, dynamic range, resolution, refresh rate, and frame rate Jain teaches: wherein the data in the first domain and the data in the second domain differ from each other in one or more of bandwidth, dynamic range, resolution, refresh rate, and frame rate (Jain, “[0231] … Different data structures may be associated with different techniques for domain-based access [wherein the data in the first domain and the data in the second domain], such as assigning different degrees of multiplexing to different data structures, assigning different latency or bandwidth [differ from each other in one or more of bandwidth, dynamic range, resolution, refresh rate, and frame rate] targets to different data structures, assigning different data structures to certain domains 310 or certain quantities or distributions of domains 310, or other techniques. When the memory controller 870 identifies such a tag or metadata accompanying an access command, the memory controller 870 can make an association between the access command and the data structure, and determine a set of domains 310, or an order for accessing a set of domains 310, based on the access command being associated with the data structure…”) Jain, Agarwal, Temam, TAO and Whatmough are related to the same field of endeavor (i.e.: memory management and execution flow design). It would have been obvious to a person having ordinary skill in the art before the effective filing date of the invention to combine the teaching of Jain with teachings of Agarwal, Temam, TAO and Whatmough to dynamically tailor latency, bandwidth and domain allocation based on data type or priority to improve overall performance and efficiency. (Jain, ¶[0230] - [0232]). Allowable subject matter Claim 31 is objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. The prior art made of record does not teach, make obvious, or suggest the claim limitations as disclosed in applicant's claims. Claim 31 recites: wherein the operator includes a first operator that performs a first artificial intelligence operation and a second operator that performs a second artificial intelligence operation different from the first artificial intelligence operation, the first operator uses, for the neural network layer, a partial area of the first memory and a partial area of the second memory as a third space for storing data before the first artificial intelligence operation and a fourth space for storing data after the first artificial intelligence operation, respectively, and the second operator uses, for the neural network layer, another partial area of the first memory and another partial area of the second memory as a fifth space for storing data before the second artificial intelligence operation and a sixth space for storing data after the second artificial intelligence operation. The closest prior art(s): Vemuri et al., Pub. No.: US11704535B1. Vemuri describes an integrated circuit (IC) for processing and accelerating data passing through a neural network is disclosed, which is reconfigurable IC that includes a digital processing engine (DPE) array having a plurality of DPEs configured to execute one or more layers of a neural network. The reconfigurable IC also includes programmable logic, which includes: an IO controller coupled to input ping-pong buffers the IO controller receiving input data from an interconnect coupled to the programmable logic, wherein the input data fills a first input buffer of the input ping-pong buffers while data stored in a second input buffer of the input ping-pong buffers is processed; However, Vemuri does not teach the operator includes two different AI operators that perform distinct AI operations. For a neural network layer, each operator uses separate partial areas of two memories to store input data before its operation and output data after its operation, with the first operator using one set of memory area and the second operator using a different set. Temam et al., Pub. No.: US10504022B2 Temam describes an accelerator includes a computing unit. The computing unit includes: a first memory bank for storing input activations or output activations; a second memory bank for storing neural network parameters used in performing computations, the second memory bank configured to store a sufficient amount of the neural network parameters on the computing unit to allow for latency below a specified level with throughput above a specified level for a given NN model and architecture; at least one cell including at least one MAC operator that receives parameters from the second memory bank and performs computations; However, Temam does not teach the operator includes two different AI operators that perform distinct AI operations. For a neural network layer, each operator uses separate partial areas of two memories to store input data before its operation and output data after its operation, with the first operator using one set of memory area and the second operator using a different set. In summary, the references made of record, fail to disclose the required claimed technical features recited by claim 31 limitations as a whole. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Olson et al., Pub. No.: US20210303307A1, (2020). Olson describes operating an accumulation process in a data processing apparatus. The accumulation process comprises a plurality of accumulations which output a respective plurality of accumulated values, each based on a stored value and a computed value generated by a data processing operation. Any inquiry concerning this communication or earlier communications from the examiner should be directed to MATIYAS T MARU whose telephone number is (571)270-0902 or via email: matiyas.maru@uspto.gov. The examiner can normally be reached Monday 8:00am - Friday 4:00pm EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Michelle Bechtold can be reached on (571)431-0762. The fax phone number for the organization were this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /M.T.M./ Examiner, Art Unit 2148 /MICHELLE T BECHTOLD/Supervisory Patent Examiner, Art Unit 2148
Read full office action

Prosecution Timeline

Show 4 earlier events
Oct 08, 2025
Examiner Interview Summary
Nov 13, 2025
Response Filed
Jan 23, 2026
Final Rejection mailed — §103
Mar 20, 2026
Response after Non-Final Action
Apr 23, 2026
Request for Continued Examination
Apr 28, 2026
Response after Non-Final Action
Aug 03, 2026
Non-Final Rejection mailed — §103
Aug 05, 2026
Interview Requested

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12705482
POINT PROCESS LEARNING METHOD, POINT PROCESS LEARNING APPARATUS AND PROGRAM
3y 3m to grant Granted Aug 11, 2026
Patent 12694283
ACCELERATING INFERENCE OF NEURAL NETWORK MODELS VIA DYNAMIC EARLY EXITS
5y 2m to grant Granted Jul 28, 2026
Patent 12688427
TRAINING ACTION SELECTION NEURAL NETWORKS USING HINDSIGHT MODELLING
4y 3m to grant Granted Jul 21, 2026
Patent 12682221
Techniques For Increasing Activation Sparsity In Artificial Neural Networks
3y 9m to grant Granted Jul 14, 2026
Patent 12664440
SYSTEM AND METHOD FOR A LARGE CODEWORD MODEL FOR DEEP LEARNING
1y 8m to grant Granted Jun 23, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
65%
Grant Probability
70%
With Interview (+4.8%)
4y 2m (~5m remaining)
Median Time to Grant
High
PTA Risk
Based on 51 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month