Prosecution Insights
Last updated: October 01, 2026
Application No. 18/194,108

NEURAL NETWORK INCLUDING LOCAL STORAGE UNIT

Non-Final OA §103§112
Filed
Mar 31, 2023
Examiner
ABOUD, ABDULLAH KHALED
Art Unit
Tech Center
Assignee
STMicroelectronics N.V.
OA Round
1 (Non-Final)
Grant Probability
Favorable
1-2
OA Rounds

Examiner Intelligence

Grants only 0% of cases
0%
Career Allowance Rate
0 granted / 0 resolved
-60.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
Avg Prosecution
20 currently pending
Career history
14
Total Applications
across all art units
This examiner has no resolved cases yet (career too new); statute-level performance unavailable. The Grant Probability card shows Tech Center averages instead.

Office Action

§103 §112
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Specification The disclosure is objected to because of the following informalities: Page 5 line 11 refers to "The convolution units 106, the activation units 108, and the pooling units 110." Reference numeral 106 is established as "feature data," 108 as the "stream switch," and 110 as the "internal storage" (see pages [3]-[7]). These numerals cannot simultaneously designate different components. The correct numeral for these units may be 112 (hardware accelerators), or each unit may require its own distinct numeral. Appropriate correction is required. The informalities itemized in this Office action are exemplary only and do not constitute an exhaustive listing of the defects present in the specification. Applicant is required to review the entire specification in detail and make appropriate corrections to all informalities of the types identified herein, including but not limited to: (1) grammatical and typographical errors; (2) missing, duplicated, or garbled words and phrases; (3) erroneous, inconsistent, or duplicated reference numerals, ensuring that each reference character designates one and the same part throughout the description and drawings; (4) inconsistencies between the written description and the drawings, including flow-diagram step numbering; and (5) internally inconsistent descriptive passages. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claim 8, 9, and 15 rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Claim 8 and 9 recites the limitation "the read transformation unit" in in the second line of the claim. There is insufficient antecedent basis for this limitation in the claim. Claim 8 and 9 depends from claim 5, which depends from claims 4 and 1; no claim in this dependency chain previously recites "a read transformation unit." A read transformation unit is first introduced in claim 6, which is not in the dependency chain of claim 8 and 9. Claim 15 recites the term "a sufficient amount" which is a relative term that renders the claim indefinite, the term "sufficient" is not defined by the claim, the specification does not provide a standard for ascertaining the requisite degree, and one of ordinary skill in the art would not be reasonably apprised of the scope of the invention. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1-7, 9-14, and 16-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Munteanu et al. (US 20200126178 A1) in view of Liu et al. (US 20210390367 A1). As to claim 1, Munteanu teaches a method, comprising: receiving, at a neural network, feature data from a memory external to the neural network; (see Munteanu paragraph [0035] "The CNN Engine 30 connects to a system bus 42 and can access main (DRAM) memory 40 into which images acquired by the system are written.", and see Munteanu paragraph [0044] "In any case, it will be seen that the image cache 31 is initially loaded with an input image/map from DRAM 40.") passing the feature data to an internal storage of the neural network; (see Munteanu paragraph [0039] "An image cache 31 exposes a data input port (din) and a data output port (dout) to the remainder of the CNN engine 30. Data is read through the data input port (din) either from DRAM 40 via read controller 36 or from the output of a convolution engine 32 via a sub-sampler 34 to an address specified at an address port of the image cache by the controller 60.") passing the feature data from the internal storage to the first hardware accelerator; and (see Munteanu paragraph [0057] "A window of N×M pixel values for each input map is read from the image cache 31 in a given clock cycle, whereas the weights for an output map need only be provided once per output map.") generating first transformed feature data by processing the feature data with the first hardware accelerator. (see Munteanu paragraph [0057] "The convolution engine 32 performs a number of scalar (dot) products followed by an activation function to produce a pixel value.") Munteanu does not explicitly teach "storing, with a write transformation unit of the internal storage, the feature data in the internal storage with a first address configuration based on a first hardware accelerator that is next in a flow of the neural network;" However, Liu teaches storing, with a write transformation unit of the internal storage, the feature data in the internal storage with a first address configuration based on a first hardware accelerator that is next in a flow of the neural network; (see Liu paragraph [0078] "MEU 200 is coupled to memory 176 and PEs 180, and converts or expands the original version of the IFM data (activations, “A”) to an IM2COL version, and provides the expanded IFM data (expanded activation sequences, “A.sub.i”) to PEs 180, as discussed in detail below....", see Liu paragraph [0081] "Generally, the number of PEs 180 within each SSA 192 and the number of expanded activation sequences ("Ai") that are provided by each MEU 200 depend on the underlying matrix dimensions of the IFM data and the weights.", see Liu paragraph [0103] "Input data selector 220 receives activation data from memory 176, and, based on a control signal from controller 210, sends the activation data to registers 230-X.sub.1, 230-X.sub.2, 230-X.sub.3, 230-X.sub.4, or registers 240-Y.sub.1, 240-Y.sub.2, 240-Y.sub.3, 240-Y.sub.4.", and see Liu paragraph [0127] "During operation in state S1, first column data of IFM 83 are loaded into registers 230-X.sub.1, 230-X.sub.2, 230-X.sub.3, 230-X.sub.4 of register set 230. During operation in state S2, second column data of IFM 83 are loaded into registers 240-Y.sub.1, 240-Y.sub.2, 240-Y.sub.3 and 240-Y.sub.4 of register set 240.") Examiner Note: The input data selector 220 under controller 210 is a write transformation unit that stores each incoming data value at a register location selected by a state-machine-controlled pattern whose configuration depends on the dimensions of the downstream PE array (the first hardware accelerator next in the flow). It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the invention of Munteanu to include a write transformation unit that stores the feature data in the internal storage with an address configuration based on the hardware accelerator next in the flow, as taught by Liu, in order to reduce the memory bandwidth required for data movement, increase power efficiency at the system level, and arrange the data to take advantage of the compute regularity of the downstream accelerator, as suggested by Liu, see paragraph [0021]. As to claim 2, Munteanu as modified by Liu teaches the method of claim 1, comprising passing the first transformed feature data to the memory. (see Munteanu paragraph [0046] "Note that once feature extraction/classification is complete, any of the generated maps/vectors can be written back to DRAM 40 via a write controller 39.") As to claim 3, Munteanu as modified by Liu teaches the method of claim 2, comprising: receiving, at the neural network, the first transformed feature data from the memory; (see Liu paragraph [0073] "Generally, HA 170 receives input data from memory 130 over communication bus 110, and transmit output data to memory 130 over communication bus 110. In certain embodiments, the input data may be associated with a layer (or portion of a layer) of the ANN model, and the output data from that layer (or portion of that layer) may be transmitted to memory 130 over communication bus 110.", and see Liu paragraph [0077] "Memory 176 may send the OFM data back to memory 130 via communication bus interface 174, after which new IFM data and associated weights are received from memory 130 via communication bus interface 174.") passing the first transformed feature data to the internal storage; (see Liu paragraph [0077] "Memory 176 is coupled to communication bus interface 174, PEs 180 and MEU 200, and receives IFM data and weights from memory 130 via communication bus interface 174, provides the IFM data (activations, "A") to MEU 200, provides associated weights ("W") to PEs 180, and receives OFM data (dot products, "O") from PEs 180.") storing, with the write transformation unit of the internal storage, the transformed feature data in the internal storage with a second address configuration based on a second hardware accelerator that is next in the flow of the neural network; (see Liu paragraph [0085] "The process then repeats for the next sequences of weights (W), activations (A) and expanded activation sequences (A.sub.1, A.sub.2, A.sub.3 and A.sub.4).", see Liu paragraph [0103] "Input data selector 220 receives activation data from memory 176, and, based on a control signal from controller 210, sends the activation data to registers 230-X.sub.1, 230-X.sub.2, 230-X.sub.3, 230-X.sub.4, or registers 240-Y.sub.1, 240-Y.sub.2, 240-Y.sub.3, 240-Y.sub.4.", and see Liu paragraph [0124] "MEU state table 320 depicts the activation values a.sub.1 to a.sub.16 stored in each register for states S1 to S9 based on the embodiments of IFM 83 and converted input data matrix 86 depicted in FIGS. 4A and 4B. States S1 to S9 are followed by states S1 and S2 of the next IFM data sequence (i.e., activation values b.sub.1 to b.sub.16.") passing the transformed feature data from the internal storage to the second hardware accelerator; and (see Liu paragraph [0081] "In these embodiments, SA 190 includes additional PEs 180, such as, for example, 4 additional PEs 180, 12 additional PEs 180, 24 additional PEs 180, etc., which may form additional SSAs 192; each additional SSA 192 may be coupled to an additional MEU 200.", and see Liu paragraph [0083] "With respect to the activations, MEU 200 reads the elements of a sequence of activations (A) from memory 176, and generates four unique expanded activation sequences (A.sub.1, A.sub.2, A.sub.3 and A.sub.4), as described below. Each PE 180 receives one of the unique activation sequences (A.sub.1, A.sub.2, A.sub.3 or A.sub.4) from MEU 200.") generating second transformed feature data by processing the first transformed feature data with the second hardware accelerator. (see Liu paragraph [0084] "Generally, PE.sub.1, PE.sub.2, PE.sub.3 and PE.sub.4 calculate and then output respective dot products O.sub.1, O.sub.2, O.sub.3 and O.sub.4 based on the weights (W) and the respective sequences of expanded activations (A.sub.1, A.sub.2, A.sub.3 and A.sub.4).") As to claim 4, Munteanu as modified by Liu teaches the method of claim 1, wherein the first hardware accelerator is a convolution accelerator. (see Munteanu paragraph [0059] "As indicated above, especially during feature extraction, the convolution engine 32 can process windows of N×M pixels provided by the image cache 31 each clock cycle.") As to claim 5, Munteanu as modified by Liu teaches the method of claim 4, wherein the first address configuration is based on a kernel size associated with the convolution accelerator. (see Liu paragraph [0081] "Generally, the number of PEs 180 within each SSA 192 and the number of expanded activation sequences (“Ai”) that are provided by each MEU 200 depend on the underlying matrix dimensions of the IFM data and the weights. For example, the embodiment depicted in FIG. 7 reflects the matrix dimensions of filter 82 (3×3), IFM 83 (4×4) and OFM 84 (2×2) depicted in FIG. 4A.", and see Liu paragraph [0109] "State machine transition diagram 300 depicts nine unique states, i.e., S1, S2, S3, S4, S5, S6, S7, S8 and S9. Each state is associated with a particular processing cycle described above, i.e., S1 is associated with the first processing cycle, S2 is associated with the second processing cycle, etc.") As to claim 6, Munteanu as modified by Liu teaches the method of claim 4, wherein passing the transformed feature data includes reading, with a read transformation unit the feature data from the internal storage with a read address pattern based on the first hardware accelerator. (see Munteanu paragraph [0060] "the image cache 31 is arranged to accelerate processing by producing rectangular windows of N×M pixels for use within the convolution engine 32 in as few clock cycles as possible and preferably in a single clock cycle.", see Munteanu paragraph [0072] "Calculates input data de-multiplexer and output data multiplexer selection signals (MS,BS);", see Munteanu paragraph [0073] "Transforms the x, y, w and h inputs into Address (ADDR) and Byte (Pixel) Enable (BE) control signals for each of the four SRAMs.", and see Munteanu paragraph [0075] "The blocks 64, 66 route each input/output pixels data to and from the correct SRAM and to the correct pixel location within each SRAM data in and data out ports.") As to claim 7, Munteanu as modified by Liu teaches the method of claim 6, wherein the read address patter is based on a kernel size associated with the convolution accelerator. (see Munteanu paragraph [0061] "A typical size for the convolution kernels for embedded applications is 5×5, but it will be appreciated that this may vary. Embodiments of the present invention can operate with kernels of any size up to 5×5 operating on windows located at any (x, y) location in the image cache 31", and see Munteanu paragraph [0066] "For embedded applications, 5×5 convolution kernels fit very well with the maximum window size limit of the cache.") As to claim 9, Munteanu as modified by Liu teaches the method of claim 5, comprising controlling the read transformation unit with a control unit of the internal storage. (see Munteanu paragraph [0068] "The internal structure of the cache is presented in FIG. 6.", see Munteanu paragraph [0072] "Calculates input data de-multiplexer and output data multiplexer selection signals (MS,BS);", and see Munteanu paragraph [0075] "The blocks 64, 66 route each input/output pixels data to and from the correct SRAM and to the correct pixel location within each SRAM data in and data out ports.") As to claim 10, Munteanu as modified by Liu teaches the method of claim 1, comprising controlling the write transformation unit with configuration registers of the neural network. (see Munteanu paragraph [0036] "In either case, the configurable succession of feature extraction and classification operations to be performed by the CNN Engine 30 is determined by the CPU 50 by setting registers within the controller 60 via the system bus 42.", and see Munteanu paragraph [0042] "In each case, a start location, comprising the base address and extent of shifting, the offset, of an image/map, map or vector within the image cache 31 is determined by the controller 60 according to the configuration received from the CPU 50.") As to claim 11, Munteanu teaches a method, comprising: receiving, at a neural network, feature data from a memory external to the neural network; (see Munteanu paragraph [0035] "The CNN Engine 30 connects to a system bus 42 and can access main (DRAM) memory 40 into which images acquired by the system are written.", and see Munteanu paragraph [0044] "In any case, it will be seen that the image cache 31 is initially loaded with an input image/map from DRAM 40.") passing the feature data to an internal storage of the neural network; (see Munteanu paragraph [0039] "Data is read through the data input port (din) either from DRAM 40 via read controller 36 or from the output of a convolution engine 32 via a sub-sampler 34 to an address specified at an address port of the image cache by the controller 60.") storing the feature data in the internal storage; (see Munteanu paragraph [0044] "the image cache 31 is initially loaded with an input image/map from DRAM 40. Then all processing can be performed only using this image cache 31 with no need to access the external DRAM for image information until classification is complete.") generating first transformed feature data by processing the feature data with the first hardware accelerator. (see Munteanu paragraph [0057] "The convolution engine 32 performs a number of scalar (dot) products followed by an activation function to produce a pixel value.") Munteanu does not explicitly teach "reading, with a read transformation unit of the neural network, the feature data to a first hardware accelerator with a read address pattern based on an operation of the first hardware accelerator" However, Liu teaches reading, with a read transformation unit of the neural network, the feature data to a first hardware accelerator with a read address pattern based on an operation of the first hardware accelerator; (see Liu paragraph [0083] "With respect to the activations, MEU 200 reads the elements of a sequence of activations (A) from memory 176, and generates four unique expanded activation sequences (A.sub.1, A.sub.2, A.sub.3 and A.sub.4), as described below. Each PE 180 receives one of the unique activation sequences (A.sub.1, A.sub.2, A.sub.3 or A.sub.4) from MEU 200.", see Liu paragraph [0088] "The sequence of weights provided by memory 176 (w.sub.1 to w.sub.9) is based on converted weight matrix 85 (one weight per cycle, read left to right), and the sequences of expanded activations provided by MEU 200 (A.sub.1, A.sub.2, A.sub.3 and A.sub.4) are based on converted input data matrix 86 (one row per cycle, each row read left to right), as depicted in FIG. 4B.", and see Liu paragraph [0106] "Output data selector 250 receives expanded activation data from register set 230 and register set 240, and sends the expanded activation data to a plurality of PEs 180 based on a control signal from controller 210. The expanded activation data are sent in a rowwise format.") It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the invention of Munteanu to include a read transformation unit that reads the feature data to the hardware accelerator with a read address pattern based on the operation of the accelerator, as taught by Liu, in order to provide the data to the accelerator in the order its computation consumes it, thereby reducing memory bandwidth, increasing power efficiency, and exploiting the compute regularity of matrix multiplication, as suggested by Liu, see paragraph [0021]. As to claim 12, Munteanu as modified by Liu teaches the method of claim 11, wherein the first hardware accelerator is a convolution accelerator. (see Munteanu paragraph [0059] "As indicated above, especially during feature extraction, the convolution engine 32 can process windows of N×M pixels provided by the image cache 31 each clock cycle.") As to claim 13, Munteanu as modified by Liu teaches the method of claim 12, wherein the read address pattern is based on a kernel size of the first hardware accelerator. (see Liu paragraph [0081] "For example, the embodiment depicted in FIG. 7 reflects the matrix dimensions of filter 82 (3×3), IFM 83 (4×4) and OFM 84 (2×2) depicted in FIG. 4A.", and see Liu paragraph [0109] "State machine transition diagram 300 depicts nine unique states, i.e., S1, S2, S3, S4, S5, S6, S7, S8 and S9. Each state is associated with a particular processing cycle described above, i.e., S1 is associated with the first processing cycle, S2 is associated with the second processing cycle, etc.") As to claim 14, Munteanu as modified by Liu teaches the method of claim 13, wherein the feature data is a feature tensor. (see Munteanu paragraph [0040] "After the first convolution and subsampling layer of feature extraction, a number of maps, in this case 5, Layer 1 Map 0 . . . Layer 1 Map 4, are generated.", and see Munteanu paragraph [0059] "In order to produce one output pixel in a given output map, the convolution engine 32 needs: one clock cycle for 2D convolution; or a number of clock cycles equal to the number of input maps for 3D convolutions.") Examiner Note: Feature data comprising an express plurality of maps processed by 3D convolution across the input maps; data having width, height, and depth is a feature tensor under the broadest reasonable interpretation. As to claim 16, Munteanu as modified by Liu teaches the method of claim 11, comprising passing the first transformed feature data to the memory. (see Munteanu paragraph [0046] "Note that once feature extraction/classification is complete, any of the generated maps/vectors can be written back to DRAM 40 via a write controller 39.") As to claim 17, Munteanu as modified by Liu teaches the method of claim 16, comprising: receiving, at the neural network, the first transformed feature data from the memory; (see Liu paragraph [0073] "In certain embodiments, the input data may be associated with a layer (or portion of a layer) of the ANN model, and the output data from that layer (or portion of that layer) may be transmitted to memory 130 over communication bus 110.", and see Liu paragraph [0077] "Memory 176 may send the OFM data back to memory 130 via communication bus interface 174, after which new IFM data and associated weights are received from memory 130 via communication bus interface 174.") passing the first transformed feature data to the internal storage; (see Liu paragraph [0077] "Memory 176 is coupled to communication bus interface 174, PEs 180 and MEU 200, and receives IFM data and weights from memory 130 via communication bus interface 174, provides the IFM data (activations, "A") to MEU 200, provides associated weights ("W") to PEs 180, and receives OFM data (dot products, "O") from PEs 180.") storing, with a write transformation unit of the internal storage, the transformed feature data in the internal storage with an address configuration based on a second hardware accelerator that is next in a flow of the neural network; (see Liu paragraph [0085] "The process then repeats for the next sequences of weights (W), activations (A) and expanded activation sequences (A.sub.1, A.sub.2, A.sub.3 and A.sub.4).", see Liu paragraph [0103] "Input data selector 220 receives activation data from memory 176, and, based on a control signal from controller 210, sends the activation data to registers 230-X.sub.1, 230-X.sub.2, 230-X.sub.3, 230-X.sub.4, or registers 240-Y.sub.1, 240-Y.sub.2, 240-Y.sub.3, 240-Y.sub.4.", and see Liu paragraph [0124] "States S1 to S9 are followed by states S1 and S2 of the next IFM data sequence (i.e., activation values b.sub.1 to b.sub.16).") passing the first transformed feature data from the internal storage to the second hardware accelerator; and (see Liu paragraph [0081] "In these embodiments, SA 190 includes additional PEs 180, such as, for example, 4 additional PEs 180, 12 additional PEs 180, 24 additional PEs 180, etc., which may form additional SSAs 192; each additional SSA 192 may be coupled to an additional MEU 200.", and see Liu paragraph [0083] "Each PE 180 receives one of the unique activation sequences (A.sub.1, A.sub.2, A.sub.3 or A.sub.4) from MEU 200.") generating second transformed feature data by processing the first transformed feature data with the second hardware accelerator. (see Liu paragraph [0084] "Generally, PE.sub.1, PE.sub.2, PE.sub.3 and PE.sub.4 calculate and then output respective dot products O.sub.1, O.sub.2, O.sub.3 and O.sub.4 based on the weights (W) and the respective sequences of expanded activations (A.sub.1, A.sub.2, A.sub.3 and A.sub.4).") As to claim 18, Munteanu teaches a device comprising a neural network, the neural network including: a stream engine configured to receive feature data from a memory external to the neural network; (see Munteanu paragraph [0035] "FIG. 3 shows a block diagram of a CNN Engine 30 implemented according to an embodiment of the present invention within an image acquisition system. The CNN Engine 30 connects to a system bus 42 and can access main (DRAM) memory 40 into which images acquired by the system are written.", and see Munteanu paragraph [0039] "Data is read through the data input port (din) either from DRAM 40 via read controller 36 or from the output of a convolution engine 32 via a sub-sampler 34 to an address specified at an address port of the image cache by the controller 60.") a hardware accelerator; and (see Munteanu paragraph [0059] "As indicated above, especially during feature extraction, the convolution engine 32 can process windows of N×M pixels provided by the image cache 31 each clock cycle.") an internal storage configured to receive the feature data from the stream engine, the internal storage including: (see Munteanu paragraph [0039] "An image cache 31 exposes a data input port (din) and a data output port (dout) to the remainder of the CNN engine 30. Data is read through the data input port (din) either from DRAM 40 via read controller 36 or from the output of a convolution engine 32 via a sub-sampler 34 to an address specified at an address port of the image cache by the controller 60.") the hardware accelerator configured to receive the feature data from the internal storage and to process the feature data to generate transformed feature data. (see Munteanu paragraph [0057] "A window of N×M pixel values for each input map is read from the image cache 31 in a given clock cycle, whereas the weights for an output map need only be provided once per output map. The convolution engine 32 performs a number of scalar (dot) products followed by an activation function to produce a pixel value.") Munteanu does not explicitly teach "a write transformation unit configured to write the feature data into the internal storage with a write address pattern based on a configuration of the hardware accelerator; and", and "a read transformation unit configured to read the feature data to the hardware accelerator with a read address pattern based on the configuration of the hardware accelerator" However, Liu teaches a write transformation unit configured to write the feature data into the internal storage with a write address pattern based on a configuration of the hardware accelerator; and (see Liu paragraph [0078] "MEU 200 is coupled to memory 176 and PEs 180, and converts or expands the original version of the IFM data (activations, "A") to an IM2COL version, and provides the expanded IFM data (expanded activation sequences, "A.sub.i") to PEs 180, as discussed in detail below.", see Liu paragraph [0081] "Generally, the number of PEs 180 within each SSA 192 and the number of expanded activation sequences ("Ai") that are provided by each MEU 200 depend on the underlying matrix dimensions of the IFM data and the weights.", see Liu paragraph [0103] "Input data selector 220 receives activation data from memory 176, and, based on a control signal from controller 210, sends the activation data to registers 230-X.sub.1, 230-X.sub.2, 230-X.sub.3, 230-X.sub.4, or registers 240-Y.sub.1, 240-Y.sub.2, 240-Y.sub.3, 240-Y.sub.4.", and see Liu paragraph [0127] "During operation in state S1, first column data of IFM 83 are loaded into registers 230-X.sub.1, 230-X.sub.2, 230-X.sub.3, 230-X.sub.4 of register set 230. During operation in state S2, second column data of IFM 83 are loaded into registers 240-Y.sub.1, 240-Y.sub.2, 240-Y.sub.3 and 240-Y.sub.4 of register set 240.") a read transformation unit configured to read the feature data to the hardware accelerator with a read address pattern based on the configuration of the hardware accelerator, (see Liu paragraph [0083] "With respect to the activations, MEU 200 reads the elements of a sequence of activations (A) from memory 176, and generates four unique expanded activation sequences (A.sub.1, A.sub.2, A.sub.3 and A.sub.4), as described below. Each PE 180 receives one of the unique activation sequences (A.sub.1, A.sub.2, A.sub.3 or A.sub.4) from MEU 200.", see Liu paragraph [0088] "The sequence of weights provided by memory 176 (w.sub.1 to w.sub.9) is based on converted weight matrix 85 (one weight per cycle, read left to right), and the sequences of expanded activations provided by MEU 200 (A.sub.1, A.sub.2, A.sub.3 and A.sub.4) are based on converted input data matrix 86 (one row per cycle, each row read left to right), as depicted in FIG. 4B.", and see Liu paragraph [0106] "Output data selector 250 receives expanded activation data from register set 230 and register set 240, and sends the expanded activation data to a plurality of PEs 180 based on a control signal from controller 210. The expanded activation data are sent in a rowwise format.") It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the invention of Munteanu to include write and read transformation units that write and read the feature data with address patterns based on the configuration of the hardware accelerator, as taught by Liu, in order to convert the data inline between the memory and the accelerator, reducing memory footprint and bandwidth and increasing power efficiency at the system level, as suggested by Liu, see paragraph [0021]. As to claim 19, Munteanu as modified by Liu teaches the device of claim 18, wherein the internal storage includes a control unit configured to control the read transformation unit. (see Munteanu paragraph [0068] "The internal structure of the cache is presented in FIG. 6.", see Munteanu paragraph [0072] "Calculates input data de-multiplexer and output data multiplexer selection signals (MS,BS);", and see Munteanu paragraph [0075] "The blocks 64, 66 route each input/output pixels data to and from the correct SRAM and to the correct pixel location within each SRAM data in and data out ports.") As to claim 20, Munteanu as modified by Liu teaches the device of claim 18, wherein the neural network is a convolutional neural network. (see Munteanu paragraph [0035] "FIG. 3 shows a block diagram of a CNN Engine 30 implemented according to an embodiment of the present invention within an image acquisition system.") Claim(s) 8 is/are rejected under 35 U.S.C. 103 as being unpatentable over Munteanu et al. (US 20200126178 A1) in view of Liu et al. (US 20210390367 A1) and Tuominen et al. (US 20090031089 A1). As to claim 8, Munteanu as modified by Liu teaches the method of claim 5, comprising: receiving the feature data in rows with the internal storage from the memory; and (see Munteanu paragraph [0044] "the image cache 31 is initially loaded with an input image/map from DRAM 40.", and see Munteanu paragraph [0053] "For example, if the system bus 42 comprises 64 bits, 8 pixels/weights/cells could be transferred across the bus 42 in one clock cycle. Thus, an 8×1 pixel window, set of weights or vector could be read/written from or into the caches 31 or 37 in one transaction.") Munteanu-Liu does not explicitly teach "reading the data from internal storage in columns with the read transformation unit" However, Tuominen teaches reading the data from internal storage in columns with the read transformation unit. (see Tuominen paragraph [0061] "A read address logic 240, which receives a common read counter value 322 from the common counter logic 320 of the transpose memory (TRAM) 105, controls the data output of the DPRAM component 230.", and see Tuominen paragraph [0097] "This means that in the first even matrix read cycle associated with the read counter value q'=0 of the even matrix reading stage, the data words representing the matrix elements ye;00 to ye;70 should have to be outputted by the transpose memory (TRAM) 105.") It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the invention of Munteanu as modified by Liu to read the row-stored data out of the internal storage in columns, as taught by Tuominen, in order to reorder the data for the downstream processing unit without any dead cycles or wait cycles, thereby maximizing the throughput of the storage, as suggested by Tuominen, see paragraph [0188]. Claim(s) 15 is/are rejected under 35 U.S.C. 103 as being unpatentable over Munteanu et al. (US 20200126178 A1) in view of Liu et al. (US 20210390367 A1) and Feinberg et al. (US 20190286975 A1). As to claim 15, Munteanu as modified by Liu teaches the method of claim 11, comprising: Munteanu-Liu does not explicitly teach "determining, with a control unit of the internal storage, whether a sufficient amount of the feature data has been received at the internal storage; and", and "reading the feature data from the internal storage to the first hardware accelerator only if the control unit determines that a sufficient amount of the feature data has been received at the internal storage" However, Feinberg teaches determining, with a control unit of the internal storage, whether a sufficient amount of the feature data has been received at the internal storage; and (see Feinberg paragraph [0073] "A controller (depicted as controller 2202 in FIG. 22 and controller 3006 in FIG. 30, but not depicted in other figures for conciseness of presentation) may be responsible for powering on and off convolver units. The controller may power on a row of convolver units once the data from row n of a horizontal stripe has been loaded into the data storage elements corresponding to the row of convolver units.") reading the feature data from the internal storage to the first hardware accelerator only if the control unit determines that a sufficient amount of the feature data has been received at the internal storage. (see Feinberg paragraph [0072] "Upon row n of horizontal stripe 902a being loaded into the second row of data storage elements (i.e., d.sub.2,1, d.sub.2,2, d.sub.2,3 and d.sub.2,4), the first row of convolver units (i.e., CU.sub.1,1, CU.sub.1,2, CU.sub.1,3 and CU.sub.1,4) corresponding to the second row of data storage elements may be activated.", and see Feinberg paragraph [0073] "In one embodiment, "active" means that a convolver unit is powered on, whereas "non-active" means that a convolver unit is powered off to save power.") It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the invention of Munteanu as modified by Liu to determine, with the control unit of the internal storage, whether a sufficient amount of feature data has been received before reading the data to the hardware accelerator, as taught by Feinberg, in order to keep the accelerator inactive until the required data has been loaded, thereby saving power, as suggested by Feinberg, see paragraph [0073]. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to ABDULLAH K ABOUD whose telephone number is (571)272-0025. The examiner can normally be reached Mon-Fri 8am-5pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Li B Zhen, can be reached at (571) 272-3768. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /ABDULLAH KHALED ABOUD/Examiner, Art Unit 2121 /Li B. Zhen/Supervisory Patent Examiner, Art Unit 2121
Read full office action

Prosecution Timeline

Mar 31, 2023
Application Filed
Aug 13, 2026
Non-Final Rejection mailed — §103, §112 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
Grant Probability
Low
PTA Risk
Based on 0 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month