Detailed Action
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
This office action is non-final and is in response to claims filed on 03/07/2026. Claims 1-10 are pending examination.
Information Disclosure Statement
The Information Disclosure Statement (IDS) submitted on 11/24/2023 is in compliance with the provisions of 37 CFR 1.97, 1.98, and MPEP § 609. It has been placed in the application file, and the information referred to therein has been considered as to the merits.
Claim Interpretation
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked.
As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph:
(A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function;
(B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and
(C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function.
Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function.
Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function.
Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action.
This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitation(s) is/are: a data buffer unit, a pooling unit, a loss computing unit, in claim 1.
Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof.
If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph.
In paragraph [0018], the specification recites “The accelerator 100 is a computing in Memory (CIM) accelerator for optimizing the latency of accessing data or moving data when the neural network is performed. The accelerator 100 can be applied to any neural network operation, such as an incremental learning operation. Further, the accelerator 100 can be implemented by an electronic system…The data buffer unit 11 is coupled to the memory 10 for buffering data outputted from the memory…The pooling unit 12 is coupled to the memory 10 for pooling data outputted from the memory 10 for acquiring a maximum pooling value… The loss computing unit 13 is coupled to the memory 10 for computing output loss”.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 8-10 rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claim 8 recites the limitation "a fifth input terminal configured to receive the second activation information”. There is insufficient antecedent basis for this limitation in the claim. For examination purposes, examiner has interpreted “a fifth input terminal configured to receive the second activation information” as “a fifth input terminal configured to receive
Claim 9 recites “the output terminal of the third macro unit is configured to output an output vector having (N+1) dimensions,”. It is unclear the output terminal is the first or second output terminal of the third macro unit. For examination purposes, examiner has interpreted “the output terminal of the third macro unit is configured to output an output vector having (N+1) dimensions,” as “the first output terminal of the third macro unit is configured to output an output vector having (N+1) dimensions,”
Claim 10 is rejected for being dependent on an above rejected claim.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim 1 is rejected as being unpatentable over Guan et al. (US 20240104360 A1) hereinafter Guan in view of Sather et al. (US 12579430 B1) hereinafter Sather.
With regards to claim 1, Guan teaches A computing in memory accelerator for applying to a neural network comprising: a memory configured to save data; (Guan [0025]: Referring to FIG. 4, a system configured for near memory computation, in accordance with aspects of the present technology, is shown… The one or more memory units 410 can include a controller 425 and a plurality of memory devices 430, 435)
a data buffer unit coupled to the memory and configured to buffer data outputted from the memory; (Guan [0025]: The controller 425 can include control logic 440, a mode register 445, a plurality of computation units 450-455, a read data buffer (RDB) 460)
a pooling unit coupled to the memory and configured to pool data for acquiring a maximum pooling value; (Guan [0036]: the memory unit can perform partial computations including one or more aggregation, combination and or the like functions on the accessed data; Guan [0039]: a computation can be the mean/max pooling aggregator)
a first macro circuit coupled to the data buffer unit; (Guan [0025]: a plurality of computation units 450-455; Guan Fig. 4: shows multiple computation units connected to the data buffer)
a second macro circuit coupled to the data buffer unit; (Guan [0025]: a plurality of computation units 450-455; Guan Fig. 4: shows multiple computation units connected to the data buffer)
a third macro circuit coupled to the data buffer unit [and the loss computing unit;] (Guan [0025]: a plurality of computation units 450-455; Guan Fig. 4: shows multiple computation units connected to the data buffer)
[and a multiplexer] coupled to the pooling unit, the first macro circuit, the second macro circuit, and the third macro circuit and configured to generate output data; (Guan [0025]: a plurality of computation units 450-455; Guan Fig. 4: shows multiple computation units; Guan [0036]: the memory unit can perform partial computations including one or more aggregation, combination and or the like functions on the accessed data; Guan [0039]: a computation can be the mean/max pooling aggregator; Guan [0030]: The result data of the optional aggregation, combination or the like function can be sent by the one or more memory units 410 as return data, at 760. In addition, when the memory access does not also include optional aggregation, combination or the like instructions and parameters, the accessed data can be returned by the one or more memory units as return data, at 760)
wherein an output terminal of [the multiplexer] is coupled to an input terminal of the memory (Guan [0030]: The result data of the optional aggregation, combination or the like function can be sent by the one or more memory units 410 as return data, at 760. In addition, when the memory access does not also include optional aggregation, combination or the like instructions and parameters, the accessed data can be returned by the one or more memory units as return data, at 760).
Guan fails to teach a loss computing unit coupled to the memory and configured to compute output loss; [a third macro circuit coupled to the data buffer unit and] the loss computing unit; and a multiplexer [coupled to the pooling unit, the first macro circuit, the second macro circuit, and the third macro circuit and configured to generate output data;] [wherein an output terminal of] the multiplexer [is coupled to an input terminal of the memory].
However, Sather teaches a loss computing unit coupled to the memory and configured to compute output loss; (Sather Column 24 Lines 57-58: an error calculator and propagator 910, a loss function generator 915)
[a third macro circuit coupled to the data buffer unit and] the loss computing unit; (Sather Column 24 Lines 57-58: an error calculator and propagator 910, a loss function generator 915)
and a multiplexer [coupled to the pooling unit, the first macro circuit, the second macro circuit, and the third macro circuit and configured to generate output data;] (Sather Column 49 Lines 18-23: The multiplexer output is the neural network node output, which is gated by a valid signal (not shown) to indicate when the post-processing unit is outputting a completed activation value to be carried by the activation write bus to the appropriate core and stored in the activation memory of that core)
[wherein an output terminal of] the multiplexer [is coupled to an input terminal of the memory] (Sather Column 49 Lines 18-23: The multiplexer output is the neural network node output, which is gated by a valid signal (not shown) to indicate when the post-processing unit is outputting a completed activation value to be carried by the activation write bus to the appropriate core and stored in the activation memory of that core).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teaching of Guan with the loss unit and multiplexer as taught by Sather. One of ordinary skill in the art would be motivated to make this combination because it would allow the system to operate to compute neural network operations in an efficient, low-power manner, according to the configuration data provided by the control circuits as taught by Sather (Sather Columns 42-43 Lines 67-2). Also, this would increase the flexibility of the system as it would allow for different outputs depending on the needs of the system.
Claim 2 is rejected as being unpatentable over Guan in view of Sather further in view of Gurram et al. (US 20210349717 A1) hereinafter Gurram.
With regards to claim 2, Guan in view of Sather teaches all of the limitations of claim 1 above. Guan further teaches wherein the first macro circuit comprises: a first macro unit comprising: a [first] input terminal configured to receive address information; (Guan [0027]: The read memory access can be extended to a read with compute (read_w_comp) access, and the write memory access can be extended to write with compute (write_w_comp) access. The extensions can include data address, data count, and data stride extension. The extensions can be embedded into a coherent interconnect (GenZ/CXL) data package, extended double data rate (DDR) command or the like; Guan Fig. 4: shows a databus that sends the control information to the units)
a [second] input terminal configured to receive read/write control information; (Guan [0027]: The read memory access can be extended to a read with compute (read_w_comp) access, and the write memory access can be extended to write with compute (write_w_comp) access. The extensions can include data address, data count, and data stride extension. The extensions can be embedded into a coherent interconnect (GenZ/CXL) data package, extended double data rate (DDR) command or the like; Guan Fig. 4: shows a databus that sends the control information to the units)
a [third] input terminal configured to receive [input feature information;] (Guan [0030]: The read data buffer (RDB) 445 and write data buffer (WDB) 460 can be multi-entry buffers used to buffer data for the computation units 440-445. The modes can include no computation, complete computation and partial computation modes. In the no computation mode, the read data buffer (RDB) 445 and write data buffer (WDB) 460 can be bypassed. In the complete computation mode, the computation units 440-445 can perform all the computations on the accessed data. In the partial computation mode, the computation units 440-445 can perform a portion of the computations on the accessed data and a partial result can be passed as the data for further computations by the compute engine 415 of the central core 405. The result data of the optional aggregation, combination or the like function can be sent by the one or more memory units 410 as return data, at 760)
a [fourth] input terminal configured to receive [activation information;] (Guan [0030]: The read data buffer (RDB) 445 and write data buffer (WDB) 460 can be multi-entry buffers used to buffer data for the computation units 440-445. The modes can include no computation, complete computation and partial computation modes. In the no computation mode, the read data buffer (RDB) 445 and write data buffer (WDB) 460 can be bypassed. In the complete computation mode, the computation units 440-445 can perform all the computations on the accessed data. In the partial computation mode, the computation units 440-445 can perform a portion of the computations on the accessed data and a partial result can be passed as the data for further computations by the compute engine 415 of the central core 405. The result data of the optional aggregation, combination or the like function can be sent by the one or more memory units 410 as return data, at 760)
a [fifth] input terminal configured to receive [weighting information;] (Guan [0030]: The read data buffer (RDB) 445 and write data buffer (WDB) 460 can be multi-entry buffers used to buffer data for the computation units 440-445. The modes can include no computation, complete computation and partial computation modes. In the no computation mode, the read data buffer (RDB) 445 and write data buffer (WDB) 460 can be bypassed. In the complete computation mode, the computation units 440-445 can perform all the computations on the accessed data. In the partial computation mode, the computation units 440-445 can perform a portion of the computations on the accessed data and a partial result can be passed as the data for further computations by the compute engine 415 of the central core 405. The result data of the optional aggregation, combination or the like function can be sent by the one or more memory units 410 as return data, at 760)
and an output terminal; (Guan [0007]: The controller can compute the aggregation, combination and or other similar operations on the attributes base on the first memory access to generate result data. The controller can output the result data based on the first memory accesses)
and a calculating activation unit comprising: a [first] input terminal [coupled] to the output terminal of the first macro unit; (Guan [0025]: a plurality of computation units 450-455; Guan Fig. 4: shows multiple computation units)
a [second] input terminal configured to receive calculating [activation] mode information; (Guan [0030]: The read data buffer (RDB) 445 and write data buffer (WDB) 460 can be multi-entry buffers used to buffer data for the computation units 440-445. The modes can include no computation, complete computation and partial computation modes. In the no computation mode, the read data buffer (RDB) 445 and write data buffer (WDB) 460 can be bypassed. In the complete computation mode, the computation units 440-445 can perform all the computations on the accessed data. In the partial computation mode, the computation units 440-445 can perform a portion of the computations on the accessed data and a partial result can be passed as the data for further computations by the compute engine 415 of the central core 405. The result data of the optional aggregation, combination or the like function can be sent by the one or more memory units 410 as return data, at 760)
and an output terminal; (Guan [0007]: The controller can compute the aggregation, combination and or other similar operations on the attributes base on the first memory access to generate result data. The controller can output the result data based on the first memory accesses).
Guan fails to teach [a] third [input terminal configured to receive] input feature information; [a] fourth [input terminal configured to receive] activation information; [a] fifth [input terminal configured to receive] weighting information; [a] second [input terminal configured to receive calculating] activation [mode information;].
However, Sather teaches [a third input terminal configured to receive] input feature information; (Sather Column 7 Lines 42-46: The convolutional layers of some embodiments use a small kernel (e.g., 2×2, 3×3, 5×5, etc.) to process blocks of input values (output values from a previous layer) in a set of two-dimensional grids (e.g., channels of pixels of an image, input feature maps) with the same set of parameters)
[a fourth input terminal configured to receive] activation information; (Sather Column 8 Lines 8-12: The input to each computation node is a subset of the input activation values, and the dot product for the computation node involves multiplying those input activation values by the weights)
[a fifth input terminal configured to receive] weighting information; (Sather Column 8 Lines 8-12: The input to each computation node is a subset of the input activation values, and the dot product for the computation node involves multiplying those input activation values by the weights)
[a second input terminal configured to receive calculating] activation [mode information;] (Sather Column 8 Lines 8-12: The input to each computation node is a subset of the input activation values, and the dot product for the computation node involves multiplying those input activation values by the weights).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teaching of Guan in view of Sather with the input feature information, activations, and weights as taught by Sather. One of ordinary skill in the art would be motivated to make this combination because it would allow the system to operate to compute neural network operations in an efficient, low-power manner, according to the configuration data provided by the control circuits as taught by Sather (Sather Columns 42-43 Lines 67-2).
While Guan in view of Sather teaches of input terminals, they fail to specifically teach of the first, second, third, fourth, and fifth input terminals of the first macro unit, the first and second input terminals of the calculating activation unit, and [a first input terminal] coupled [to the output terminal].
However, Gurram teaches the first, second, third, fourth, and fifth input terminals of the first macro unit (Gurram [0207]: The components of the hardware system 1800 can be included in any ALU, FPU, system of ALUs, or processing elements described herein. The system 1800 includes input hardware 1802; Gurram Fig. 18: Shows multiple inputs)
the first and second input terminals of the calculating activation unit (Gurram [0207]: The components of the hardware system 1800 can be included in any ALU, FPU, system of ALUs, or processing elements described herein. The system 1800 includes input hardware 1802; Gurram Fig. 18: Shows multiple inputs)
[a] first [input terminal] coupled [to the output terminal] (Gurram [0098]: A compute engine cluster 332 can include a set of compute engine tiles 340A-340D... can also be interconnected via a set of tile interconnects 323A-323F; Gurram Fig. 3C: shows the compute tiles connected together; Gurram [0207]: The components of the hardware system 1800 can be included in any ALU, FPU, system of ALUs, or processing elements described herein. The system 1800 includes input hardware 1802; Gurram Fig. 18: Shows multiple inputs).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teaching of Guan in view of Sather with the input terminals and the units being coupled together as taught by Gurram. One of ordinary skill in the art would be motivated to make this combination because multi-threaded operation enables an efficient execution environment in the face of higher latency memory accesses as taught by Gurram (Gurram [0110]).
Claims 3-4 are rejected as being unpatentable over Guan in view of Sather further in view of Gurram further in view of Chilwal et al. (US 10769329 B1) hereinafter Chilwal.
With regards to claim 3, Guan in view of Sather further in view of Gurram teaches all of the limitations of claim 1 above. Guan further teaches the output terminal of the first macro unit is configured to output [an output vector having (N+1) dimensions,] (Guan [0007]: The controller can compute the aggregation, combination and or other similar operations on the attributes base on the first memory access to generate result data. The controller can output the result data based on the first memory accesses)
the first macro unit further comprises [a clock control terminal and a reset terminal,] (Guan [0025]: a plurality of computation units 450-455; Guan Fig. 4: shows multiple computation units).
Guan fails to teach wherein the input feature information comprises an input feature vector having (M+1) dimensions, [the output terminal of the first macro unit is configured to output] an output vector having (N+1) dimensions, and M and N are two positive integers.
However, Sather teaches wherein the input feature information comprises an input feature vector having (M+1) dimensions, (Sather Columns 7-8 Lines 67-2: The array can be conceptualized as a set of two-dimensional grids, also referred to as input feature maps or input channels for the layer)
[the output terminal of the first macro unit is configured to output] an output vector having (N+1) dimensions, (Sather Column 8 Lines 8-12 and 30-36: The input to each computation node is a subset of the input activation values, and the dot product for the computation node involves multiplying those input activation values by the weights... To generate the output activations, each of the filters 205 is applied to numerous subsets of the input activation values 200. Specifically, in a typical convolution layer, each 3×3×3 filter is moved across the three-dimensional array of activation values, and the dot product between the 27 activations in the current subset and the 27 weight values in the filter is computed)
and M and N are two positive integers (Sather Column 8 Lines 8-12 and 30-36: The input to each computation node is a subset of the input activation values, and the dot product for the computation node involves multiplying those input activation values by the weights... To generate the output activations, each of the filters 205 is applied to numerous subsets of the input activation values 200. Specifically, in a typical convolution layer, each 3×3×3 filter is moved across the three-dimensional array of activation values, and the dot product between the 27 activations in the current subset and the 27 weight values in the filter is computed; Sather Fig. 2: shows the inputs being a 6x6 matrix and the outputs being a 4x4 matrix).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teaching of Guan in view of Sather further in view of Gurram with the vectors and dimensions as taught by Sather. One of ordinary skill in the art would be motivated to make this combination because it would allow the system to operate to compute neural network operations in an efficient, low-power manner, according to the configuration data provided by the control circuits as taught by Sather (Sather Columns 42-43 Lines 67-2).
Guan in view of Sather fails to teach [the first macro unit further comprises] a clock control terminal and a reset terminal.
However, Chilwal teaches [the first macro unit further comprises] a clock control terminal and a reset terminal, (Chilwal Column 4 Lines 24-29: a signal routing circuit that is reconfigurable, by way of asserting corresponding combinations of the signal path control signals, to generate multiple alternative data, clock and set/reset signal paths between the two flip-flop/latch elements and generic input output nodes of the retention model).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teaching of Guan in view of Sather further in view of Gurram with the clock control and reset terminals as taught by Chilwal. One of ordinary skill in the art would be motivated to make this combination because it would facilitate efficient processing as taught by Chilwal (Chilwal Column 5 Lines 12-13). Also, this would increase the reliability of the system as the circuits would be timed by a clock, ensuring the data is moved in a synchronized way.
With regards to claim 4, Guan in view of Sather further in view of Gurram further in view of Chilwal teaches all of the limitations of claim 1 above. Guan further teaches wherein the first macro unit generates [(M+1)x (N+1) weightings after the weighting information is received,] (Guan [0025]: a plurality of computation units 450-455; Guan Fig. 4: shows multiple computation units)
an output of the first macro unit is generated (Guan [0007]: The controller can compute the aggregation, combination and or other similar operations on the attributes base on the first memory access to generate result data. The controller can output the result data based on the first memory accesses).
Guan fails to teach [wherein the first macro unit generates] (M+1)x (N+1) weightings after the weighting information is received, and after (M+1) weightings of each column are linearly combined with the input feature vector having (M+1) dimensions.
However, Sather teaches (M+1)x (N+1) weightings after the weighting information is received, (Sather Column 8 Lines 8-13: The input to each computation node is a subset of the input activation values, and the dot product for the computation node involves multiplying those input activation values by the weights that make up one of the filters of the layer. As shown, in this example the layer includes six filters 205, each of which is 3×3×3)
and after (M+1) weightings of each column are linearly combined with the input feature vector having (M+1) dimensions (Sather Column 8 Lines 8-12 and 30-36: The input to each computation node is a subset of the input activation values, and the dot product for the computation node involves multiplying those input activation values by the weights... To generate the output activations, each of the filters 205 is applied to numerous subsets of the input activation values 200. Specifically, in a typical convolution layer, each 3×3×3 filter is moved across the three-dimensional array of activation values, and the dot product between the 27 activations in the current subset and the 27 weight values in the filter is computed).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teaching of Guan in view of Sather further in view of Gurram further in view of Chilwal with the weightings as taught by Sather. One of ordinary skill in the art would be motivated to make this combination because it would allow the system to operate to compute neural network operations in an efficient, low-power manner, according to the configuration data provided by the control circuits as taught by Sather (Sather Columns 42-43 Lines 67-2).
Claims 5 and 8 are rejected as being unpatentable over Guan in view of Sather further in view of Gurram further in view of Moon et al. (US 20240346312 A1) hereinafter Moon further in view of Vilcu et al. (US 20240046069 A1) hereinafter Vilcu.
With regards to claim 5, Guan in view of Sather teaches all of the limitations of claim 1 above. Guan further teaches wherein the second macro circuit comprises: a second macro unit comprising: a [first] input terminal configured to receive address information; (Guan [0027]: The read memory access can be extended to a read with compute (read_w_comp) access, and the write memory access can be extended to write with compute (write_w_comp) access. The extensions can include data address, data count, and data stride extension. The extensions can be embedded into a coherent interconnect (GenZ/CXL) data package, extended double data rate (DDR) command or the like; Guan Fig. 4: shows a databus that sends the control information to the units)
a [second] input terminal configured to receive read/write control information; (Guan [0027]: The read memory access can be extended to a read with compute (read_w_comp) access, and the write memory access can be extended to write with compute (write_w_comp) access. The extensions can include data address, data count, and data stride extension. The extensions can be embedded into a coherent interconnect (GenZ/CXL) data package, extended double data rate (DDR) command or the like; Guan Fig. 4: shows a databus that sends the control information to the units)
a [third] input terminal configured to receive [input feature information;] (Guan [0030]: The read data buffer (RDB) 445 and write data buffer (WDB) 460 can be multi-entry buffers used to buffer data for the computation units 440-445. The modes can include no computation, complete computation and partial computation modes. In the no computation mode, the read data buffer (RDB) 445 and write data buffer (WDB) 460 can be bypassed. In the complete computation mode, the computation units 440-445 can perform all the computations on the accessed data. In the partial computation mode, the computation units 440-445 can perform a portion of the computations on the accessed data and a partial result can be passed as the data for further computations by the compute engine 415 of the central core 405. The result data of the optional aggregation, combination or the like function can be sent by the one or more memory units 410 as return data, at 760)
a [fourth] input terminal configured to receive [activation information;] (Guan [0030]: The read data buffer (RDB) 445 and write data buffer (WDB) 460 can be multi-entry buffers used to buffer data for the computation units 440-445. The modes can include no computation, complete computation and partial computation modes. In the no computation mode, the read data buffer (RDB) 445 and write data buffer (WDB) 460 can be bypassed. In the complete computation mode, the computation units 440-445 can perform all the computations on the accessed data. In the partial computation mode, the computation units 440-445 can perform a portion of the computations on the accessed data and a partial result can be passed as the data for further computations by the compute engine 415 of the central core 405. The result data of the optional aggregation, combination or the like function can be sent by the one or more memory units 410 as return data, at 760)
a [fifth] input terminal configured to receive [weighting update information;] (Guan [0030]: The read data buffer (RDB) 445 and write data buffer (WDB) 460 can be multi-entry buffers used to buffer data for the computation units 440-445. The modes can include no computation, complete computation and partial computation modes. In the no computation mode, the read data buffer (RDB) 445 and write data buffer (WDB) 460 can be bypassed. In the complete computation mode, the computation units 440-445 can perform all the computations on the accessed data. In the partial computation mode, the computation units 440-445 can perform a portion of the computations on the accessed data and a partial result can be passed as the data for further computations by the compute engine 415 of the central core 405. The result data of the optional aggregation, combination or the like function can be sent by the one or more memory units 410 as return data, at 760)
a [sixth] input terminal configured to receive [weighting information;] (Guan [0030]: The read data buffer (RDB) 445 and write data buffer (WDB) 460 can be multi-entry buffers used to buffer data for the computation units 440-445. The modes can include no computation, complete computation and partial computation modes. In the no computation mode, the read data buffer (RDB) 445 and write data buffer (WDB) 460 can be bypassed. In the complete computation mode, the computation units 440-445 can perform all the computations on the accessed data. In the partial computation mode, the computation units 440-445 can perform a portion of the computations on the accessed data and a partial result can be passed as the data for further computations by the compute engine 415 of the central core 405. The result data of the optional aggregation, combination or the like function can be sent by the one or more memory units 410 as return data, at 760)
and an output terminal; (Guan [0007]: The controller can compute the aggregation, combination and or other similar operations on the attributes base on the first memory access to generate result data. The controller can output the result data based on the first memory accesses)
and a calculating activation unit comprising: a [first] input terminal [coupled] to the output terminal of the second macro unit; (Guan [0025]: a plurality of computation units 450-455; Guan Fig. 4: shows multiple computation units)
a [second] input terminal configured to receive calculating [activation] mode information; (Guan [0030]: The read data buffer (RDB) 445 and write data buffer (WDB) 460 can be multi-entry buffers used to buffer data for the computation units 440-445. The modes can include no computation, complete computation and partial computation modes. In the no computation mode, the read data buffer (RDB) 445 and write data buffer (WDB) 460 can be bypassed. In the complete computation mode, the computation units 440-445 can perform all the computations on the accessed data. In the partial computation mode, the computation units 440-445 can perform a portion of the computations on the accessed data and a partial result can be passed as the data for further computations by the compute engine 415 of the central core 405. The result data of the optional aggregation, combination or the like function can be sent by the one or more memory units 410 as return data, at 760)
a [third] input terminal configured to receive [first variation information;] (Guan [0030]: The read data buffer (RDB) 445 and write data buffer (WDB) 460 can be multi-entry buffers used to buffer data for the computation units 440-445. The modes can include no computation, complete computation and partial computation modes. In the no computation mode, the read data buffer (RDB) 445 and write data buffer (WDB) 460 can be bypassed. In the complete computation mode, the computation units 440-445 can perform all the computations on the accessed data. In the partial computation mode, the computation units 440-445 can perform a portion of the computations on the accessed data and a partial result can be passed as the data for further computations by the compute engine 415 of the central core 405. The result data of the optional aggregation, combination or the like function can be sent by the one or more memory units 410 as return data, at 760)
a first output terminal; (Guan [0007]: The controller can compute the aggregation, combination and or other similar operations on the attributes base on the first memory access to generate result data. The controller can output the result data based on the first memory accesses)
and a [second] output terminal [configured to output second variation information;] (Guan [0007]: The controller can compute the aggregation, combination and or other similar operations on the attributes base on the first memory access to generate result data. The controller can output the result data based on the first memory accesses)
a weighting gradient calculation unit comprising: a [first] input terminal [coupled] to the [second] output terminal of the calculating activation and derivative unit; (Guan [0025]: a plurality of computation units 450-455; Guan Fig. 4: shows multiple computation units)
a [second] input terminal configured to receive [input feature information;] (Guan [0030]: The read data buffer (RDB) 445 and write data buffer (WDB) 460 can be multi-entry buffers used to buffer data for the computation units 440-445. The modes can include no computation, complete computation and partial computation modes. In the no computation mode, the read data buffer (RDB) 445 and write data buffer (WDB) 460 can be bypassed. In the complete computation mode, the computation units 440-445 can perform all the computations on the accessed data. In the partial computation mode, the computation units 440-445 can perform a portion of the computations on the accessed data and a partial result can be passed as the data for further computations by the compute engine 415 of the central core 405. The result data of the optional aggregation, combination or the like function can be sent by the one or more memory units 410 as return data, at 760)
a [third] input terminal configured to receive an output control signal; (Guan [0008]: The control logic can be configured to receive a memory access including an aggregation and or combination operation, and access and compute the attributes based on the operation included in the memory access. The control logic of the controller can configure one or more of the plurality of computation units of the controller to compute the aggregation or combination operation on the attributes, based on the operation of the memory access, to generate result data. The control logic of the controller can then output the result data; Guan [0027]: The read memory access can be extended to a read with compute (read_w_comp) access, and the write memory access can be extended to write with compute (write_w_comp) access. The extensions can include data address, data count, and data stride extension. The extensions can be embedded into a coherent interconnect (GenZ/CXL) data package, extended double data rate (DDR) command or the like; Guan Fig. 4: shows a databus that sends the control information to the units)
a [fourth] input terminal configured to receive a computing control signal; (Guan [0008]: The control logic can be configured to receive a memory access including an aggregation and or combination operation, and access and compute the attributes based on the operation included in the memory access. The control logic of the controller can configure one or more of the plurality of computation units of the controller to compute the aggregation or combination operation on the attributes, based on the operation of the memory access, to generate result data. The control logic of the controller can then output the result data; Guan [0027]: The read memory access can be extended to a read with compute (read_w_comp) access, and the write memory access can be extended to write with compute (write_w_comp) access. The extensions can include data address, data count, and data stride extension. The extensions can be embedded into a coherent interconnect (GenZ/CXL) data package, extended double data rate (DDR) command or the like; Guan Fig. 4: shows a databus that sends the control information to the units)
and an output terminal configured to output [third variation information;] (Guan [0007]: The controller can compute the aggregation, combination and or other similar operations on the attributes base on the first memory access to generate result data. The controller can output the result data based on the first memory accesses)
[a second input terminal coupled to] the output terminal of the weighting gradient calculation unit; (Guan [0007]: The controller can compute the aggregation, combination and or other similar operations on the attributes base on the first memory access to generate result data. The controller can output the result data based on the first memory accesses)
[an output terminal coupled to] the [third] input terminal of the second macro unit; (Guan [0025]: a plurality of computation units 450-455; Guan Fig. 4: shows multiple computation units).
Guan fails to teach [a] third [input terminal configured to receive] input feature information; [a] fourth [input terminal configured to receive] activation information; [a] sixth [input terminal configured to receive] weighting information; [a] second [input terminal configured to receive calculating] activation [mode information;] [a] second [input terminal configured to receive calculating] activation [mode information;] and an input multiplexer comprising: a first input terminal configured to receive the input feature information; a second input terminal coupled to the output terminal of [the weighting gradient calculation unit;] an output terminal [coupled to the third input terminal of the second macro unit;] and a control terminal configured to receive a selection signal.
However, Sather teaches [a third input terminal configured to receive] input feature information; (Sather Column 7 Lines 42-46: The convolutional layers of some embodiments use a small kernel (e.g., 2×2, 3×3, 5×5, etc.) to process blocks of input values (output values from a previous layer) in a set of two-dimensional grids (e.g., channels of pixels of an image, input feature maps) with the same set of parameters)
[a fourth input terminal configured to receive] activation information; (Sather Column 8 Lines 8-12: The input to each computation node is a subset of the input activation values, and the dot product for the computation node involves multiplying those input activation values by the weights)
[a sixth input terminal configured to receive] weighting information; (Sather Column 8 Lines 8-12: The input to each computation node is a subset of the input activation values, and the dot product for the computation node involves multiplying those input activation values by the weights)
[a second input terminal configured to receive calculating] activation [mode information;] (Sather Column 8 Lines 8-12: The input to each computation node is a subset of the input activation values, and the dot product for the computation node involves multiplying those input activation values by the weights)
[a second input terminal configured to receive calculating] activation [mode information;] (Sather Column 8 Lines 8-12: The input to each computation node is a subset of the input activation values, and the dot product for the computation node involves multiplying those input activation values by the weights).
and an input multiplexer comprising: a first input terminal configured to receive the input feature information; (Sather Column 48 Lines 7-8: sent to a multiplexer 1515, and a set of configuration bits is used to select between these two possible inputs; The convolutional layers of some embodiments use a small kernel (e.g., 2×2, 3×3, 5×5, etc.) to process blocks of input values (output values from a previous layer) in a set of two-dimensional grids (e.g., channels of pixels of an image; Sather Column 7 Lines 42-46: input feature maps) with the same set of parameters)
a second input terminal coupled to the output terminal of [the weighting gradient calculation unit;] (Sather Column 48 Lines 7-8: sent to a multiplexer 1515, and a set of configuration bits is used to select between these two possible inputs)
an output terminal [coupled to the third input terminal of the second macro unit;] (Sather Column 48 Lines 7-8: sent to a multiplexer 1515, and a set of configuration bits is used to select between these two possible inputs)
and a control terminal configured to receive a selection signal (Sather Column 48 Lines 7-8: sent to a multiplexer 1515, and a set of configuration bits is used to select between these two possible inputs).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teaching of Guan in view of Sather with the input feature information, activations, weights, and multiplexer as taught by Sather. One of ordinary skill in the art would be motivated to make this combination because it would allow the system to operate to compute neural network operations in an efficient, low-power manner, according to the configuration data provided by the control circuits as taught by Sather (Sather Columns 42-43 Lines 67-2). Also, this would increase the flexibility of the system as it would allow for different outputs depending on the needs of the system.
While Guan in view of Sather teaches of input terminals, they fail to specifically teach of the first, second, third, fourth, fifth, and sixth input terminals of the second macro unit, the first, second, and third input terminals and second output terminal of the calculating activation and derivative unit, [a first input terminal] coupled [to the output terminal], the first, second, third, and fourth input terminals of the weighting gradient calculation unit, [a] first [input terminal] coupled [to the] second [output terminal].
However, Gurram teaches the first, second, third, fourth, fifth, and sixth input terminals of the second macro unit (Gurram [0207]: The components of the hardware system 1800 can be included in any ALU, FPU, system of ALUs, or processing elements described herein. The system 1800 includes input hardware 1802; Gurram Fig. 18: Shows multiple inputs)
The first, second, and third input terminals and second output terminal of the calculating activation and derivative unit (Gurram [0207]: he components of the hardware system 1800 can be included in any ALU, FPU, system of ALUs, or processing elements described herein. The system 1800 includes input hardware 1802... and output hardware 1810; Gurram Fig. 18: Shows multiple inputs and multiple outputs)
[a] first [input terminal] coupled [to the output terminal] (Gurram [0098]: A compute engine cluster 332 can include a set of compute engine tiles 340A-340D... can also be interconnected via a set of tile interconnects 323A-323F; Gurram Fig. 3C: shows the compute tiles connected together; Gurram [0207]: The components of the hardware system 1800 can be included in any ALU, FPU, system of ALUs, or processing elements described herein. The system 1800 includes input hardware 1802; Gurram Fig. 18: Shows multiple inputs)
the first, second, third, and fourth input terminals of the weighting gradient calculation unit (Gurram [0207]: The components of the hardware system 1800 can be included in any ALU, FPU, system of ALUs, or processing elements described herein. The system 1800 includes input hardware 1802; Gurram Fig. 18: Shows multiple inputs)
[a] first [input terminal] coupled [to the] second [output terminal] (Gurram [0207]: he components of the hardware system 1800 can be included in any ALU, FPU, system of ALUs, or processing elements described herein. The system 1800 includes input hardware 1802... and output hardware 1810; Gurram Fig. 18: Shows multiple inputs and multiple outputs).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teaching of Guan in view of Sather with the input and output terminals and the units being coupled together as taught by Gurram. One of ordinary skill in the art would be motivated to make this combination because multi-threaded operation enables an efficient execution environment in the face of higher latency memory accesses as taught by Gurram (Gurram [0110]).
Guan in view of Sather fails to teach [a fifth input terminal configured to receive] weighting update information; [a third input terminal configured to receive] first variation information; [and a second output terminal configured to output] second variation information.
However, Moon teaches [a fifth input terminal configured to receive] weighting update information; (Moon [0053]: the gradient to be used for updating the weights)
[a third input terminal configured to receive] first variation information; (Moon [0053]: the gradient to be used for updating the weights)
[and a second output terminal configured to output] second variation information; (Moon [0053]: the derivative of the activation function of each layer).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teaching of Guan in view of Sather with the weighting update information and first and second variation information as taught by Moon. One of ordinary skill in the art would be motivated to make this combination because accordingly, the training of the neural network model for personalization can be efficiently performed, in particular on-device, without overhead as taught by Moon (Moon [0101]).
Guan in view of Sather further in view of Moon fails to teach [and an output terminal configured to output] third variation information.
However, Vilcu teaches [and an output terminal configured to output] third variation information; (Vilcu [0127]: The partial differential of the sum of weighted activations with respect to the weight).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teaching of Guan in view of Sather further in view of Moon with the third variation information as taught by Vilcu. One of ordinary skill in the art would be motivated to make this combination because this method processes batches of training data and losses to efficiently train a neural network as taught by Vilcu (Vilcu [0133]).
With regards to claim 8, Guan in view of Sather teaches all of the limitations of claim 1 above. Guan further teaches wherein the third macro circuit comprises: a third macro unit comprising: a [first] input terminal configured to receive address information; (Guan [0027]: The read memory access can be extended to a read with compute (read_w_comp) access, and the write memory access can be extended to write with compute (write_w_comp) access. The extensions can include data address, data count, and data stride extension. The extensions can be embedded into a coherent interconnect (GenZ/CXL) data package, extended double data rate (DDR) command or the like; Guan Fig. 4: shows a databus that sends the control information to the units)
a [second] input terminal configured to receive read/write control information; (Guan [0027]: The read memory access can be extended to a read with compute (read_w_comp) access, and the write memory access can be extended to write with compute (write_w_comp) access. The extensions can include data address, data count, and data stride extension. The extensions can be embedded into a coherent interconnect (GenZ/CXL) data package, extended double data rate (DDR) command or the like; Guan Fig. 4: shows a databus that sends the control information to the units)
a [third] input terminal configured to receive [input feature information;] (Guan [0030]: The read data buffer (RDB) 445 and write data buffer (WDB) 460 can be multi-entry buffers used to buffer data for the computation units 440-445. The modes can include no computation, complete computation and partial computation modes. In the no computation mode, the read data buffer (RDB) 445 and write data buffer (WDB) 460 can be bypassed. In the complete computation mode, the computation units 440-445 can perform all the computations on the accessed data. In the partial computation mode, the computation units 440-445 can perform a portion of the computations on the accessed data and a partial result can be passed as the data for further computations by the compute engine 415 of the central core 405. The result data of the optional aggregation, combination or the like function can be sent by the one or more memory units 410 as return data, at 760)
a [fourth] input terminal configured to receive [first activation information;] (Guan [0030]: The read data buffer (RDB) 445 and write data buffer (WDB) 460 can be multi-entry buffers used to buffer data for the computation units 440-445. The modes can include no computation, complete computation and partial computation modes. In the no computation mode, the read data buffer (RDB) 445 and write data buffer (WDB) 460 can be bypassed. In the complete computation mode, the computation units 440-445 can perform all the computations on the accessed data. In the partial computation mode, the computation units 440-445 can perform a portion of the computations on the accessed data and a partial result can be passed as the data for further computations by the compute engine 415 of the central core 405. The result data of the optional aggregation, combination or the like function can be sent by the one or more memory units 410 as return data, at 760)
a [fifth] input terminal configured to receive [the second activation information;] (Guan [0030]: The read data buffer (RDB) 445 and write data buffer (WDB) 460 can be multi-entry buffers used to buffer data for the computation units 440-445. The modes can include no computation, complete computation and partial computation modes. In the no computation mode, the read data buffer (RDB) 445 and write data buffer (WDB) 460 can be bypassed. In the complete computation mode, the computation units 440-445 can perform all the computations on the accessed data. In the partial computation mode, the computation units 440-445 can perform a portion of the computations on the accessed data and a partial result can be passed as the data for further computations by the compute engine 415 of the central core 405. The result data of the optional aggregation, combination or the like function can be sent by the one or more memory units 410 as return data, at 760)
a [sixth] input terminal; (Guan [0030]: The read data buffer (RDB) 445 and write data buffer (WDB) 460 can be multi-entry buffers used to buffer data for the computation units 440-445. The modes can include no computation, complete computation and partial computation modes. In the no computation mode, the read data buffer (RDB) 445 and write data buffer (WDB) 460 can be bypassed. In the complete computation mode, the computation units 440-445 can perform all the computations on the accessed data. In the partial computation mode, the computation units 440-445 can perform a portion of the computations on the accessed data and a partial result can be passed as the data for further computations by the compute engine 415 of the central core 405. The result data of the optional aggregation, combination or the like function can be sent by the one or more memory units 410 as return data, at 760)
a [seventh] input terminal configured to receive [weighting update information;] (Guan [0030]: The read data buffer (RDB) 445 and write data buffer (WDB) 460 can be multi-entry buffers used to buffer data for the computation units 440-445. The modes can include no computation, complete computation and partial computation modes. In the no computation mode, the read data buffer (RDB) 445 and write data buffer (WDB) 460 can be bypassed. In the complete computation mode, the computation units 440-445 can perform all the computations on the accessed data. In the partial computation mode, the computation units 440-445 can perform a portion of the computations on the accessed data and a partial result can be passed as the data for further computations by the compute engine 415 of the central core 405. The result data of the optional aggregation, combination or the like function can be sent by the one or more memory units 410 as return data, at 760)
an [eighth] input terminal configured to receive [weighting information;] (Guan [0030]: The read data buffer (RDB) 445 and write data buffer (WDB) 460 can be multi-entry buffers used to buffer data for the computation units 440-445. The modes can include no computation, complete computation and partial computation modes. In the no computation mode, the read data buffer (RDB) 445 and write data buffer (WDB) 460 can be bypassed. In the complete computation mode, the computation units 440-445 can perform all the computations on the accessed data. In the partial computation mode, the computation units 440-445 can perform a portion of the computations on the accessed data and a partial result can be passed as the data for further computations by the compute engine 415 of the central core 405. The result data of the optional aggregation, combination or the like function can be sent by the one or more memory units 410 as return data, at 760)
a first output terminal configured to output [first variation information;] (Guan [0007]: The controller can compute the aggregation, combination and or other similar operations on the attributes base on the first memory access to generate result data. The controller can output the result data based on the first memory accesses)
and a [second] output terminal; (Guan [0007]: The controller can compute the aggregation, combination and or other similar operations on the attributes base on the first memory access to generate result data. The controller can output the result data based on the first memory accesses)
and a calculating activation unit comprising: a [first] input terminal [coupled] to the [second] output terminal of the third macro unit; (Guan [0025]: a plurality of computation units 450-455; Guan Fig. 4: shows multiple computation units)
a [second] input terminal configured to receive calculating [activation] mode information; (Guan [0030]: The read data buffer (RDB) 445 and write data buffer (WDB) 460 can be multi-entry buffers used to buffer data for the computation units 440-445. The modes can include no computation, complete computation and partial computation modes. In the no computation mode, the read data buffer (RDB) 445 and write data buffer (WDB) 460 can be bypassed. In the complete computation mode, the computation units 440-445 can perform all the computations on the accessed data. In the partial computation mode, the computation units 440-445 can perform a portion of the computations on the accessed data and a partial result can be passed as the data for further computations by the compute engine 415 of the central core 405. The result data of the optional aggregation, combination or the like function can be sent by the one or more memory units 410 as return data, at 760)
a [third] input terminal configured to receive [first variation information;] (Guan [0030]: The read data buffer (RDB) 445 and write data buffer (WDB) 460 can be multi-entry buffers used to buffer data for the computation units 440-445. The modes can include no computation, complete computation and partial computation modes. In the no computation mode, the read data buffer (RDB) 445 and write data buffer (WDB) 460 can be bypassed. In the complete computation mode, the computation units 440-445 can perform all the computations on the accessed data. In the partial computation mode, the computation units 440-445 can perform a portion of the computations on the accessed data and a partial result can be passed as the data for further computations by the compute engine 415 of the central core 405. The result data of the optional aggregation, combination or the like function can be sent by the one or more memory units 410 as return data, at 760)
a first output terminal; (Guan [0007]: The controller can compute the aggregation, combination and or other similar operations on the attributes base on the first memory access to generate result data. The controller can output the result data based on the first memory accesses)
and a [second] output terminal [configured to output second variation information;] (Guan [0007]: The controller can compute the aggregation, combination and or other similar operations on the attributes base on the first memory access to generate result data. The controller can output the result data based on the first memory accesses)
a weighting gradient calculation unit comprising: a [first] input terminal [coupled] to the [second] output terminal of the calculating activation and derivative unit; (Guan [0025]: a plurality of computation units 450-455; Guan Fig. 4: shows multiple computation units)
a [second] input terminal configured to receive [input feature information;] (Guan [0030]: The read data buffer (RDB) 445 and write data buffer (WDB) 460 can be multi-entry buffers used to buffer data for the computation units 440-445. The modes can include no computation, complete computation and partial computation modes. In the no computation mode, the read data buffer (RDB) 445 and write data buffer (WDB) 460 can be bypassed. In the complete computation mode, the computation units 440-445 can perform all the computations on the accessed data. In the partial computation mode, the computation units 440-445 can perform a portion of the computations on the accessed data and a partial result can be passed as the data for further computations by the compute engine 415 of the central core 405. The result data of the optional aggregation, combination or the like function can be sent by the one or more memory units 410 as return data, at 760)
a [third] input terminal configured to receive an output control signal; (Guan [0008]: The control logic can be configured to receive a memory access including an aggregation and or combination operation, and access and compute the attributes based on the operation included in the memory access. The control logic of the controller can configure one or more of the plurality of computation units of the controller to compute the aggregation or combination operation on the attributes, based on the operation of the memory access, to generate result data. The control logic of the controller can then output the result data; Guan [0027]: The read memory access can be extended to a read with compute (read_w_comp) access, and the write memory access can be extended to write with compute (write_w_comp) access. The extensions can include data address, data count, and data stride extension. The extensions can be embedded into a coherent interconnect (GenZ/CXL) data package, extended double data rate (DDR) command or the like; Guan Fig. 4: shows a databus that sends the control information to the units)
a [fourth] input terminal configured to receive a computing control signal; (Guan [0008]: The control logic can be configured to receive a memory access including an aggregation and or combination operation, and access and compute the attributes based on the operation included in the memory access. The control logic of the controller can configure one or more of the plurality of computation units of the controller to compute the aggregation or combination operation on the attributes, based on the operation of the memory access, to generate result data. The control logic of the controller can then output the result data; Guan [0027]: The read memory access can be extended to a read with compute (read_w_comp) access, and the write memory access can be extended to write with compute (write_w_comp) access. The extensions can include data address, data count, and data stride extension. The extensions can be embedded into a coherent interconnect (GenZ/CXL) data package, extended double data rate (DDR) command or the like; Guan Fig. 4: shows a databus that sends the control information to the units)
and an output terminal configured to output [third variation information;] (Guan [0007]: The controller can compute the aggregation, combination and or other similar operations on the attributes base on the first memory access to generate result data. The controller can output the result data based on the first memory accesses)
[a second input terminal coupled to] the output terminal of the weighting gradient calculation unit; (Guan [0007]: The controller can compute the aggregation, combination and or other similar operations on the attributes base on the first memory access to generate result data. The controller can output the result data based on the first memory accesses)
[an output terminal coupled to] the [third] input terminal of the second macro unit; (Guan [0025]: a plurality of computation units 450-455; Guan Fig. 4: shows multiple computation units)
[and a derivative input multiplexer comprising: a first input terminal configured to receive second variation information] outputted from the loss computing unit; (Guan [0007]: The controller can compute the aggregation, combination and or other similar operations on the attributes base on the first memory access to generate result data. The controller can output the result data based on the first memory accesses)
[a second input terminal coupled to the second] output terminal of the calculating activation and derivative unit; (Guan [0007]: The controller can compute the aggregation, combination and or other similar operations on the attributes base on the first memory access to generate result data. The controller can output the result data based on the first memory accesses)
[and an output terminal coupled to the sixth] input terminal of the third macro unit (Guan [0025]: a plurality of computation units 450-455; Guan Fig. 4: shows multiple computation units).
Guan fails to teach [a] third [input terminal configured to receive] input feature information; [a] fourth [input terminal configured to receive] first activation information; [a] fifth [input terminal configured to receive] the second activation information; [a] eighth [input terminal configured to receive] weighting information; [a] second [input terminal configured to receive calculating] activation [mode information;] [a] second [input terminal configured to receive calculating] activation [mode information;] an input multiplexer comprising: a first input terminal configured to receive the input feature information; a second input terminal coupled to the output terminal of [the weighting gradient calculation unit;] an output terminal [coupled to the third input terminal of the third macro unit;] and a control terminal configured to receive a selection signal and a derivative input multiplexer comprising: a first input terminal configured to receive second variation information [outputted from the loss computing unit;] a second input terminal coupled to the second [output terminal of the calculating activation and derivative unit;] a control terminal configured to receive a selection signal; and an output terminal coupled to the sixth [input terminal of the third macro unit].
However, Sather teaches [a third input terminal configured to receive] input feature information; (Sather Column 7 Lines 42-46: The convolutional layers of some embodiments use a small kernel (e.g., 2×2, 3×3, 5×5, etc.) to process blocks of input values (output values from a previous layer) in a set of two-dimensional grids (e.g., channels of pixels of an image, input feature maps) with the same set of parameters)
[a fourth input terminal configured to receive] first activation information; (Sather Column 8 Lines 8-12: The input to each computation node is a subset of the input activation values, and the dot product for the computation node involves multiplying those input activation values by the weights)
[a fifth input terminal configured to receive] the second activation information; (Sather Column 8 Lines 8-12: The input to each computation node is a subset of the input activation values, and the dot product for the computation node involves multiplying those input activation values by the weights; Sather Fig. 2: shows multiple activation matrices)
[an eighth input terminal configured to receive] weighting information; (Sather Column 8 Lines 8-12: The input to each computation node is a subset of the input activation values, and the dot product for the computation node involves multiplying those input activation values by the weights)
[a second input terminal configured to receive calculating] activation [mode information;] (Sather Column 8 Lines 8-12: The input to each computation node is a subset of the input activation values, and the dot product for the computation node involves multiplying those input activation values by the weights)
[a second input terminal configured to receive calculating] activation [mode information;] (Sather Column 8 Lines 8-12: The input to each computation node is a subset of the input activation values, and the dot product for the computation node involves multiplying those input activation values by the weights).
an input multiplexer comprising: a first input terminal configured to receive the input feature information; (Sather Column 48 Lines 7-8: sent to a multiplexer 1515, and a set of configuration bits is used to select between these two possible inputs; The convolutional layers of some embodiments use a small kernel (e.g., 2×2, 3×3, 5×5, etc.) to process blocks of input values (output values from a previous layer) in a set of two-dimensional grids (e.g., channels of pixels of an image; Sather Column 7 Lines 42-46: input feature maps) with the same set of parameters)
a second input terminal coupled to the output terminal of [the weighting gradient calculation unit;] (Sather Column 48 Lines 7-8: sent to a multiplexer 1515, and a set of configuration bits is used to select between these two possible inputs)
an output terminal [coupled to the third input terminal of the third macro unit;] (Sather Column 48 Lines 7-8: sent to a multiplexer 1515, and a set of configuration bits is used to select between these two possible inputs)
and a control terminal configured to receive a selection signal (Sather Column 48 Lines 7-8: sent to a multiplexer 1515, and a set of configuration bits is used to select between these two possible inputs)
and a derivative input multiplexer comprising: a first input terminal configured to receive [second variation information outputted from the loss computing unit;] (Sather Column 51 Lines 51-53: The multiplexers 1710 each have eight inputs and receive a set of select bits (the weight selector input) from the core controller)
a second input terminal coupled to the [second output terminal of the calculating activation and derivative unit;] (Sather Column 51 Lines 51-53: The multiplexers 1710 each have eight inputs and receive a set of select bits (the weight selector input) from the core controller)
a control terminal configured to receive a selection signal; (Sather Column 51 Lines 51-53: sent to a multiplexer 1515, and a set of configuration bits is used to select between these two possible inputs)
and an output terminal coupled to the [sixth input terminal of the third macro unit] (Sather Column 51 Lines 51-53: sent to a multiplexer 1515, and a set of configuration bits is used to select between these two possible inputs).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teaching of Guan in view of Sather with the input feature information, activations, weights, and multiplexers as taught by Sather. One of ordinary skill in the art would be motivated to make this combination because it would allow the system to operate to compute neural network operations in an efficient, low-power manner, according to the configuration data provided by the control circuits as taught by Sather (Sather Columns 42-43 Lines 67-2). Also, this would increase the flexibility of the system as it would allow for different outputs depending on the needs of the system.
While Guan in view of Sather teaches of input terminals, they fail to specifically teach of the first, second, third, fourth, fifth, sixth, seventh, eighth input terminals, and second output terminal of the third macro unit, the first, second, and third input terminals and second output terminal of the calculating activation and derivative unit, [a first input terminal] coupled [to the] second [output terminal], the first, second, third, and fourth input terminals of the weighting gradient calculation unit, [a] first [input terminal] coupled [to the] second [output terminal].
However, Gurram teaches the first, second, third, fourth, fifth, sixth, seventh, eighth input terminals, and second output terminal of the third macro unit (Gurram [0207]: he components of the hardware system 1800 can be included in any ALU, FPU, system of ALUs, or processing elements described herein. The system 1800 includes input hardware 1802... and output hardware 1810; Gurram Fig. 18: Shows multiple inputs and multiple outputs)
The first, second, and third input terminals and second output terminal of the calculating activation and derivative unit (Gurram [0207]: he components of the hardware system 1800 can be included in any ALU, FPU, system of ALUs, or processing elements described herein. The system 1800 includes input hardware 1802... and output hardware 1810; Gurram Fig. 18: Shows multiple inputs and multiple outputs)
[a] first [input terminal] coupled [to the output terminal] (Gurram [0098]: A compute engine cluster 332 can include a set of compute engine tiles 340A-340D... can also be interconnected via a set of tile interconnects 323A-323F; Gurram Fig. 3C: shows the compute tiles connected together; Gurram [0207]: The components of the hardware system 1800 can be included in any ALU, FPU, system of ALUs, or processing elements described herein. The system 1800 includes input hardware 1802; Gurram Fig. 18: Shows multiple inputs)
the first, second, third, and fourth input terminals of the weighting gradient calculation unit (Gurram [0207]: The components of the hardware system 1800 can be included in any ALU, FPU, system of ALUs, or processing elements described herein. The system 1800 includes input hardware 1802; Gurram Fig. 18: Shows multiple inputs)
[a] first [input terminal] coupled [to the] second [output terminal] (Gurram [0207]: he components of the hardware system 1800 can be included in any ALU, FPU, system of ALUs, or processing elements described herein. The system 1800 includes input hardware 1802... and output hardware 1810; Gurram Fig. 18: Shows multiple inputs and multiple outputs).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teaching of Guan in view of Sather with the input and output terminals and the units being coupled together as taught by Gurram. One of ordinary skill in the art would be motivated to make this combination because multi-threaded operation enables an efficient execution environment in the face of higher latency memory accesses as taught by Gurram (Gurram [0110]).
Guan in view of Sather fails to teach [a seventh input terminal configured to receive] weighting update information; [a first output terminal configured to output] first variation information; [a third input terminal configured to receive] first variation information; [and a second output terminal configured to output] second variation information [and a derivative input multiplexer comprising: a first input terminal configured to receive] second variation information [outputted from the loss computing unit;].
However, Moon teaches [a seventh input terminal configured to receive] weighting update information; (Moon [0053]: the gradient to be used for updating the weights)
[a first output terminal configured to output] first variation information; (Moon [0053]: the gradient to be used for updating the weights)
[a third input terminal configured to receive] first variation information; (Moon [0053]: the gradient to be used for updating the weights)
[and a second output terminal configured to output] second variation information; (Moon [0053]: the derivative of the activation function of each layer)
[and a derivative input multiplexer comprising: a first input terminal configured to receive] second variation information [outputted from the loss computing unit;] (Moon [0053]: the derivative of the activation function of each layer).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teaching of Guan in view of Sather with the weighting update information and first and second variation information as taught by Moon. One of ordinary skill in the art would be motivated to make this combination because accordingly, the training of the neural network model for personalization can be efficiently performed, in particular on-device, without overhead as taught by Moon (Moon [0101]).
Guan in view of Sather further in view of Moon fails to teach [and an output terminal configured to output] third variation information.
However, Vilcu teaches [and an output terminal configured to output] third variation information; (Vilcu [0127]: The partial differential of the sum of weighted activations with respect to the weight).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teaching of Guan in view of Sather further in view of Moon with the third variation information as taught by Vilcu. One of ordinary skill in the art would be motivated to make this combination because this method processes batches of training data and losses to efficiently train a neural network as taught by Vilcu (Vilcu [0133]).
Claims 6-7 and 9-10 are rejected as being unpatentable over Guan in view of Sather further in view of Gurram further in view of Moon further in view of Vilcu further in view of Chilwal.
With regards to claim 6, Guan in view of Sather further in view of Gurram further in view of Moon further in view of Vilcu teaches all of the limitations of claim 5 above. Guan further teaches the output terminal of the second macro unit is configured to output [an output vector having (N+1) dimensions,] (Guan [0007]: The controller can compute the aggregation, combination and or other similar operations on the attributes base on the first memory access to generate result data. The controller can output the result data based on the first memory accesses).
Guan fails to teach wherein the input feature information of the second macro unit comprises an input feature vector having (M+1) dimensions, [the output terminal of the second macro unit is configured to output] an output vector having (N+1) dimensions, and M and N are two positive integers.
However, Sather teaches wherein the input feature information of the second macro unit comprises an input feature vector having (M+1) dimensions, (Sather Columns 7-8 Lines 67-2: The array can be conceptualized as a set of two-dimensional grids, also referred to as input feature maps or input channels for the layer)
[the output terminal of the second macro unit is configured to output] an output vector having (N+1) dimensions, (Sather Column 8 Lines 8-12 and 30-36: The input to each computation node is a subset of the input activation values, and the dot product for the computation node involves multiplying those input activation values by the weights... To generate the output activations, each of the filters 205 is applied to numerous subsets of the input activation values 200. Specifically, in a typical convolution layer, each 3×3×3 filter is moved across the three-dimensional array of activation values, and the dot product between the 27 activations in the current subset and the 27 weight values in the filter is computed)
and M and N are two positive integers (Sather Column 8 Lines 8-12 and 30-36: The input to each computation node is a subset of the input activation values, and the dot product for the computation node involves multiplying those input activation values by the weights... To generate the output activations, each of the filters 205 is applied to numerous subsets of the input activation values 200. Specifically, in a typical convolution layer, each 3×3×3 filter is moved across the three-dimensional array of activation values, and the dot product between the 27 activations in the current subset and the 27 weight values in the filter is computed; Sather Fig. 2: shows the inputs being a 6x6 matrix and the outputs being a 4x4 matrix).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teaching of Guan in view of Sather further in view of Gurram further in view of Moon further in view of Vilcu with the vectors and dimensions as taught by Sather. One of ordinary skill in the art would be motivated to make this combination because it would allow the system to operate to compute neural network operations in an efficient, low-power manner, according to the configuration data provided by the control circuits as taught by Sather (Sather Columns 42-43 Lines 67-2).
Guan in view of Sather fails to teach [the second macro unit further comprises] a clock control terminal and a reset terminal.
However, Chilwal teaches [the second macro unit further comprises] a clock control terminal and a reset terminal, (Chilwal Column 4 Lines 24-29: a signal routing circuit that is reconfigurable, by way of asserting corresponding combinations of the signal path control signals, to generate multiple alternative data, clock and set/reset signal paths between the two flip-flop/latch elements and generic input output nodes of the retention model).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teaching of Guan in view of Sather further in view of Gurram further in view of Moon further in view of Vilcu with the clock control and reset terminals as taught by Chilwal. One of ordinary skill in the art would be motivated to make this combination because it would facilitate efficient processing as taught by Chilwal (Chilwal Column 5 Lines 12-13). Also, this would increase the reliability of the system as the circuits would be timed by a clock, ensuring the data is moved in a synchronized way.
With regards to claim 7, Guan in view of Sather further in view of Gurram further in view of Moon further in view of Vilcu further in view of Chilwal teaches all of the limitations of claim 6 above. Guan further teaches wherein the [third] input terminal of the second macro unit is further used for receiving [(M+1) weighting difference information,] (Guan [0030]: The read data buffer (RDB) 445 and write data buffer (WDB) 460 can be multi-entry buffers used to buffer data for the computation units 440-445. The modes can include no computation, complete computation and partial computation modes. In the no computation mode, the read data buffer (RDB) 445 and write data buffer (WDB) 460 can be bypassed. In the complete computation mode, the computation units 440-445 can perform all the computations on the accessed data. In the partial computation mode, the computation units 440-445 can perform a portion of the computations on the accessed data and a partial result can be passed as the data for further computations by the compute engine 415 of the central core 405. The result data of the optional aggregation, combination or the like function can be sent by the one or more memory units 410 as return data, at 760)
the second macro unit generates [(M+1)x (N+1) weightings after the weighting information is received,] (Guan [0025]: a plurality of computation units 450-455; Guan Fig. 4: shows multiple computation units)
an output of the second macro unit is generated, (Guan [0007]: The controller can compute the aggregation, combination and or other similar operations on the attributes base on the first memory access to generate result data. The controller can output the result data based on the first memory accesses).
Guan fails to teach [the second macro unit generates] (M+1)x (N+1) weightings after the weighting information is received, after (M+1) weightings of each column are linearly combined with the input feature vector having (M+1) dimensions, and the (M+1) weightings of each column of the (M+1)x (N+1) weightings are updated according to the (M+1) weighting difference information.
However, Sather teaches [the second macro unit generates] (M+1)x (N+1) weightings after the weighting information is received, (Sather Column 8 Lines 8-13: The input to each computation node is a subset of the input activation values, and the dot product for the computation node involves multiplying those input activation values by the weights that make up one of the filters of the layer. As shown, in this example the layer includes six filters 205, each of which is 3×3×3)
after (M+1) weightings of each column are linearly combined with the input feature vector having (M+1) dimensions, (Sather Column 8 Lines 8-12 and 30-36: The input to each computation node is a subset of the input activation values, and the dot product for the computation node involves multiplying those input activation values by the weights... To generate the output activations, each of the filters 205 is applied to numerous subsets of the input activation values 200. Specifically, in a typical convolution layer, each 3×3×3 filter is moved across the three-dimensional array of activation values, and the dot product between the 27 activations in the current subset and the 27 weight values in the filter is computed)
and the (M+1) weightings of each column of the (M+1)x (N+1) weightings [are updated according to the (M+1) weighting difference information] (Sather Column 8 Lines 8-12 and 30-36: The input to each computation node is a subset of the input activation values, and the dot product for the computation node involves multiplying those input activation values by the weights... To generate the output activations, each of the filters 205 is applied to numerous subsets of the input activation values 200. Specifically, in a typical convolution layer, each 3×3×3 filter is moved across the three-dimensional array of activation values, and the dot product between the 27 activations in the current subset and the 27 weight values in the filter is computed).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teaching of Guan in view of Sather further in view of Gurram further in view of Moon further in view of Vilcu further in view of Chilwal with the vectors and dimensions as taught by Sather. One of ordinary skill in the art would be motivated to make this combination because it would allow the system to operate to compute neural network operations in an efficient, low-power manner, according to the configuration data provided by the control circuits as taught by Sather (Sather Columns 42-43 Lines 67-2).
Guan in view of Sather fails to teach [wherein the] third [input terminal of the second macro unit].
However, Gurram teaches [wherein the] third [input terminal of the second macro unit] (Gurram [0207]: The components of the hardware system 1800 can be included in any ALU, FPU, system of ALUs, or processing elements described herein. The system 1800 includes input hardware 1802; Gurram Fig. 18: Shows multiple inputs).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teaching of Guan in view of Sather further in view of Gurram further in view of Moon further in view of Vilcu further in view of Chilwal with the input terminals as taught by Gurram. One of ordinary skill in the art would be motivated to make this combination because multi-threaded operation enables an efficient execution environment in the face of higher latency memory accesses as taught by Gurram (Gurram [0110]).
Guan in view of Sather further in view of Gurram fails to teach [wherein the third input terminal of the second macro unit is further used for receiving] (M+1) weighting difference information, [and the (M+1) weightings of each column of the (M+1)x (N+1) weightings are] updated according to the (M+1) weighting difference information.
However, Moon teaches [wherein the third input terminal of the second macro unit is further used for receiving] (M+1) weighting difference information, (Moon [0053]: the gradient to be used for updating the weights)
[and the (M+1) weightings of each column of the (M+1)x (N+1) weightings are] updated according to the (M+1) weighting difference information (Moon [0053]: the gradient to be used for updating the weights).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teaching of Guan in view of Sather further in view of Gurram further in view of Moon further in view of Vilcu further in view of Chilwal with the weighting difference information as taught by Moon. One of ordinary skill in the art would be motivated to make this combination because accordingly, the training of the neural network model for personalization can be efficiently performed, in particular on-device, without overhead as taught by Moon (Moon [0101]).
With regards to claim 9, Guan in view of Sather further in view of Gurram further in view of Moon further in view of Vilcu teaches all of the limitations of claim 8 above. Guan further teaches the output terminal of the third macro unit is configured to output [an output vector having (N+1) dimensions,] (Guan [0007]: The controller can compute the aggregation, combination and or other similar operations on the attributes base on the first memory access to generate result data. The controller can output the result data based on the first memory accesses)
the first output terminal of the third macro unit is configured to output [(M+1) first derivatives,] (Guan [0007]: The controller can compute the aggregation, combination and or other similar operations on the attributes base on the first memory access to generate result data. The controller can output the result data based on the first memory accesses)
the [sixth] input terminal of the third macro unit is configured to receive [(N+1) second derivatives outputted from the derivative input multiplexer,] (Guan [0030]: The read data buffer (RDB) 445 and write data buffer (WDB) 460 can be multi-entry buffers used to buffer data for the computation units 440-445. The modes can include no computation, complete computation and partial computation modes. In the no computation mode, the read data buffer (RDB) 445 and write data buffer (WDB) 460 can be bypassed. In the complete computation mode, the computation units 440-445 can perform all the computations on the accessed data. In the partial computation mode, the computation units 440-445 can perform a portion of the computations on the accessed data and a partial result can be passed as the data for further computations by the compute engine 415 of the central core 405. The result data of the optional aggregation, combination or the like function can be sent by the one or more memory units 410 as return data, at 760).
Guan fails to teach wherein the input feature information of the third macro unit comprises an input feature vector having (M+1) dimensions, [the output terminal of the third macro unit is configured to output] an output vector having (N+1) dimensions, [the] sixth [input terminal of the third macro unit is configured to receive] (N+1) second derivatives outputted from the derivative input multiplexer, and M and N are two positive integers.
However, Sather teaches wherein the input feature information of the second macro unit comprises an input feature vector having (M+1) dimensions, (Sather Columns 7-8 Lines 67-2: The array can be conceptualized as a set of two-dimensional grids, also referred to as input feature maps or input channels for the layer)
[the output terminal of the third macro unit is configured to output] an output vector having (N+1) dimensions, (Sather Column 8 Lines 8-12 and 30-36: The input to each computation node is a subset of the input activation values, and the dot product for the computation node involves multiplying those input activation values by the weights... To generate the output activations, each of the filters 205 is applied to numerous subsets of the input activation values 200. Specifically, in a typical convolution layer, each 3×3×3 filter is moved across the three-dimensional array of activation values, and the dot product between the 27 activations in the current subset and the 27 weight values in the filter is computed)
[the sixth input terminal of the third macro unit is configured to receive (N+1) second derivatives] outputted from the derivative input multiplexer, (Sather Column 51 Lines 51-53: The multiplexers 1710 each have eight inputs and receive a set of select bits (the weight selector input) from the core controller)
and M and N are two positive integers (Sather Column 8 Lines 8-12 and 30-36: The input to each computation node is a subset of the input activation values, and the dot product for the computation node involves multiplying those input activation values by the weights... To generate the output activations, each of the filters 205 is applied to numerous subsets of the input activation values 200. Specifically, in a typical convolution layer, each 3×3×3 filter is moved across the three-dimensional array of activation values, and the dot product between the 27 activations in the current subset and the 27 weight values in the filter is computed; Sather Fig. 2: shows the inputs being a 6x6 matrix and the outputs being a 4x4 matrix).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teaching of Guan in view of Sather further in view of Gurram further in view of Moon further in view of Vilcu with the vectors and dimensions as taught by Sather. One of ordinary skill in the art would be motivated to make this combination because it would allow the system to operate to compute neural network operations in an efficient, low-power manner, according to the configuration data provided by the control circuits as taught by Sather (Sather Columns 42-43 Lines 67-2).
Guan in view of Sather fails to teach [the] sixth [input terminal of the third macro unit].
However, Gurram teaches [the] sixth [input terminal of the third macro unit] (Gurram [0207]: The components of the hardware system 1800 can be included in any ALU, FPU, system of ALUs, or processing elements described herein. The system 1800 includes input hardware 1802; Gurram Fig. 18: Shows multiple inputs).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teaching of Guan in view of Sather further in view of Gurram further in view of Moon further in view of Vilcu with the input terminals as taught by Gurram. One of ordinary skill in the art would be motivated to make this combination because multi-threaded operation enables an efficient execution environment in the face of higher latency memory accesses as taught by Gurram (Gurram [0110]).
Guan in view of Sather further in view of Gurram fails to teach [the first output terminal of the third macro unit is configured to output] (M+1) first derivatives, [the sixth input terminal of the third macro unit is configured to receive] (N+1) second derivatives [outputted from the derivative input multiplexer,].
However, Moon teaches [the first output terminal of the third macro unit is configured to output] (M+1) first derivatives, (Moon [0053]: the gradient to be used for updating the weights)
[the sixth input terminal of the third macro unit is configured to receive] (N+1) second derivatives [outputted from the derivative input multiplexer,] (Moon [0053]: the derivative of the activation function of each layer).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teaching of Guan in view of Sather further in view of Gurram further in view of Moon further in view of Vilcu with the derivatives as taught by Moon. One of ordinary skill in the art would be motivated to make this combination because accordingly, the training of the neural network model for personalization can be efficiently performed, in particular on-device, without overhead as taught by Moon (Moon [0101]).
Guan in view of Sather further in view of Gurran further in view of Moon fails to teach [the third macro unit further comprises] a clock control terminal and a reset terminal.
However, Chilwal teaches [the third macro unit further comprises] a clock control terminal and a reset terminal, (Chilwal Column 4 Lines 24-29: a signal routing circuit that is reconfigurable, by way of asserting corresponding combinations of the signal path control signals, to generate multiple alternative data, clock and set/reset signal paths between the two flip-flop/latch elements and generic input output nodes of the retention model).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teaching of Guan in view of Sather further in view of Gurram further in view of Moon further in view of Vilcu with the clock control and reset terminals as taught by Chilwal. One of ordinary skill in the art would be motivated to make this combination because it would facilitate efficient processing as taught by Chilwal (Chilwal Column 5 Lines 12-13). Also, this would increase the reliability of the system as the circuits would be timed by a clock, ensuring the data is moved in a synchronized way.
With regards to claim 10, Guan in view of Sather further in view of Gurram further in view of Moon further in view of Vilcu further in view of Chilwal teaches all of the limitations of claim 6 above. Guan further teaches wherein the [third] input terminal of the third macro unit is further used for receiving [(M+1) weighting difference information,] (Guan [0030]: The read data buffer (RDB) 445 and write data buffer (WDB) 460 can be multi-entry buffers used to buffer data for the computation units 440-445. The modes can include no computation, complete computation and partial computation modes. In the no computation mode, the read data buffer (RDB) 445 and write data buffer (WDB) 460 can be bypassed. In the complete computation mode, the computation units 440-445 can perform all the computations on the accessed data. In the partial computation mode, the computation units 440-445 can perform a portion of the computations on the accessed data and a partial result can be passed as the data for further computations by the compute engine 415 of the central core 405. The result data of the optional aggregation, combination or the like function can be sent by the one or more memory units 410 as return data, at 760)
the third macro unit generates [(M+1)x (N+1) weightings after the weighting information is received,] (Guan [0025]: a plurality of computation units 450-455; Guan Fig. 4: shows multiple computation units)
an output of the third macro unit is generated, (Guan [0007]: The controller can compute the aggregation, combination and or other similar operations on the attributes base on the first memory access to generate result data. The controller can output the result data based on the first memory accesses)
[and after the (N+1) second derivatives are linearly combined with (N+1) weightings of each row] by the third macro unit, (Guan [0025]: a plurality of computation units 450-455; Guan Fig. 4: shows multiple computation units)
[the (M+1) first derivatives outputted from the] first output terminal of the third macro unit are generated (Guan [0007]: The controller can compute the aggregation, combination and or other similar operations on the attributes base on the first memory access to generate result data. The controller can output the result data based on the first memory accesses).
Guan fails to teach [the third macro unit generates] (M+1)x (N+1) weightings after the weighting information is received, after (M+1) weightings of each column are linearly combined with the input feature vector having (M+1) dimensions, the (M+1) weightings of each column of the (M+1)x (N+1) weightings are updated according to the (M+1) weighting difference information and after the (N+1) second derivatives are linearly combined with (N+1) weightings of each row [by the third macro unit].
However, Sather teaches [the second macro unit generates] (M+1)x (N+1) weightings after the weighting information is received, (Sather Column 8 Lines 8-13: The input to each computation node is a subset of the input activation values, and the dot product for the computation node involves multiplying those input activation values by the weights that make up one of the filters of the layer. As shown, in this example the layer includes six filters 205, each of which is 3×3×3)
after (M+1) weightings of each column are linearly combined with the input feature vector having (M+1) dimensions, (Sather Column 8 Lines 8-12 and 30-36: The input to each computation node is a subset of the input activation values, and the dot product for the computation node involves multiplying those input activation values by the weights... To generate the output activations, each of the filters 205 is applied to numerous subsets of the input activation values 200. Specifically, in a typical convolution layer, each 3×3×3 filter is moved across the three-dimensional array of activation values, and the dot product between the 27 activations in the current subset and the 27 weight values in the filter is computed)
and the (M+1) weightings of each column of the (M+1)x (N+1) weightings [are updated according to the (M+1) weighting difference information] (Sather Column 8 Lines 8-12 and 30-36: The input to each computation node is a subset of the input activation values, and the dot product for the computation node involves multiplying those input activation values by the weights... To generate the output activations, each of the filters 205 is applied to numerous subsets of the input activation values 200. Specifically, in a typical convolution layer, each 3×3×3 filter is moved across the three-dimensional array of activation values, and the dot product between the 27 activations in the current subset and the 27 weight values in the filter is computed)
[and after the (N+1) second derivatives] are linearly combined with (N+1) weightings of each row [by the third macro unit] (Sather Column 8 Lines 8-12 and 30-36: The input to each computation node is a subset of the input activation values, and the dot product for the computation node involves multiplying those input activation values by the weights... To generate the output activations, each of the filters 205 is applied to numerous subsets of the input activation values 200. Specifically, in a typical convolution layer, each 3×3×3 filter is moved across the three-dimensional array of activation values, and the dot product between the 27 activations in the current subset and the 27 weight values in the filter is computed).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teaching of Guan in view of Sather further in view of Gurram further in view of Moon further in view of Vilcu further in view of Chilwal with the vectors and dimensions as taught by Sather. One of ordinary skill in the art would be motivated to make this combination because it would allow the system to operate to compute neural network operations in an efficient, low-power manner, according to the configuration data provided by the control circuits as taught by Sather (Sather Columns 42-43 Lines 67-2).
Guan in view of Sather fails to teach [wherein the] third [input terminal of the third macro unit].
However, Gurram teaches [wherein the] third [input terminal of the third macro unit] (Gurram [0207]: The components of the hardware system 1800 can be included in any ALU, FPU, system of ALUs, or processing elements described herein. The system 1800 includes input hardware 1802; Gurram Fig. 18: Shows multiple inputs).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teaching of Guan in view of Sather further in view of Gurram further in view of Moon further in view of Vilcu further in view of Chilwal with the input terminals as taught by Gurram. One of ordinary skill in the art would be motivated to make this combination because multi-threaded operation enables an efficient execution environment in the face of higher latency memory accesses as taught by Gurram (Gurram [0110]).
Guan in view of Sather further in view of Gurram fails to teach [wherein the third input terminal of the second macro unit is further used for receiving] (M+1) weighting difference information, [and the (M+1) weightings of each column of the (M+1)x (N+1) weightings are] updated according to the (M+1) weighting difference information and after the (N+1) second derivatives [are linearly combined with (N+1) weightings of each row by the third macro unit,] the (M+1) first derivatives [outputted from the first output terminal of the third macro unit are generated].
However, Moon teaches [wherein the third input terminal of the second macro unit is further used for receiving] (M+1) weighting difference information, (Moon [0053]: the gradient to be used for updating the weights)
[and the (M+1) weightings of each column of the (M+1)x (N+1) weightings are] updated according to the (M+1) weighting difference information (Moon [0053]: the gradient to be used for updating the weights)
and after the (N+1) second derivatives [are linearly combined with (N+1) weightings of each row by the third macro unit,] (Moon [0053]: the gradient to be used for updating the weights)
the (M+1) first derivatives [outputted from the first output terminal of the third macro unit are generated] (Moon [0053]: the derivative of the activation function of each layer).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teaching of Guan in view of Sather further in view of Gurram further in view of Moon further in view of Vilcu further in view of Chilwal with the weighting difference information and derivatives as taught by Moon. One of ordinary skill in the art would be motivated to make this combination because accordingly, the training of the neural network model for personalization can be efficiently performed, in particular on-device, without overhead as taught by Moon (Moon [0101]).
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Jakob O Gudas whose telephone number is (571)272-0695. The examiner can normally be reached Monday-Thursday: 7:30AM-5:00PM Friday: 7:30AM-4:00PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, James Trujillo can be reached at (571) 272-3677. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/J.O.G./Examiner, Art Unit 2151
/James Trujillo/Supervisory Patent Examiner, Art Unit 2151