Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 103
The text of those sections of Title 35, U.S. Code not included in this action can be found in a prior Office action.
Claim(s) 1, 3-4, 8, 10-11, 15, 17, and 21 is/are rejected under 35 U.S.C. 103 as being unpatentable over Gokmen et al. (From IDS: US Patent No. 9,646,243, published May 2017, hereinafter “Gokmen”) in view of Chiu et al. (NPL: A Binarized Neural Network Accelerator with Differential Crosspoint Memristor Array for Energy-Efficient MAC Operations, published May 2019, hereinafter “Chiu”) and further in view of Song et al. (NPL: PipeLayer: A Pipeline ReRAM-Based Accelerator for Deep Learning, published May 2017, hereinafter “Song”) and Boybat Kara et al. (US Pub. No. 2019/0122105, published April 2019, hereinafter “Boybat Kara”).
Regarding claim 1, Gokmen teaches a method for data processing, the method comprising:
receiving, by a neural network system, training data, (Gokmen, Page 41 Col. 15 Lines 23-28 – “For example, as shown in FIG. 8, the current I.sub.4 generated by column wire 814 is according to the equation I.sub.4=V.sub.1σ.sub.41+V.sub.2σ.sub.42+V.sub.3σ.sub.43. Thus, array 800 computes the forward matrix multiplication by multiplying the values stored in the RPUs by the row wire inputs, which are defined by voltages V.sub.1, V.sub.2, V.sub.3. ” – teaches receiving training data (row wire inputs defined by voltages)), wherein the neural network system comprises a plurality of neural network arrays, each neural network array of the plurality of neural network arrays comprises a plurality of in-memory computing units, and each in-memory computing unit of the plurality of in-memory computing units is configured to store a weight value of a neuron in a corresponding neural network array (Gokmen, Page 41 Col. 15 Lines 3-14 – “FIG. 8 is a diagram of a two-dimensional (2D) crossbar array 800 that performs forward matrix multiplication, backward matrix multiplication and weight updates according to the present description. Crossbar array 800 is formed from a set of conductive row wires 802, 804, 806 and a set of conductive column wires 808, 810, 812, 814 that intersect the set of conductive row wires 802, 804, 806. The intersections between the set of row wires and the set of column wires are separated by RPUs, which are shown in FIG. 8 as resistive elements each having its own adjustable/updateable resistive weight, depicted as σ.sub.11, σ.sub.21, σ.sub.31, σ.sub.41, σ.sub.12, σ.sub.22, σ.sub.32, σ.sub.42, σ.sub.13, σ.sub.23, σ.sub.33 and σ.sub.43, respectively.” – teaches the neural network system comprising a plurality of neural network arrays (crossbar array 800), where each neural network array comprises a plurality of in-memory computing units (conductive row wires, RPUs), and each in-memory compute unit of the plurality is configured to store a weight value of a neuron in a corresponding neural network array (adjustable/updateable resistive weight));
generating, by the neural network system, first output data based on the training data (Gokmen, Page 41 Col. 15 Lines 23-28 – “For example, as shown in FIG. 8, the current I.sub.4 generated by column wire 814 is according to the equation I.sub.4=V.sub.1σ.sub.41+V.sub.2σ.sub.42+V.sub.3σ.sub.43. Thus, array 800 computes the forward matrix multiplication by multiplying the values stored in the RPUs by the row wire inputs, which are defined by voltages V.sub.1, V.sub.2, V.sub.3.” – teaches generating first output data based on the training data (computes forward matrix multiplication of values stored in RPUs by row wire inputs));
Gokmen fails to explicitly teach wherein the plurality of neural network arrays comprises a first neural network array, a second neural network array, and a third neural network array, wherein input data of the first neural network array comprises output data of the second neural network array and wherein the third neural network array and the second neural network array are configured to implement computing of a convolutional layer in the neural network system in parallel.
However, analogous to the field of the claimed invention, Chiu teaches:
wherein the plurality of neural network arrays comprises a first neural network array, a second neural network array, and a third neural network array, (Chiu, Section II Paragraph 1 – “Figure I(a) shows the proposed differential crosspoint (DX) memristor array”, Section III Paragraph 1 – “A DX unit (DXU) includes a DX array with peripheral circuits, such as WL drivers, control logic, column multiplexers and sense amplifiers.” and in Section IV Subsection A Paragraph 1 – “Figure 6 shows an example of mapping one convolution layer to multiple DXUs. A convolution layer with 643 x3x64 filters can be mapped onto nine DXUs.” – teaches wherein the plurality of neural network arrays (DX memristor arrays) comprises a first neural network array, a second neural network array, and a third neural network array (maps a convolution layer onto nine DXUs, thus teaching wherein the plurality of neural network arrays comprises a first, second, and third neural network array)),
wherein input data of the first neural network array comprises output data of the second neural network array (Chiu, Fig. 5 and in Section IV Paragraph 1 – “The input and output of the DXU are stored in input buffers (IB) and output buffers (OB). The input/output transactions are sent to and from the DXUs through shared buses.” – teaches wherein input data of the first neural network array comprises output data of the second neural network array (input/output transactions are sent to and from the DXUs through shared buses, thus the input data of the first neural network array comprises output data of the second neural network array)) and wherein the third neural network array and the second neural network array are configured to implement computing of a convolutional layer in the neural network system in parallel (Chiu, Section III Paragraph 2 – “The control logic takes commands from the scheduler (Figure 5) to issue parallel MAC operations.”, Section IV Subsection A Paragraph 1 – “Figure 6 shows an example of mapping one convolution layer to multiple DXUs. A convolution layer with 643 x3x64 filters can be mapped onto nine DXUs.”, and in Fig. 5 – teaches wherein the third neural network array and the second neural network array are configured to implement computing of a convolutional layer (maps convolutional layer onto nine DXUs, where each DXU includes a DX array, thus teaching wherein the third neural network array and the second neural network array are configured to implement computing of a convolutional layer) in the neural network system in parallel (control logic takes commands to issue parallel MAC operations, and Fig. 5 shows the DXUs implementing computing of a convolutional layer in parallel));
Therefore, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to incorporate the first, second, and third neural network arrays and parallel architecture of Chiu to the neural network system, training data, and compute units of Gokmen. Doing so would enable parallel multiply-and-accumulate (MAC) operations using crosspoint memristor arrays to further improve efficiency and provide quicker MAC operations (Chiu, Abstract).
The combination of Gokmen and Chiu fails to explicitly teach calculating, by the neural network system, a deviation between the first output data and target output data; and adjusting, by the neural network system, based on the deviation, a weight value stored in at least one in-memory computing unit in at least one neural network array in the plurality of neural network arrays, wherein the at least one neural network array is configured to implement computing of at least a portion of one neural network layer in the neural network system, wherein the adjusting, based on the deviation, the weight value stored in the at least one in-memory computing unit in the at least one neural network array in the plurality of neural network arrays comprises: independently adjusting, based on the first sub-deviation and input data of the second neural network array, a weight value stored in at least one in-memory computing unit in the second neural network array; and independently adjusting, based on the second sub-deviation and input data of the third neural network array, a weight value stored in at least one in-memory computing unit in the third neural network array.
However, analogous to the field of the claimed invention, Song teaches:
calculating, by the neural network system, a deviation between the first output data and target output data (Song, Section 2.2 Paragraph 2 – “In training phase, a cost function is defined to quantitatively evaluate how well the outputs of a neural network compare to the standard labels. We use y and t to represent the output of a neural network and the standard label respectively. An L2 norm loss function is defined as J(W, b)=12∥y−t∥22 and J(W, b)=−∑1(yi=tj)logp(yi=tj) is the softmax i,j loss function.” – teaches calculating a deviation (loss) between first output data (output y) and target output data (label t)); and
adjusting, by the neural network system, based on the deviation, a weight value stored in at least one in-memory computing unit in at least one neural network array in the plurality of neural network arrays, wherein the at least one neural network array is configured to implement computing of at least a portion of one neural network layer in the neural network system (Song, Section 2.2 Paragraphs 3-4 “The error δ for each layer is defined as: δl≜∂J/∂bl. If we use an L2 norm loss function, for the last (output) layer L, the error is δLf′(uL)∘(y−t) where o represents a Hadamard product, i.e. element-wise multiplications. For other layers excluding the output layer, the error is δl=(Wl+1)1δl+1∘f′(ul). And with a ReLU activation function, the error can be rewritten as δl=(Wl+1)Tδl+1∘f′(dl). So that the backward partial derivatives to Wl is ∂J∂Wl=dl−1(δl)T. And the backward partial derivatives to bl is ∂J∂bl=δl. Now we can use the gradient descent method to update the weights of neural network.” – teaches adjusting weight values (uses gradient descent based on the loss to update the weights of the neural network) based on the deviation, and in Section 3.1 Paragraph 4 – “In T5, two computations happen in parallel: (1) partial derivatives (∇W3) is computed by previous results in d2 and δ3; (2) errors (δ2) of the second layer is computed from δ3. Both of the computations depend on δ3, which is computed in T4.∇W3 is stored in memory subarrays, which will be used to update weights in A3 and A32 later.” - teaches the adjusted weights being stored in at least one in-memory compute unit in at least one neural network array (stored in memory subarrays), where the neural network arrays are configured to implement the computing of at least a portion of one neural network layer in the neural network system (computes errors of layers to determine an update for the weights of the layers)),
wherein the adjusting, based on the deviation, the weight value stored in the at least one in-memory computing unit in the at least one neural network array in the plurality of neural network arrays comprises:
independently adjusting, based on the first sub-deviation and input data of the second neural network array, a weight value stored in at least one in-memory computing unit in the second neural network array (Song, Section 2.2 Paragraph 4 – “And with a ReLU activation function, the error can be rewritten as δl=(Wl+1)Tδl+1∘f′(dl). So that the backward partial derivatives to Wl is ∂J∂Wl=dl−1(δl)T. And the backward partial derivatives to bl is ∂J∂bl=δl Now we can use the gradient descent method to update the weights of neural network.” – teaches independently adjusting a weight value stored in at least one in-memory computing unit (Wl – where l may be set to 2 to represent the weight stored in at least one in-memory computing unit in the second neural network array) based on input data (dl-1 – the output of the previous layer, therefore the input of the current layer) and the deviation (δl)); and
independently adjusting, based on the second sub-deviation and input data of the third neural network array, a weight value stored in at least one in-memory computing unit in the third neural network array (Song, Section 2.2 Paragraph 4 – “And with a ReLU activation function, the error can be rewritten as δl=(Wl+1)Tδl+1∘f′(dl). So that the backward partial derivatives to Wl is ∂J∂Wl=dl−1(δl)T. And the backward partial derivatives to bl is ∂J∂bl=δl Now we can use the gradient descent method to update the weights of neural network.” – teaches independently adjusting a weight value stored in at least one in-memory computing unit (Wl – where l may be set to 3 to represent the weight stored in at least one in-memory computing unit in the third neural network array) based on input data (dl-1 – the output of the previous layer, therefore the input of the current layer) and the deviation (δl)).
Therefore, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to incorporate the deviation calculations and weight adjustments of Song to the neural network system, neural network arrays, and in-memory computing units of Gokmen and Chiu in order to update weights stored in in-memory computing units of a neural network array. Doing so would reduce data movements in memory hierarchy (Song, Introduction) and support training that involves weight updates in ReRAM-based architectures (Song, Section 2.3).
The combination of Gokmen, Chiu, and Song fails to explicitly teach wherein an initial value of the weight value is determined by performing offline training, and dividing the deviation into at least different two sub-deviations, wherein a first sub-deviation in the at least two sub-deviations corresponds to the output data of the second neural network array, and a second sub-deviation in the at least two sub-deviations corresponds to output data of the third neural network array, wherein the second sub-deviation is different from the first sub-deviation.
However, analogous to the field of resistive random-access memory and parallel acceleration, Boybat Kara teaches:
wherein an initial value of the weight value of a given neuron is determined by performing offline training (Boybat Kara, [0034] – “The second array a.sub.2 of array-set 6 implements the layer of synapses s.sub.jk between the second and third neuron layers of ANN 1. Structure corresponds directly to that of array a.sub.1. Hence, devices 10 of array a.sub.2 store weights Ŵ.sub.jk for synapses s.sub.jk, with row lines r.sub.j representing connections between respective layer 2 neurons n.sub.2j and synapses s.sub.jk, and column lines c.sub.k representing connections between respective output layer neurons n.sub.3k and synapses s.sub.jk.” and in [0044] – “The weights Ŵ stored in memristive arrays 6 may be initialized to predetermined values, or may be randomly distributed for the start of the training process.” and in [0051] – “Also, weight updates may be performed after backpropagation in every iteration of the training scheme (“online training”), or after a certain number K of iterations (“batch training”).” – teaches wherein an initial value of the weight value of a given neuron is determined by performing offline training (weights stored in memristive arrays 6 correspond to synapses between neurons, weights may be initialized to predetermined values, weight updates may be performed after certain number K iterations of batch training, which is offline training)), and
dividing the deviation into at least different two sub-deviations, wherein a first sub-deviation in the at least two sub-deviations corresponds to the output data of the second neural network array, and a second sub-deviation in the at least two sub-deviations corresponds to output data of the third neural network array, wherein the second sub-deviation is different from the first sub-deviation (Boybat Kara, [0043] – “In the forward propagation operation (step 30), the input data for a current training sample is forward-propagated through ANN 1 from the input to the output neuron layer. This operation, detailed further below, involves calculating outputs, denoted by x.sub.1i, x.sub.2j and x.sub.3k respectively, for neurons n.sub.1i, n.sub.2j, and n.sub.3k in DPU 4, and application of input signals to memristive arrays 6 to obtain array output signals used in these calculations. In the subsequent back-propagation operation (step 31), DPU 4 calculates error values (denoted by 631) for respective output neurons n.sub.3k and propagates these error values back through ANN 1 from the output layer to the penultimate layer in the backpropagation direction. This involves application of input signals to memristive arrays 6 to obtain array output signals, and calculation of error values for neurons in all other layers except the input neuron layer, in this case errors δ.sub.2j for the layer 2 neurons n.sub.2j. In a subsequent weight update operation (step 32), the DPU computes digital weight-correction values ΔW for respective memristive devices 10 using values computed in the forward and backpropagation steps. The DPU 4 then controls memcomputing unit 3 to applying programming signals to the devices to update the stored weights Ŵ in dependence on the respective digital weight-correction values ΔW. ” and in [0044] – “Weight update step 32 may be performed for every iteration, or after a predetermined number of backpropagation operations, and may involve update of all or a selected subset of the weights Ŵ as described further below.” – teaches dividing the deviation into at least different two sub-deviations (errors δ.sub.2j for the layer 2 neurons n.sub.2j split into per-device weight correction values, DPU distributes different weight-correction values for respective memristive devices and updates stores weights in dependence on respective weight-correction values), wherein a first sub-deviation corresponds to the output data of the second neural network array (errors δ.sub.2j for the layer 2 neurons n.sub.2j and array output signals are used to compute respective weight correction values ΔW for each array) and a second sub-deviation corresponding to the output data of the third neural network array (errors δ.sub.2j for the layer 2 neurons n.sub.2j and array output signals are used to compute respective weight correction values ΔW for each memristive array, each respective weight correction value is based on the deviation δ and respective array output signal), wherein the second sub-deviation is different from the first sub-deviation (updates stored weights W in dependence on respective digital weight correction values ΔW for respective memristive devices, thus the sub-deviations are different))
Therefore, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to incorporate the offline learning and sub-deviations of Boybat Kara to the neural network arrays, deviations, and weight adjustment of Gokmen, Chiu, and Song in order to determine initial values for the weights and adjust their values according to respective deviations. Doing so would exploit the capabilities of memristive arrays and dramatically reduce the computational complexity associated with ANN training (Boybat Kara, [0005]).
Claims 8 and 15 incorporate substantively all the limitations of claim 1 in a system and a chip, and are rejected on similar grounds as above. Gokmen teaches the memory, processors, and data interface of these claims 8 and 15 at Pg. 34, Col. 1 Line 62 - Col. 2 Line 3 – “The processor configures the RPU array corresponding to the convolution layer based on dimensions associated with convolution kernels of the convolution layer.”, Pg. 34, Col. 2, Lines 20-25 – “a computer program product for training a convolution layer of a convolutional neural network (CNN) using resistive processing unit (RPU) arrays includes a computer readable storage medium.”, and Fig. 19 – teaches processor, memory, and a neuron data interface.
Regarding claim 3, the combination of Gokmen, Chiu, Song, and Boybat Kara teaches the method according to claim 1,
wherein the first neural network array comprises a neural network array configured to implement computing of a fully-connected layer in the neural network system (Gokmen, Page 45 Col. 24 Lines 50-57 –“The neuron control system 1900, as described herein, trains the neural network with the convolutional and fully connected layers by setting up the RPU arrays with the dimensions as described herein. Further, the neuron control system 1900 converts each convolution layer in the CNN training into a fully connected layer, by converting the convolution computations into matrix multiplications as described above.” – teaches wherein the first neural network array (RPU array) comprises a neural network array configured to implement computing of a fully-connected layer in the neural network system (converts each convolution layer in the CNN into a fully-connected layer, and trains the neural network by setting up RPU arrays configured to implement the computing of these layers by means of matrix-multiplication)).
Claims 10 and 17 are similar to claim 3, hence similarly rejected.
Regarding claim 4, the combination of Gokmen, Chiu, Song, and Boybat Kara teaches the method according to claim 3, wherein the adjusting, based on the deviation, the weight value stored in the at least one in-memory computing unit in the at least one neural network arrays in the plurality of neural network arrays further comprises:
adjusting, based on input data of the first neural network array and the deviation, a weight value stored in at least one in-memory computing unit in the first neural network array (Song, Section 2.2 Paragraph 4 – “And with a ReLU activation function, the error can be rewritten as δl=(Wl+1)Tδl+1∘f′(dl). So that the backward partial derivatives to Wl is ∂J∂Wl=dl−1(δl)T. And the backward partial derivatives to bl is ∂J∂bl=δl Now we can use the gradient descent method to update the weights of neural network.” – teaches adjusting a weight value stored in at least one in-memory computing unit (Wl) based on input data (dl-1 – the output of the previous layer, therefore the input of the current layer) and the deviation (δl)).
Therefore, it would have been obvious, to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to incorporate the weight adjustment of Song to further modify the neural network arrays comprising in-memory computing units of Gokmen, Chiu, Song, and Boybat Kara in order to determine and adjust the in-memory compute unit stored weights of a fully-connected layer of a neural network system. Doing so would greatly reduce data movements and energy consumption by avoiding data being transferred across memory hierarchy (Song, Section 3.1) and support training that involves weight updates in ReRAM-based architectures (Song, Section 2.3).
Claims 11 and 21 are similar to claim 4, hence similarly rejected.
Response to Arguments
Applicant’s arguments, see pp. 2-4 of Remarks, filed 7 July 2026, with respect to the rejection(s) of claim(s) 1, 8, and 15 under 35 U.S.C. 103 have been fully considered and are persuasive. Therefore, the rejection has been withdrawn. However, upon further consideration, a new ground(s) of rejection is made over Gokmen in view of Chiu et al. (NPL: A Binarized Neural Network Accelerator with Differential Crosspoint Memristor Array for Energy-Efficient MAC Operations, published May 2019), and further in view of Song and Boybat Kara. Chiu teaches the limitations of claim 1 regarding “wherein the plurality of neural network arrays comprises… and wherein the third neural network array and the second neural network array… implement computing… in parallel”. Song teaches the amended limitations of claim 1 regarding “independently adjusting… a weight value stored in… the second neural network array” and “independently adjusting… a weight value stored in… the third neural network array”. Boybat Kara teaches the amended limitations of claim 1 regarding “dividing the deviation into at least two different sub-deviations… wherein the second sub-deviation is different from the first sub-deviation”.
Applicant’s arguments regarding Song’s mechanism including copies storing the same weights are moot in view of the new grounds of the rejection. Applicant further argues on pp. 3-4 of Remarks that Boybat Kara’s respective weight correction values are different from dividing a deviation into different sub-deviations respectively corresponding to output data of the neural network arrays. Examiner respectfully disagrees. Boybat Kara at [0043] states “In the subsequent back-propagation operation (step 31), DPU 4 calculates error values (denoted by 631) for respective output neurons n.sub.3k and propagates these error values back through ANN 1 from the output layer to the penultimate layer in the backpropagation direction. This involves application of input signals to memristive arrays 6 to obtain array output signals, and calculation of error values for neurons in all other layers except the input neuron layer, in this case errors δ.sub.2j for the layer 2 neurons n.sub.2j. In a subsequent weight update operation (step 32), the DPU computes digital weight-correction values ΔW for respective memristive devices 10 using values computed in the forward and backpropagation steps. The DPU 4 then controls memcomputing unit 3 to applying programming signals to the devices to update the stored weights Ŵ in dependence on the respective digital weight-correction values ΔW.” – which teaches dividing a deviation (errors δ.sub.2j for the layer 2 neurons) into different sub-deviations (weight-correction values ΔW), wherein a first sub-deviation corresponds to output data of the second neural network array and a second sub-deviation corresponds to output data of the third neural network array (computes respective weight-correction values ΔW using values computed in the forward and backpropagation steps, which includes application of input signals to memristive arrays 6 to obtain array output signals, and thus the weight-correction values ΔW correspond to output data of the neural network arrays). Boybat Kara states that weights Ŵ are updated depending on respective weight-correction values ΔW, thus each ΔW is different depending on the respective array output signals and weights.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Sun et al. (NPL: Cascaded Architecture for Memristor Crossbar Array Based Larger-Scale Neuromorphic Computing, published May 2019) teaches a memristor-based cascaded framework with basic computation units (BCU). Teaches a plurality of BCUs configured to implement computing of a convolutional layer in parallel. Teaches a plurality of neural network arrays comprising a plurality of in-memory computing units. Teaches initial weights that are predetermined.
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to LOUIS C NYE whose telephone number is 571-272-0636. The examiner can normally be reached Monday - Friday 9:00AM - 5:00PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, MATT ELL can be reached at 571-270-3264. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/LOUIS CHRISTOPHER NYE/Examiner, Art Unit 2141
/MATTHEW ELL/Supervisory Patent Examiner, Art Unit 2141