Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Priority
Acknowledgment is made of applicant’s claim for foreign priority under 35 U.S.C. 119 (a)-(d). The certified copy has been filed in parent Application No. 10-2022-0187082, filed on December 28, 2022.
Information Disclosure Statement
The information disclosure statement (IDS) submitted on December 6, 2024 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Objections
Claims 8 and 9 are objected to because of the following informalities:
In claim 8, line 1, "the method of claim 1" should read "the method of claim 6."
Further, in claim 8, Eq. 1, “window” and “A” are defined in the specification but not in the claims.
Claim 9 is dependent on claim 8, and includes its limitations. Claim 9 does not address the limitations of claim 8; hence claim 9 is objected to for the same reason as claim 8.
Appropriate correction is required.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-11 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1: Claims 1-5 are device claims. Claims 6-11 are method type claims. Therefore, the claims are directed to either a process, machine, manufacture or composition of matter.
Claim 1:
Regarding claim 1, in step 1 of the 101-analysis set forth in MPEP 2106, the claim recites “a synaptic array device,” and a device or machine is one of the four statutory categories of invention.
In step 2A prong 1 of the 101-analysis set forth in MPEP 2106, the examiner has determined that the following limitations recite a process that, under the broadest reasonable interpretation, covers a mathematical concept but for recitation of generic computer components:
“(deriving) a moving average value by averaging accumulated values of the gradient values of the gradient values received from the second synaptic array” (this is a mathematical concept; the equation per the specification is
m
o
v
i
n
g
a
v
e
r
a
g
e
=
m
o
v
i
n
g
a
v
e
r
a
g
e
*
w
i
n
d
o
w
-
1
w
i
n
d
o
w
+
A
*
1
w
i
n
d
o
w
, see MPEP §2106.04(a)(2)(I)),
If claim limitations, under the broadest reasonable interpretation, covers performance of the limitations as a mathematical concept but for the recitation of generic computer components, then it falls within the mathematical concepts grouping of abstract ideas. Accordingly, the claim “recites” an abstract idea.
In step 2A prong 2 of the 101-analysis set forth in MPEP 2106, the examiner has determined that the following additional elements do not integrate this judicial exception into a practical application:
i. “A synaptic array device comprising” (A synaptic array device is considered a generic computer component being used as a tool to perform the judicial exception – see MPEP § 2106.05(f)).
ii. “a first synaptic array representing weight values” (A synaptic array is considered a generic computer component being used as a tool to perform the judicial exception – see MPEP § 2106.05(f)).
iii. “a second synaptic array” (A synaptic array is considered a generic computer component being used as a tool to perform the judicial exception – see MPEP § 2106.05(f)).
iv. “receiving the error gradient of the weights of the first synaptic array and representing gradient values refined in row units” (Receiving the error gradients is considered insignificant extra-solution activity of mere data gathering – see MPEP § 2106.05(g)).
v. “a third synaptic array” (A synaptic array is considered a generic computer component being used as a tool to perform the judicial exception – see MPEP § 2106.05(f)).
vi. “receiving the gradient values refined in row units from the second synaptic array and passing the portion of the received gradient values exceeding a threshold to the first synaptic array” (Receiving and passing the gradient values is considered insignificant extra-solution activity of mere data gathering – see MPEP § 2106.05(g)).
vii. “wherein the third synaptic array” (A synaptic array is considered a generic computer component being used as a tool to perform the judicial exception – see MPEP § 2106.05(f)).
viii. “passes the derived moving average value to the second synaptic array” (Passing the moving average value is considered insignificant extra-solution activity of mere data gathering –see MPEP § 2106.05(g)).
In step 2B of the 101-analysis set forth in the 2019 PEG, the examiner has determined that the claim does not include additional elements that are sufficient to amount significantly more than the judicial exception.
As discussed above, the additional elements iv, vi, and viii recite insignificant extra-solution activity of mere gathering, which is a well-understood routine and conventional activity, receiving or transmitting data over a network, e.g., using the Internet to gather data, see Symantec, 838 F.3d at 1321, 120 USPQ2d at 1362. In addition, additional elements i, ii, iii, v, and vii recite a generic computer component being used to perform the judicial exception, which are not indicative of significantly more.
Considering that the additional elements individually and in combination, and the claim as a whole, the additional elements do not provide significantly more than the abstract idea. Therefore, the claim is not patent eligible.
Claim 2:
Regarding claim 2, it is dependent upon claim 1, and thereby incorporates the limitations of, and corresponding analysis to claim 1. Further, claim 2 recites the following additional element:
“The device of claim 1, wherein the first synaptic array and the second synaptic array use analog array devices, and the third synaptic array is allocated on the digital domain.” (In step 2A, prong 2, this is considered mere instructions to apply an exception using a generic computer, see MPEP § 2106.05(f)). (In step 2B, this is also considered mere instructions to apply an exception using a generic computer, see MPEP § 2106.05(f)).
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible.
Claim 3:
Regarding claim 3, it is dependent upon claim 1, and thereby incorporates the limitations of, and corresponding analysis to claim 1. Further, claim 3 recites the following additional element:
i. “The device of claim 1, wherein, as a training process is repeated on the first synaptic array, the second synaptic array, and the third synaptic array, the second synaptic array converges to '0', and the gradient value becomes close to '0'.” (In step 2A, prong 2, this is considered mere instructions to apply an exception using a generic computer, see MPEP § 2106.05(f)). (In step 2B, this is also considered mere instructions to apply an exception using a generic computer, see MPEP § 2106.05(f)).
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible.
Claim 4:
Regarding claim 4, it is dependent upon claim 1, and thereby incorporates the limitations of, and corresponding analysis to claim 1. Further, claim 4 recites the following additional element:
“The device of claim 1, wherein, when the moving average value is passed to the second synaptic array, the moving average value is updated continuously to be set as an offset.” (In step 2A, prong 2, this is considered mere instructions to apply an exception using a generic computer, see MPEP § 2106.05(f)). (In step 2B, this is also considered mere instructions to apply an exception using a generic computer, see MPEP § 2106.05(f)).
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible.
Claim 5:
Regarding claim 5, it is dependent upon claim 1, and thereby incorporates the limitations of, and corresponding analysis to claim 1. Further, claim 5 recites the following additional element:
“wherein the moving average value is calculated in the form of adding an existing average value and a new value of the second synaptic array at a specific ratio, wherein the specific ratio is maintained constant or varied to adjust the degree of convergence.” (this is a mathematical concept, see MPEP §2106.04(a)(2)(I))
If claim limitations, under the broadest reasonable interpretation, covers performance of the limitations as a mathematical concept but for the recitation of generic computer components, then it falls within the mathematical concepts grouping of abstract ideas. Accordingly, the claim “recites” an abstract idea.
“The device of claim 1” (in step 2A, prong 2, this is considered mere instructions to apply an exception using a generic computer, see MPEP § 2106.05(f). In step 2B, this is also considered mere instructions to apply an exception using a generic computer, see MPEP § 2106.05(f).)
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible.
Claim 6:
Regarding claim 6, in step 1 of the 101-analysis set forth in MPEP 2106, the claim recites “an artificial neural network learning method,” and a method or process is one of the four statutory categories of invention.
In step 2A prong 1 of the 101-analysis set forth in MPEP 2106, the examiner has determined that the following limitations recite a process that, under the broadest reasonable interpretation, covers a mathematical concept but for recitation of generic computer components:
i. “deriving a moving average value by averaging accumulated values of the gradient values of the gradient values received from the second synaptic array” (this is a mathematical equation, see MPEP §2106.04(a)(2)(I)),
If claim limitations, under the broadest reasonable interpretation, covers performance of the limitations as a mathematical concept but for the recitation of generic computer components, then it falls within the mathematical concepts grouping of abstract ideas. Accordingly, the claim “recites” an abstract idea.
In step 2A prong 2 of the 101-analysis set forth in MPEP 2106, the examiner has determined that the following additional elements do not integrate this judicial exception into a practical application:
ii. “an artificial neural network learning method using a synaptic array device” (A method using a synaptic array device is considered mere instructions to apply the exception using a generic computer component – see MPEP § 2106.05(f)).
iii. “a first synaptic array and a second synaptic array using analog devices” (A synaptic array device is considered mere instructions to apply the exception using a generic computer component – see MPEP § 2106.05(f)).
iv. “a third synaptic array allocated on the digital domain” (A synaptic array device is considered mere instructions to apply the exception using a generic computer component – see MPEP § 2106.05(f)).
v. “passing the error gradient of the weights of the first synaptic array to the second synaptic array” (Passing the error gradients is considered insignificant extra-solution activity of mere data gathering – see MPEP § 2106.05(g)).
vi. “refining the error gradient in row units through the second synaptic array” (Refining the error gradient is considered insignificant extra-solution activity of mere data gathering – see MPEP § 2106.05(g)).
vii. “passing the corresponding gradient value to the third synaptic array” (Passing the gradient value is considered insignificant extra-solution activity of mere data gathering – see MPEP § 2106.05(g)).
viii. “passing the moving average value again to the second array” (Passing the moving average value is considered insignificant extra-solution activity of mere data gathering – see MPEP § 2106.05(g)).
In step 2B of the 101-analysis set forth in the 2019 PEG, the examiner has determined that the claim does not include additional elements that are sufficient to amount significantly more than the judicial exception.
As discussed above, the additional elements ii-iv recite mere instructions to apply the exception using a generic computer component. The additional elements v-viii recite insignificant extra-solution activity of mere gathering, which is a well-understood routine and conventional activity, receiving or transmitting data over a network, e.g., using the Internet to gather data, see Symantec, 838 F.3d at 1321, 120 USPQ2d at 1362.
Considering that the additional elements individually and in combination, and the claim as a whole, the additional elements do not provide significantly more than the abstract idea. Therefore, the claim is not patent eligible.
Claim 7:
Regarding claim 7, it is dependent upon claim 6, and thereby incorporates the limitations of, and corresponding analysis to claim 6. Further, claim 7 recites the following additional element:
“The method of claim 6, wherein the second synaptic array is initialized to a symmetry point or to a value different from the symmetry point.”
(In step 2A, prong 2, this is considered mere instructions to apply an exception using a generic computer, see MPEP § 2106.05(f)). (In step 2B, this is also considered mere instructions to apply an exception using a generic computer, see MPEP § 2106.05(f)).
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible.
Claim 8:
Regarding claim 7, it is dependent upon claim 6, and thereby incorporates the limitations of, and corresponding analysis to claim 6. Further, claim 7 recites the following additional element:
“The method of claim (6), wherein the moving average value is calculated by Eq. 1 below:
m
o
v
i
n
g
a
v
e
r
a
g
e
=
m
o
v
i
n
g
a
v
e
r
a
g
e
*
w
i
n
d
o
w
-
1
w
i
n
d
o
w
+
A
*
1
w
i
n
d
o
w
.” (this is a mathematical equation, see MPEP §2106.04(a)(2)(I))
If claim limitations, under the broadest reasonable interpretation, covers performance of the limitations as a mathematical concept but for the recitation of generic computer components, then it falls within the mathematical concepts grouping of abstract ideas. Accordingly, the claim “recites” an abstract idea.
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible.
Claim 9:
Regarding claim 9, it is dependent upon claim 8, and thereby incorporates the limitations of, and corresponding analysis to claim 8. Further, claim 9 recites the following additional element:
“The method of claim 8, wherein the window value of Eq. 1 is adjusted to control the degree of convergence, wherein the window value is constant or a value that varies continuously by using a moving average value or a function employing the values of the first and second synaptic arrays depending on epochs.” (this is a mathematical concept, see MPEP §2106.04(a)(2)(I))
If claim limitations, under the broadest reasonable interpretation, covers performance of the limitations as a mathematical concept but for the recitation of generic computer components, then it falls within the mathematical concepts grouping of abstract ideas. Accordingly, the claim “recites” an abstract idea.
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible.
Claim 10:
Regarding claim 10, it is dependent upon claim 6, and thereby incorporates the limitations of, and corresponding analysis to claim 6. Further, claim 10 recites the following additional element:
“The method of claim 6, wherein the passing of the corresponding gradient value to the third synaptic array updates the moving average value continuously to process the moving average value as an offset.”
(In step 2A, prong 2, this is considered mere instructions to apply an exception using a generic computer, see MPEP § 2106.05(f)). (In step 2B, this is also considered mere instructions to apply an exception using a generic computer, see MPEP § 2106.05(f)).
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible.
Claim 11:
Regarding claim 11, it is dependent upon claim 10, and thereby incorporates the limitations of, and corresponding analysis to claim 10. Furthermore, claim 11 recites the following additional element:
“The method of claim 10, wherein, whenever the update is performed, periodic attenuation is introduced, wherein the periodic attenuation defines a gamma parameter between 0 and 1 and multiplies the moving average value by the gamma parameter to reduce the moving average value.” (this is a mathematical concept, see MPEP §2106.04(a)(2)(I))
If claim limitations, under the broadest reasonable interpretation, covers performance of the limitations as a mathematical concept but for the recitation of generic computer components, then it falls within the mathematical concepts grouping of abstract ideas. Accordingly, the claim “recites” an abstract idea.
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 1, 5, 6, and 9 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claim 1 recites “the error gradient of the weights” in line 3. There is insufficient antecedent basis for this limitation in the claim. The examiner recommends changing the limitation to “an error gradient of the weights.”
Claim 5 recites the limitation "the degree of convergence" in lines 3-4. There is insufficient antecedent basis for this limitation in the claim. The examiner recommends changing the limitation to “a degree of convergence.”
Claim 6 recites the limitation “the error gradient of the weights” in line 5. There is insufficient antecedent basis for this limitation in the claim. The examiner recommends changing the limitation to “an error gradient of the weights.”
Claim 9 recites the limitation "the degree of convergence" in line 2. There is insufficient antecedent basis for this limitation in the claim. The examiner recommends changing the limitation to “a degree of convergence.”
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1 is rejected under 35 U.S.C. 103 as being unpatentable over Ambrogio et al. (US 20230306252 A1) (hereafter Ambrogio) in view of Lee et al. (https://www.frontiersin.org/journals/neuroscience/articles/10.3389/fnins.2021.767953/full) (hereafter Lee).
Claim 1
Regarding claim 1, Ambrogio teaches “A synaptic array device comprising a first synaptic array representing weight values;”
See Ambrogio in paragraph 0036 where it describes “a synaptic weight matrix which comprises synaptic weights that represent connection strengths between the neurons in one layer with the neurons in another layer.” Here, Ambrogio teaches an array of synaptic weight values.
Further, Ambrogio teaches “a second synaptic array;”
See Ambrogio in paragraph 0043 where it describes “the artificial neural network training process 130 will generate a plurality of trained synaptic weight matrices for a given artificial neural network which is trained in the digital domain, wherein each synaptic weight matrix comprises a matrix of trained (target) weight values W.sub.T.” Here, Ambrogio establishes that there is more than one synaptic array in the device.
Further, Ambrogio teaches “receiving the error gradient of the weights of the first synaptic array and representing gradient values refined in row units;”
See Ambrogio in paragraph 0039 where it describes “In general, in some embodiments, training an artificial neural network involves using a set of training data and performing a process of recursively adjusting the parameters/weights of the synaptic device arrays that connect the neuron layers, to fit the set of training data in order to maximize a likelihood function that minimizes error. The training process can be implemented using non-linear optimization techniques such as gradient-based techniques which utilize an error back-propagation process. For example, in some embodiments, a stochastic gradient descent (SGD) process is utilized to train artificial neural networks using the backpropagation method in which an error gradient with respect to each model parameter (e.g., weight) is calculated using the backpropagation algorithm.” Here, Ambrogio describes calculating an error gradient. Further see Ambrogio in paragraph 0063 where it describes “In some embodiments, the PWM circuitry and associated pulse driver circuitry of the peripheral circuitry 320 and 330 is configured to generate and apply PWM read pulses to the rows and columns of the array of RPU cells 310 in response to digital input vector values (read input values) that are received during different operations (e.g., programming operations, forward pass computations, etc.). In some embodiments, the PWM circuitry is configured to receive a digital input vector (to be applied to rows or columns) and convert the elements of the digital input vector into analog input vector values that are represented by input voltage voltages of varying pulse width. In some embodiments, a time-encoding scheme is used when input vectors are represented by fixed amplitude V.sub.IN=1V pulses with a tunable duration (e.g., pulse duration is a multiple of ins and is proportional to the value of the input vector). The input voltages applied to the rows (or columns) generate output MAC values on the columns (or rows) which are represented by output currents, wherein the output currents are processed by the readout circuitry.” Here, Ambrogio describes receiving gradients in row units.
Further, Ambrogio teaches “a third synaptic array;”
See Ambrogio in paragraph 0043 where it describes “the artificial neural network training process 130 will generate a plurality of trained synaptic weight matrices for a given artificial neural network which is trained in the digital domain, wherein each synaptic weight matrix comprises a matrix of trained (target) weight values W.sub.T.” Here, Ambrogio establishes that there is more than one synaptic array in the device.
Further, Ambrogio teaches “receiving the gradient values refined in row units from the second synaptic array”
See Ambrogio in paragraph 0039 where it describes “In general, in some embodiments, training an artificial neural network involves using a set of training data and performing a process of recursively adjusting the parameters/weights of the synaptic device arrays that connect the neuron layers, to fit the set of training data in order to maximize a likelihood function that minimizes error. The training process can be implemented using non-linear optimization techniques such as gradient-based techniques which utilize an error back-propagation process. For example, in some embodiments, a stochastic gradient descent (SGD) process is utilized to train artificial neural networks using the backpropagation method in which an error gradient with respect to each model parameter (e.g., weight) is calculated using the backpropagation algorithm.” Here, Ambrogio describes calculating an error gradient. Further see Ambrogio in paragraph 0063 where it describes “In some embodiments, the PWM circuitry and associated pulse driver circuitry of the peripheral circuitry 320 and 330 is configured to generate and apply PWM read pulses to the rows and columns of the array of RPU cells 310 in response to digital input vector values (read input values) that are received during different operations (e.g., programming operations, forward pass computations, etc.). In some embodiments, the PWM circuitry is configured to receive a digital input vector (to be applied to rows or columns) and convert the elements of the digital input vector into analog input vector values that are represented by input voltage voltages of varying pulse width. In some embodiments, a time-encoding scheme is used when input vectors are represented by fixed amplitude V.sub.IN=1V pulses with a tunable duration (e.g., pulse duration is a multiple of ins and is proportional to the value of the input vector). The input voltages applied to the rows (or columns) generate output MAC values on the columns (or rows) which are represented by output currents, wherein the output currents are processed by the readout circuitry.” Here, Ambrogio describes receiving gradients in row units.
However, Ambrogio did not explicitly teach “passing the portion of the received gradient values exceeding a threshold to the first synaptic array,” or “the third synaptic array derives a moving average value by averaging accumulated values of the gradient values received from the second synaptic array and passes the derived moving average value to the second synaptic array.”
However, Lee in the same field of art teaches “passing the portion of the received gradient values exceeding a threshold to the first synaptic array,”
See Lee in page 4, col 1 where it describes “Finally, Tiki-Taka algorithm allows the array A to participate in the forward pass by introducing a parameter,
γ
∈
0
,
1
:
w
i
j
=
γ
w
i
j
A
+
w
i
j
C
. Therefore, the effective weight
w
i
j
depends on
γ
: if
γ
=
1
,
w
i
j
A
+
w
i
j
C
replace the effective weight vector. Otherwise, if
γ
=
0
, then the effective weight is
w
i
j
=
w
i
j
C
“ Here, Lee teaches sending weight values back to another array if
γ
equals 1, in other words if they are greater than the threshold of 0.
Further, Lee in the same field of art teaches “the third synaptic array derives a moving average value by averaging accumulated values of the gradient values received from the second synaptic array and passes the derived moving average value to the second synaptic array.”
See Lee in page 4, col 1 where it describes “On the other hand, Tiki-Taka algorithm requires one additional array, namely A, and it stores
∆
W
by accumulating gradient vectors,
∇
L
. The weight vectors stored in the array A are denoted as WA and the array, C, stores the weight vectors denoted as WC. When compared to SGD, Tiki-Taka algorithm accumulates the gradient
∇
i
j
L
, in
w
i
j
A
. Then, the accumulated gradient is transferred to update the weight,
w
i
j
C
, at every ns steps.” See also Lee in page 9 figure 6 where it describes “Raw data of t(k) and WA was averaged using exponential moving average and approximated with natural smoothing spline method.” Here, Lee establishes an array of weight vectors stored in the array A as WA that is averaged with an accumulated moving average. The values are then transferred to update the weight.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the base reference of Ambrogio with the teachings of Lee by using Ambrogio’s teachings of a synaptic array device consisting of a synaptic array representing weight values, a synaptic array receiving an error gradient, and a synaptic array receiving gradient values, and incorporate with Lee’s teachings of transferring values that meet a criterion to another synaptic array and deriving a moving average value.
One of ordinary skill in the art would be motivated to do so because by integrating Lee’s frameworks into the methods of Ambrogio, one with ordinary skill in the art would be able to “resolve the performance degradation issue in neural network training caused by the update asymmetry in the crosspoint elements” (Lee, pages 3-4).
Claim 2 and 6 are rejected under 35 U.S.C. 103 as being unpatentable over Ambrogio in view of Lee, and further in view of Birdwell et al. (US 20150106316 A1) (hereinafter Birdwell).
Claim 2
Regarding claim 2, Ambrogio in view of Lee teaches the limitations in claim 1.
Ambrogio teaches “the third synaptic array is allocated on the digital domain.”
See Ambrogio in paragraph 0035 where it describes “the artificial neural network training process 130 implements methods for training an artificial neural network model in the digital domain. The artificial neural network model can be any type of neural network including, but not limited to, a feed-forward neural network (e.g., a Deep Neural Network (DNN), a Convolutional Neural Network (CNN), etc.), a Recurrent Neural Network (RNN) (e.g., a Long Short-Term Memory (LSTM) neural network), etc.”
Neither Ambrogio or Lee appear to teach “the first synaptic array and the second synaptic array use analog array devices.”
However, Birdwell in the same field of art teaches “the first synaptic array and the second synaptic array use analog array devices.”
See Birdwell in paragraph 0172 where it describes “Referring again to FIG. 10A, in one embodiment, the FPGA DANNA array of circuit elements connects to, for example, a PCIe interface 1010 that is used for external programming and adaptive "learning" algorithms that may monitor and control the configuration and characteristics of the network and may have array elements 1021, 1023, 1031, and 1033 that may be located on an edge of the array and may have external inputs or outputs. Array elements, including 1021, 1023, 1031, and 1033 may preferably also have inputs or outputs that are internal to the array. Array circuit elements may be digital or analog, but as will be seen in the discussions of FIGS. 9A and 9B, the FPGA implementation of a circuit element selectively being a neuron and a synapse comprises mostly digital circuit components such as registers of digital data. Analog circuit elements that may be used in constructing a circuit element of an array (besides the implementations shown in FIGS. 9A and 9B by way of example) include a memristor, a capacitor, inductive device such as a relay, an optic device, a magnetic device and the like which may implement, for example, a memory or storage. "Analog signal storage device" as used in the specification and claims shall mean any analog storage device that is known in the art including but not limited to memristor, phase change memory, spin transport electronic device, optical storage, capacitive storage, inductive storage (e.g. a relay), magnetic storage or other known analog memory device.”
Ambrogio, Lee and Birdwell are both considered to be analogous to the claimed invention because they are in the same field of artificial neural network devices that store weight values. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Ambrogio and Lee to incorporate the teachings of Birdwell and incorporate both analog and digital devices in the synaptic array device.
One of ordinary skill in the art would be motivated to do so because by integrating Birdwell’s frameworks into the methods of Ambrogio and Lee, one with ordinary skill in the art would be able to build “an improved neuroscience-inspired network architecture which overcomes the problems exhibited by known architectures for real-time monitoring of array elements.” (Birdwell, paragraph 0081).
Claim 6
Regarding claim 6, Ambrogio teaches “In an artificial neural network learning method using a synaptic array device comprising… a third synaptic array allocated on the digital domain;”
See Ambrogio in paragraph 0034 where it describes “In some embodiments, as shown in FIG. 1A, the digital processing system 110 performs various processes including, but not limited to, an artificial neural network training process 130, a neural core configuration process 132, an analog crossbar array calibration process 134, and an inference/classification process 136. Further, in some embodiments, as shown in FIG. 1B, the analog crossbar array calibration process 134 comprises a first calibration process 134-1, a second calibration process 134-2, and a third calibration process 134-3. As explained in further detail below, the analog crossbar array calibration process 134 implements methods for calibrating the analog RPU hardware of the neural cores 122 to reduce hardware computation errors that arise due to the programming errors and non-idealities of the analog RPU hardware.” Further see Ambrogio in paragraph 0035 where it describes “the artificial neural network training process 130 implements methods for training an artificial neural network model in the digital domain. The artificial neural network model can be any type of neural network including, but not limited to, a feed-forward neural network (e.g., a Deep Neural Network (DNN), a Convolutional Neural Network (CNN), etc.), a Recurrent Neural Network (RNN) (e.g., a Long Short-Term Memory (LSTM) neural network), etc.” Here, Ambrogio describes an artificial neural network learning method allocated on the digital domain.
Further Ambrogio teaches “passing the error gradient of the weights of the first synaptic array to the second synaptic array;”
See Ambrogio in paragraph 0039 where it describes “In general, in some embodiments, training an artificial neural network involves using a set of training data and performing a process of recursively adjusting the parameters/weights of the synaptic device arrays that connect the neuron layers, to fit the set of training data in order to maximize a likelihood function that minimizes error. The training process can be implemented using non-linear optimization techniques such as gradient-based techniques which utilize an error back-propagation process. For example, in some embodiments, a stochastic gradient descent (SGD) process is utilized to train artificial neural networks using the backpropagation method in which an error gradient with respect to each model parameter (e.g., weight) is calculated using the backpropagation algorithm.” Here, Ambrogio describes a process for calculating error gradient. Further see Ambrogio in paragraph 0040 where it cites “The forward pass processes input data in a forward direction (from the input layer to the output layer) through the layers of the network, and generates predictions and calculates errors between the predictions and the ground truth. The backward pass backpropagates errors in a backward direction (from the output layer to the input layer) through the artificial neural network to obtain gradients to update model weights.” Here, Ambrogio describes passing the error through synaptic arrays.
Further Ambrogio teaches “refining the error gradient in row units through the second synaptic array;”
See Ambrogio in paragraph 0047 where it describes “re-programming the weight values of the weight matrix stored in the analog RPU array, until a convergence criterion is met for each output line in which a difference (error err) between a target offset value, and an actual offset of the given output line does not exceed an error threshold value E.” Here, Ambrogio cites refining the values in the synaptic array.
However, Ambrogio did not explicitly teach “passing the error gradient of the weights of the first synaptic array to the second synaptic array” or “passing the corresponding gradient value to the third synaptic array” or “deriving a moving average value by averaging accumulated values of the gradient values received from the second synaptic array to the third synaptic array and passing the moving average value again to the third synaptic array.”
However, Lee teaches “passing the error gradient of the weights of the first synaptic array to the second synaptic array;”
See Lee in page 2, col 2 where it describes “Stochastic gradient descent (SGD) is a widely used, first-order optimization method for training neural networks. During forward pass, a neural network infers output with expected values. Based on the loss, gradients of weights in a neural network are calculated during backward pass and later updated during the weight update phase by error back-propagation algorithm. The amount of weight update is proportional to the calculated gradient with the scaling factor called learning rate,
η
:
w
i
j
←
w
i
j
-
η
∇
i
j
L
=
w
i
j
-
η
x
δ
j
, where
w
i
j
indicates a weight from pre-synaptic neuron i to post-synaptic neuron j, L is the loss of a neural network, and
∇
i
j
,
k
L
indicates the derivative of L with respect to
w
i
j
.
x
i
is the input activation of pre-synaptic neuron i, and
δ
j
is the error of post-synaptic neurons.” Here, Lee teaches calculating error gradient so that it can be transferred to another array.
Further Lee teaches “passing the corresponding gradient value to the third synaptic array;”
See Lee in page 2, col 2 where it describes “Stochastic gradient descent (SGD) is a widely used, first-order optimization method for training neural networks. During forward pass, a neural network infers output with expected values. Based on the loss, gradients of weights in a neural network are calculated during backward pass and later updated during the weight update phase by error back-propagation algorithm. The amount of weight update is proportional to the calculated gradient with the scaling factor called learning rate,
η
:
w
i
j
←
w
i
j
-
η
∇
i
j
L
=
w
i
j
-
η
x
δ
j
, where
w
i
j
indicates a weight from pre-synaptic neuron i to post-synaptic neuron j, L is the loss of a neural network, and
∇
i
j
,
k
L
indicates the derivative of L with respect to
w
i
j
.
x
i
is the input activation of pre-synaptic neuron i, and
δ
j
is the error of post-synaptic neurons.” Here, Lee teaches calculating error gradient so that it can be transferred to another array.
Further Lee teaches “deriving a moving average value by averaging accumulated values of the gradient values received from the second synaptic array to the third synaptic array and passing the moving average value again to the second synaptic array.”
See Lee in page 4, col 1 where it describes “On the other hand, Tiki-Taka algorithm requires one additional array, namely A, and it stores
∆
W
by accumulating gradient vectors,
∇
L
. The weight vectors stored in the array A are denoted as WA and the array, C, stores the weight vectors denoted as WC. When compared to SGD, Tiki-Taka algorithm accumulates the gradient
∇
i
j
L
, in
w
i
j
A
. Then, the accumulated gradient is transferred to update the weight,
w
i
j
C
, at every ns steps.” See also Lee in page 9 figure 6 where it describes “Raw data of t(k) and WA was averaged using exponential moving average and approximated with natural smoothing spline method.” Here, Lee establishes an array of weight vectors stored in the array A as WA that is averaged with an accumulated moving average. The values are then transferred to update the weight.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the base reference of Ambrogio with the teachings of Lee by using Ambrogio’s teachings of a synaptic array device consisting of a synaptic array representing weight values, a synaptic array receiving an error gradient, and a synaptic array receiving gradient values, and incorporate with Lee’s teachings of transferring values that meet a criterion to another synaptic array and deriving a moving average value.
One of ordinary skill in the art would be motivated to do so because by integrating Lee’s frameworks into the methods of Ambrogio, one with ordinary skill in the art would be able to “resolve the performance degradation issue in neural network training caused by the update asymmetry in the crosspoint elements” (Lee, page 3 col 2, page 4 col 1).
Neither Ambrogio or Lee appear to teach “the first synaptic array and the second synaptic array use analog array devices.”
However, Birdwell in the same field of art teaches “the first synaptic array and the second synaptic array use analog array devices.”
See Birdwell in paragraph 0172 where it describes “Referring again to FIG. 10A, in one embodiment, the FPGA DANNA array of circuit elements connects to, for example, a PCIe interface 1010 that is used for external programming and adaptive "learning" algorithms that may monitor and control the configuration and characteristics of the network and may have array elements 1021, 1023, 1031, and 1033 that may be located on an edge of the array and may have external inputs or outputs. Array elements, including 1021, 1023, 1031, and 1033 may preferably also have inputs or outputs that are internal to the array. Array circuit elements may be digital or analog, but as will be seen in the discussions of FIGS. 9A and 9B, the FPGA implementation of a circuit element selectively being a neuron and a synapse comprises mostly digital circuit components such as registers of digital data. Analog circuit elements that may be used in constructing a circuit element of an array (besides the implementations shown in FIGS. 9A and 9B by way of example) include a memristor, a capacitor, inductive device such as a relay, an optic device, a magnetic device and the like which may implement, for example, a memory or storage. "Analog signal storage device" as used in the specification and claims shall mean any analog storage device that is known in the art including but not limited to memristor, phase change memory, spin transport electronic device, optical storage, capacitive storage, inductive storage (e.g. a relay), magnetic storage or other known analog memory device."”
Ambrogio, Lee and Birdwell are both considered to be analogous to the claimed invention because they are in the same field of artificial neural network devices that store weight values. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Ambrogio and Lee to incorporate the teachings of Birdwell and incorporate both analog and digital devices in the synaptic array device.
One of ordinary skill in the art would be motivated to do so because by integrating Birdwell’s frameworks into the methods of Ambrogio and Lee, one with ordinary skill in the art would be able to build “an improved neuroscience-inspired network architecture which overcomes the problems exhibited by known architectures for real-time monitoring of array elements.” (Birdwell, paragraph 0081).
Claim 3 is rejected under 35 U.S.C. 103 as being unpatentable over Ambrogio in view of Lee, and further in view of David et al. (US 11055617 B1) (hereinafter David).
Claim 3
Regarding claim 3, Ambrogio in view of Lee teaches the limitations in claim 1.
Ambrogio teaches “as a training process is repeated on the first synaptic array, the second synaptic array, and the third synaptic array,”
See Ambrogio in paragraph 0043 where it describes “the artificial neural network training process 130 will generate a plurality of trained synaptic weight matrices for a given artificial neural network which is trained in the digital domain, wherein each synaptic weight matrix comprises a matrix of trained (target) weight values W.sub.T.” Here, Ambrogio establishes that there is more than one synaptic array in the device. Further see Ambrogio in paragraph 0035 where it describes “the artificial neural network training process 130 implements methods for training an artificial neural network model in the digital domain. The artificial neural network model can be any type of neural network including, but not limited to, a feed-forward neural network (e.g., a Deep Neural Network (DNN), a Convolutional Neural Network (CNN), etc.), a Recurrent Neural Network (RNN) (e.g., a Long Short-Term Memory (LSTM) neural network), etc. In general, an artificial neural network comprises a plurality of layers (neuron layers), wherein each layer comprises multiple neurons. The neuron layers include an input layer, an output layer, and one or more hidden model layers between the input and output layers, wherein the number of neuron layer and configuration of the neuron layers (e.g., number of constituent artificial neurons) will vary depending on the type of neural network that is implemented.” Here Ambrogio teaches a training process that is repeated on an artificial neural network with multiple synaptic arrays (a plurality of trained synaptic weight matrices).
Further Ambrogio teaches “the second synaptic array converges”
See Ambrogio in paragraph 0040 where it teaches “As is known in the art, a backpropagation process comprises three repeating processes including (i) a forward process, (ii) a backward process, and (iii) a model parameter update process. During the training process, training data are randomly sampled into mini-batches, and the mini-batches are input to the artificial neural network to traverse the model in two phases: forward and backward passes. The forward pass processes input data in a forward direction (from the input layer to the output layer) through the layers of the network, and generates predictions and calculates errors between the predictions and the ground truth. The backward pass backpropagates errors in a backward direction (from the output layer to the input layer) through the artificial neural network to obtain gradients to update model weights. The forward and backward cycles mainly involve performing matrix-vector multiplication operations in forward and backward directions. The weight update involves performing incremental weight updates for weight values of the synaptic weight matrices of the artificial neural network being trained. The processing of a given mini-batch via the forward and backward phases is referred to as an iteration, and an epoch is defined as performing the forward-backward pass through an entire training dataset. The training process iterates multiple epochs until the model converges to given convergence criterion.” Here, Ambrogio teaches iterating the process until the model converges to a convergence criterion.
Neither Ambrogio or Lee appear to teach “the second synaptic array converges to ‘0’, and the gradient value becomes close to ‘0’.”
However, David in the same field of art teaches “the second synaptic array converges to ‘0’, and the gradient value becomes close to ‘0’.”
See David in col. 10, lines 54-67, col. 11, lines 1-2 where it teaches “Additionally or alternatively, because partial-activation is only an approximation of the full neural network, partial-activation may be used in an initial stage (e.g., the first P iterations, or until the output converges during predictions or the error converges during training to within a threshold) and thereafter the fully-activated neural network may be run (e.g., the next or last Q iterations, or until the output converges or the error converges to zero) to confirm the initial partial approximation. Additionally or alternatively, partial activation may be used in specific layers (e.g., deeper middle layers which often have a less direct effect on the final result), while a fully connected network may be used for the remaining layers (e.g., layers near the input and output layers which often have a greater direct effect on the final result). Other hybrid combinations of partial and fully activated neural networks may be used.” Here, David teaches doing multiple iterations over a neural network until either the array converges to zero or the error converges to zero.
Ambrogio, Lee, and David are both considered to be analogous to the claimed invention because they are in the same field of artificial neural network devices that store weight values. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Ambrogio and Lee to incorporate the teachings of Gardner and iterate the process on the synaptic array devices until they converge to 0 or something close to it. Doing this will minimize the margin of error.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the base reference of Ambrogio and Lee with the teachings of David by using Ambrogio and Lee’s teachings of a synaptic array device that iterates until reaching a convergence criterion, and incorporate with Gardner’s teachings of setting the convergence criterion to zero. One of ordinary skill in the art would be motivated to do so because by integrating David’s frameworks into the methods of Ambrogio and Lee, one with ordinary skill in the art would be able to minimize the ensure that “the error is minimized or converges” (David, col. 1, lines 38-39).
Claims 4 and 5 are rejected under 35 U.S.C. 103 as being unpatentable over Ambrogio in view of Lee, and further in view of Marius (https://stackoverflow.com/questions/11074665/calculate-cumulative-average-mean/17882721#17882721).
Claim 4
Regarding claim 4, Ambrogio in view of Lee teaches the limitations in claim 1.
Lee teaches “when the moving average value is passed to the second synaptic array,”
See Lee in page 4, col 1 where it describes “On the other hand, Tiki-Taka algorithm requires one additional array, namely A, and it stores
∆
W
by accumulating gradient vectors,
∇
L
. The weight vectors stored in the array A are denoted as WA and the array, C, stores the weight vectors denoted as WC. When compared to SGD, Tiki-Taka algorithm accumulates the gradient
∇
i
j
L
, in
w
i
j
A
. Then, the accumulated gradient is transferred to update the weight,
w
i
j
C
, at every ns steps.” See also Lee in page 9 figure 6 where it describes “Raw data of t(k) and WA was averaged using exponential moving average and approximated with natural smoothing spline method.” Here, Lee establishes an array of weight vectors stored in the array A as WA that is averaged with an accumulated moving average. The values are then transferred to update the weight.
Neither Ambrogio or Lee teach “the moving average value is updated continuously to be set as an offset.”
However, Marius in the same field of art teaches “the moving average value is updated continuously to be set as an offset.”
See Marius in page 2 where it teaches “In analogy to the cumulative sum of a list I propose this: The cumulative average avg of a vector x would contain the averages from 1st position till position i. One method is just to compute the mean for each position by summing over all previous values and dividing by their number. By rewriting the definition of the arithmetic mean as a recursive formula. One gets avg(1) = x(1) and avg(i) = (i-1)/i*avg(i-1) + x(i)/I; (i>1) Evaluating this expression for every element of your vector (or list, one-dimensional array or however you call it) gives you the cumulative average.” Here, Marius describes a cumulative average, which moves as new data is added, and is therefore a moving average. This cumulative average is continuously updated by adding previous values.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the base reference of Ambrogio and Lee with the teachings of Marius by using Ambrogio’s teachings of a synaptic array device and incorporate with Lee’s teachings of incorporating a moving average value, and further Marius’s teachings of updating the moving average value.
One of ordinary skill in the art would be motivated to do so because by integrating Marius’s frameworks into the methods of Ambrogio and Lee, one with ordinary skill in the art would provide a useful tool “if you have to calculate an average over very large or very many integers and would run into an overflow if you had to store their cumulative sum” (Marius, page 2) in order to accommodate continuously updating the moving average.
Claim 5
Regarding claim 5, Ambrogio in view of Lee teaches the limitations in claim 1.
Neither Ambrogio or Lee teach “the moving average value is calculated in the form of adding an existing average value and a new value of the second synaptic array at a specific ratio, wherein the specific ratio is maintained constant or varied to adjust the degree of convergence.”
However, Marius in the same field of art teaches “the moving average value is calculated in the form of adding an existing average value and a new value of the second synaptic array at a specific ratio, wherein the specific ratio is maintained constant or varied to adjust the degree of convergence.”
See Marius in page 2 where it teaches “In analogy to the cumulative sum of a list I propose this: The cumulative average avg of a vector x would contain the averages from 1st position till position i. One method is just to compute the mean for each position by summing over all previous values and dividing by their number. By rewriting the definition of the arithmetic mean as a recursive formula. One gets avg(1) = x(1) and avg(i) = (i-1)/i*avg(i-1) + x(i)/I; (i>1) Evaluating this expression for every element of your vector (or list, one-dimensional array or however you call it) gives you the cumulative average.” Here, Marius describes updating the cumulative (moving) average value by adding an existing average value and a new value from the array x() (which in this case represents the values in the synaptic array). The specific ratio will be
1
i
to
i
-
1
i
, as determined by the equation.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the base reference of Ambrogio and Lee with the teachings of Marius by using Ambrogio’s teachings of a synaptic array device and incorporate with Lee’s teachings of incorporating a moving average value, and further Marius’s teachings of deriving the moving average value.
One of ordinary skill in the art would be motivated to do so because by integrating Marius’s frameworks into the methods of Ambrogio and Lee, one with ordinary skill in the art would provide a useful tool “if you have to calculate an average over very large or very many integers and would run into an overflow if you had to store their cumulative sum” (Marius, page 2) in order to accommodate continuously updating the moving average.
Claim 7 is rejected under 35 U.S.C. 103 as being unpatentable over Ambrogio in view of Lee in view of Birdwell and further in view of Kim et al. (https://arxiv.org/abs/1907.10228) (hereinafter Kim).
Claim 7
Regarding claim 7, Ambrogio in view of Lee and Birdwell teaches the limitations in claim 6.
Neither Ambrogio, Lee, or Birdwell teach “where the second synaptic array is initialized to a symmetry point or to a value different from the symmetry point.”
However, Kim in the same field of art teaches “the second synaptic array is initialized to a symmetry point or to a value different from the symmetry point.”
See Kim in page 10 where it teaches “In the Soft-Bound model, there is one point where the graphs of potentiation and depression cross, as shown in Fig. 2d. We call the crossing point the symmetry point since the absolute values of
∆
w
0
+
and
∆
w
0
-
are the same in this point. If the Soft-Bound model has a symmetry point it is always unique. The reason why the symmetry point is important is that the conductance state of the Soft-Bound device tends to converge to the symmetry point with random programming pulses: When the current weight is stronger than the symmetry point, potentiation is stronger than depression. Therefore, the weight tends to increase towards the symmetry point. In contrast, when the weight is larger than the symmetry point, depression is stronger than potentiation resulting in an overall decrease of the weight. Hence, if the device is updated randomly, the state of the device will converge to the symmetry point.” Here, Kim defines a symmetry point to which the device is intended to converge.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the base reference of Ambrogio, Lee, and Birdwell with the teachings of Kim by using Ambrogio’s teachings of a synaptic array device and incorporate with Lee’s teachings of incorporating a moving average value, and further Kim’s teachings of defining a symmetry point and initializing an array to it.
One of ordinary skill in the art would be motivated to do so because by integrating Kim’s frameworks into the teachings of Amborogio, Lee, and Birdwell, one would be able to ensure that “network performance dramatically improves for imbalanced synapse devices” (Kim, page 3).
Claims 8 and 10 are rejected under 35 U.S.C. 103 as being unpatentable over Ambrogio in view of Lee and Birdwell, and further in view of Marius.
Claim 8
Regarding claim 8, Ambrogio in view of Lee and Birdwell teaches the limitations in claim 6.
Neither Ambrogio, Lee, or Birdwell teach “the method of claim (6), wherein the moving average value is calculated by Eq. 1 below:
m
o
v
i
n
g
a
v
e
r
a
g
e
=
m
o
v
i
n
g
a
v
e
r
a
g
e
*
w
i
n
d
o
w
-
1
w
i
n
d
o
w
+
A
*
1
w
i
n
d
o
w
”
However, Marius in the same field of art teaches “the method of claim (6), wherein the moving average value is calculated by Eq. 1 below:
m
o
v
i
n
g
a
v
e
r
a
g
e
=
m
o
v
i
n
g
a
v
e
r
a
g
e
*
w
i
n
d
o
w
-
1
w
i
n
d
o
w
+
A
*
1
w
i
n
d
o
w
”
See Marius in page 2 where it teaches “In analogy to the cumulative sum of a list I propose this: The cumulative average avg of a vector x would contain the averages from 1st position till position i. One method is just to compute the mean for each position by summing over all previous values and dividing by their number. By rewriting the definition of the arithmetic mean as a recursive formula. One gets avg(1) = x(1) and avg(i) = (i-1)/i*avg(i-1) + x(i)/I; (i>1) Evaluating this expression for every element of your vector (or list, one-dimensional array or however you call it) gives you the cumulative average.” Here, Marius describes an equation for how to calculate a cumulative average; an average that “moves” with the new data, or in other words, a moving average. In the equation, avg(x) represents an existing moving average. The range of averages from the position to position i would represent a “window” in which i is the highest position in the vector and is equivalent to a window value, and x(i) is a constant equivalent to A. An existing moving average is multiplied by (i-1)/I, to which x(i)/(i) is added.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the base reference of Ambrogio, Lee, and Birdwell with the teachings of Marius by using Ambrogio’s teachings of a synaptic array device and incorporate with Lee’s teachings of incorporating a moving average value, and further Marius’s teachings of deriving the moving average value.
One of ordinary skill in the art would be motivated to do so because by integrating Marius’s frameworks into the methods of Ambrogio, Lee, and Birdwell, one with ordinary skill in the art would provide a useful tool “if you have to calculate an average over very large or very many integers and would run into an overflow if you had to store their cumulative sum” (Marius, page 2) in order to accommodate the moving average.
Claim 10
Regarding claim 10, Ambrogio in view of Lee and Birdwell teaches the limitations in claim 6.
Lee teaches “wherein the passing of the corresponding gradient value to the third synaptic array”
See Lee in page 2, col 2 where it describes “Stochastic gradient descent (SGD) is a widely used, first-order optimization method for training neural networks. During forward pass, a neural network infers output with expected values. Based on the loss, gradients of weights in a neural network are calculated during backward pass and later updated during the weight update phase by error back-propagation algorithm. The amount of weight update is proportional to the calculated gradient with the scaling factor called learning rate,
η
:
w
i
j
←
w
i
j
-
η
∇
i
j
L
=
w
i
j
-
η
x
δ
j
, where
w
i
j
indicates a weight from pre-synaptic neuron i to post-synaptic neuron j, L is the loss of a neural network, and
∇
i
j
,
k
L
indicates the derivative of L with respect to
w
i
j
.
x
i
is the input activation of pre-synaptic neuron i, and
δ
j
is the error of post-synaptic neurons.” Here, Lee teaches calculating error gradient so that it can be transferred to another array.
Neither Ambrogio, Lee, or Birdwell teach “updates the moving average continuously to process the moving average value as an offset.”
However, Marius in the same field of art teaches “updates the moving average continuously to process the moving average value as an offset.”
See Marius in page 2 where it teaches “In analogy to the cumulative sum of a list I propose this: The cumulative average avg of a vector x would contain the averages from 1st position till position i. One method is just to compute the mean for each position by summing over all previous values and dividing by their number. By rewriting the definition of the arithmetic mean as a recursive formula. One gets avg(1) = x(1) and avg(i) = (i-1)/i*avg(i-1) + x(i)/I; (i>1) Evaluating this expression for every element of your vector (or list, one-dimensional array or however you call it) gives you the cumulative average.” Here, Marius describes a cumulative average, which moves as new data is added, and is therefore a moving average. This cumulative average is continuously updated by adding previous values.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the base reference of Ambrogio, Lee, and Birdwell with the teachings of Marius by using Ambrogio’s teachings of a synaptic array device and incorporate with Lee’s teachings of incorporating a moving average value, and further Marius’s teachings of updating the moving average value.
One of ordinary skill in the art would be motivated to do so because by integrating Marius’s frameworks into the methods of Ambrogio, Lee, and Birdwell, one with ordinary skill in the art would provide a useful tool “if you have to calculate an average over very large or very many integers and would run into an overflow if you had to store their cumulative sum” (Marius, page 2) in order to accommodate continuously updating the moving average.
Claim 9 is rejected under 35 U.S.C. 103 as being unpatentable over Ambrogio in view of Lee, Birdwell, and Marius, and further in view of Kowalczyk (https://www.researchgate.net/profile/Kamil-Kowalczyk-3/publication/332655133_Changes_in_mean_sea_level_on_the_coast_of_Baltic_Sea_on_tide_gouge_data_from_years_1811_2015/links/5cd02ae7458515712e95a7f2/Changes-in-mean-sea-level-on-the-coast-of-Baltic-Sea-on-tide-gouge-data-from-years-1811-2015.pdf?_sg%5B0%5D=started_experiment_milestone&origin=journalDetail).
Claim 9
Regarding claim 9, Ambrogio in view of Lee, Birdwell, and Marius teaches the limitations in claim 8.
Neither Ambrogio, Lee, Birdwell, or Marius teach “wherein the window value of Eq. 1 is adjusted to control the degree of convergence, wherein the window value is constant or a value that varies continuously by using a moving average value or a function employing the values of the first and second synaptic arrays depending on epochs.”
However, Kowalczyk in the same field of art teaches “wherein the window value of Eq. 1 is adjusted to control the degree of convergence, wherein the window value is constant or a value that varies continuously by using a moving average value or a function employing the values of the first and second synaptic arrays depending on epochs.”
See Kowalczyk in page 196, col 1 where it teaches “The most common methods for smoothing data include the moving average and median methods. The moving average method is more sensitive to gross errors. In this method, each element of the series is replaced by a weighted average of the neighboring elements, the number of which is determined by the “window”. The “window” of 11 years old is often used. An alternative is to adopt a period from the Fourier analysis or another “window” value, e.g. equal to the longest period occurring in observation data, to approx. 19 years (18.6 years is the duration of the period of lunar node precession.” Here, Kowalczyk teaches using a window value when calculating the moving average, in which the window is either a constant (such as 11 years) or another value in the data. Further see Kowalczyk in page 207, col 2 where it teaches “After smoothing the moving average (19-year), the obtained standard deviation decreased significantly. The trend line obtained fits very well to the data.” Here, Kowalczyk teaches that when the window value is adjusted, the data converges.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to combine the base reference of Ambrogio, Lee, Birdwell, and Marius with the teachings of Kowalczyk by using Ambrogio’s teachings of a synaptic array device, and incorporate with Lee’s teaching of incorporating a moving average value, Birdwell’s teaching of using analog array devices, Marius’s teachings of calculating a moving average, and further Kowalczyk’s teaching of a window value.
One of ordinary skill in the art would be motivated to do so because by integrating Kowalczyk’s frameworks into the methods of Ambrogio, Lee, Birdwell, and Marius, one of ordinary skill in the art would be able to “minimize the impact of these errors on the final result” (Kowalczyk, page 196, col 1).
Claim 11 is rejected under 35 U.S.C. 103 as being unpatentable over Ambrogio in view of Lee, Birdwell, and Marius, and further in view of Hansun (https://ieeexplore.ieee.org/abstract/document/6708545).
Claim 11
Regarding claim 11, Ambrogio in view of Lee, Birdwell, and Marius teaches the limitations in claim 10.
Neither Ambrogio, Lee, Birdwell, or Marius teach “whenever the update is performed, periodic attenuation is introduced, wherein the periodic attenuation defines a gamma parameter between 0 and 1 and multiplies the moving average value by the gamma parameter to reduce the moving average value.”
However, Hansun in the same field of art teaches “whenever the update is performed, periodic attenuation is introduced, wherein the periodic attenuation defines a gamma parameter between 0 and 1 and multiplies the moving average value by the gamma parameter to reduce the moving average value.”
See Hansun in pages 1, col 2 and page 2, col 1 where it teaches “An Exponential Moving Average (EMA) is a type of WMA which assigns a weighting factor to each value in the data series according to its age. Like WMA, in EMA the most recent data gets the greatest weight and each data value gets a smaller weight as we go back chronologically. But unlike WMA, in EMA the weighting for each older data point decreases exponentially, so its’ never reaching zero. The EMA for a time series can be calculated recursively as
S
1
=
Y
1
, for
t
>
1
,
S
t
=
a
∙
Y
t
+
1
-
α
∙
S
t
-
1
, where
Y
t
is the value at time period t, and
α
represents the degree of weighting decrease, a constant smoothing factor between 0 and 1.” Here, Hansun teaches periodically multiplying the moving average value by a factor between 0 and 1.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the base reference of Ambrogio, Lee, Birdwell, and Marius with the teachings of Hansun by using Ambrogio, Lee, Birdwell, and Marius’s teachings of a synaptic array device incorporating a moving average value, and further Hansun’s teachings of setting a parameter to reduce the moving average value.
One of ordinary skill in the art would be motivated to do so because by integrating Hansun’s frameworks into the methods of Ambrogio and Lee, one with ordinary skill in the art would be able to have “an improvement of SMA that gives a weighting factor for each point” (Hansun, page 1, col 1) in which “MSE (mean square error) and MAPE (mean absolute percentage error) values for the proposed method are quite small” (Hansun, page 3, col 2).
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ANDREW CHARLES YORKS whose telephone number is (571)270-1803. The examiner can normally be reached M-F, 9am to 5pm ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Usmaan Saeed can be reached at (571) 272-4046. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/ANDREW CHARLES YORKS/Examiner, Art Unit 2146
/SHOURJO DASGUPTA/Primary Examiner, Art Unit 2144