DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
This action is in response to communications filed on 05/08/2026.
Claims 7 and 14 have been canceled.
Claims 21-22 have been added.
Claims 1-6, 8-13, and 15-22 are pending and have been examined.
Information Disclosure Statement
The information disclosure statement (IDS) submitted was filed on 03/12/2026. The submission is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Response to Arguments
Previous rejections under 35 USC 112 have been withdrawn in view of amendments.
Applicant’s arguments with respect to the newly amended features have been considered but are moot in view of new grounds of rejection. However, it is noted that Cherupally does not only describe a forward pass as alleged by applicant. As can be seen in paragraph 41 and figure 2A of Cherupally, the loss function is used to update the weights by a backward pass during “IMC hardware noise-aware training”, which means that the loss function integrates the injected noise described by Cherupally, but, in the interest of advancing prosecution, see Bunandar et al. (US 20220172052 A1) below.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 1-3, 15-18, and 21-22 are rejected under 35 U.S.C. 103 as being unpatentable over Stevens et al. (US 20220067513 A1) in view of Bunandar et al. (US 20220172052 A1) and Cherupally et al. (US 20220318628 A1).
As per independent claim 1, Stevens teaches a system, comprising:
at least one processor circuit (e.g. in paragraphs 42-43, “processor”); and
at least one memory that stores program code configured to be executed by the at least one processor circuit (e.g. in paragraphs 42-43, “non-transitory machine readable media comprising machine-executable instructions”), the program code comprising:
a neural network model trainer configured to: receive a configuration file that specifies characteristics of hardware including multiply-and-accumulation circuits utilized to implement nodes of a particular layer of a neural network (e.g. in paragraphs 32, 34, 42, 45, 88 and 90-92, ““Neural network” refers to an algorithm or computational system based on a collection of connected units or nodes called artificial neurons… neural networks…acquire differences during training (e.g., be trained to have different weights from one another)… comprise “layers” that perform operations on vector inputs to generate vector or scalar outputs… instructions...comprises settings and values (such as resistance, impedance, capacitance, inductance, current/voltage ratings, etc.)… each local processing element may utilize one or more collectors (e.g., small register-files): one in front of a weight buffer, another one in front of an accumulation buffer, and another in front of an input activation buffer… Each of the vector multiply-accumulate units 1202 includes a weight collector 1320 buffer having a configurable depth (e.g., number of distinct registers or addresses in a register file used by the vector multiply-accumulate units 1202 during computations) of WD and a width V×N×WP (WP is also called the weight precision)… Some or all of WD, WP, IAP, AD, and AP may be configurable”);
generate an inference model based at least on a training session of the neural network (e.g. in paragraphs 32 and 70, “inference on a trained model [i.e. generated] with less precise data representations to increase performance (improve throughput or latency per inference) and reduce computational energy expended per inference”),
but does not specifically teach the hardware including analog multiply-and-accumulation circuits and during a training session of the neural network: inject noise into an output value generated by a first node of the nodes, the noise modeled at least on the characteristics, the noise emulating noise generated at an output of an analog-to-digital converter of the analog multiply-and- accumulation circuits and the inference model associating a first weight parameter to the node that is based at least on a loss function that integrates the injected noise, the first weight parameter learned through training on a dataset with the loss function injected with an estimation of the intrinsic electrical noise of the analog multiply- and-accumulation circuit.
However, Bunandar teaches during a training session of the neural network: inject noise into an output value generated by a first node of the nodes, the noise modeled at least on characteristics, the noise emulating noise generated at an output of an analog-to-digital converter of an analog circuit and an inference model associating a first weight parameter to a node that is based at least on a loss function that integrates the injected noise, the first weight parameter learned through training on a dataset with the loss function injected with an estimation of the intrinsic electrical noise of the analog circuit (e.g. in paragraphs 50, 53, 70, 106, 123, 143, 160, and 167, “analog-to-digital converter (ADC)…used to convert analog signals output by the analog processor… incorporate noise (e.g., from the analog processor and/or an ADC)… mitigate the effect of noise by injecting noise representative of noise introduced as a result of using an analog processor into outputs of a machine learning model. For example, some embodiments inject noise representative of noise introduced as a result of using an analog processor into outputs of layers of a neural network during training… determine one or more gradients of a loss function with respect to parameters of the machine learning model 112 (e.g., neural network weights) based on the determined outputs… obtain a noise sample from a Gaussian distribution with…a standard deviation commensurate to noise expected from an analog processor. An expected standard deviation may be determined from previously collected outputs of the analog processor… use the gradient to update parameters (e.g., weights and/or biases) of the neural network”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Stevens to include the teachings of Bunandar because one of ordinary skill in the art would have recognized the benefit of optimizing a neural network,
but does not specifically teach wherein the analog circuit includes analog multiply-and-accumulation circuits.
However, Cherupally teaches an analog circuit including analog multiply-and- accumulation circuits (e.g. in paragraphs 3, 6-7, 9-10, 57, and 59, “multiply-and-accumulate (MAC)… (IMC)-based deep neural network (DNN) hardware is provided… performing the analog MAC computation… the injected noise data is directly based on measurements of actual IMC prototype chips… injected noise is directly from IMC chip measurement results on the quantized ADC outputs” and figure 2A). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of the combination to include the teachings of Cherupally because one of ordinary skill in the art would have recognized the benefit of incorporating well-known types of circuits (also amounts a simple substitution that yields predictable results [e.g. see KSR Int'l Co v. Teleflex Inc., 550 US 398,82 USPQ2d 1385,1396 (U.S. 2007) and MPEP § 2143(B)]),
As per claim 2, the rejection of claim 1 is incorporated and the combination further teaches wherein the particular layer comprises at least one of a fully-connected layer; or a convolutional layer (e.g. Stevens, in paragraph 82, “convolutional and fully-connected layers”).
As per claim 3, the rejection of claim 1 is incorporated and the combination further teaches wherein the characteristics comprise at least one of: a bit width for input data provided as an input for each of the analog multiply- and-accumulation circuits; a bit width for a second weight parameter provided as an input for each of the analog multiply-and-accumulation circuits; a bit width for output data output by analog-to-digital converters of the analog multiply-and-accumulation circuits; or a vector size supported by the analog multiply-and-accumulation circuits (e.g. Stevens, in paragraphs 90-92, “weight collector 1320 buffer having a configurable depth (e.g., number of distinct registers or addresses in a register file used by the vector multiply-accumulate units 1202 during computations) of WD and a width V×N×WP... input activations have width IAP. Each of the vector multiply-accumulate units 1202 also includes an accumulation collector 1322 having a configurable operational depth AD and width N×AP… WP×N×V bits wide and is able to supply different weight vectors… values of V and N may be adjusted”).
As per independent claim 15, Stevens teaches a method, comprising:
receiving a configuration file that specifies characteristics of a multiply-and-accumulation circuit utilized to implement a node of a particular layer of a neural network (e.g. in paragraphs 32, 34, 42, 45, 88 and 90-92, ““Neural network” refers to an algorithm or computational system based on a collection of connected units or nodes called artificial neurons… neural networks…acquire differences during training (e.g., be trained to have different weights from one another)… comprise “layers” that perform operations on vector inputs to generate vector or scalar outputs… instructions...comprises settings and values (such as resistance, impedance, capacitance, inductance, current/voltage ratings, etc.)… each local processing element may utilize one or more collectors (e.g., small register-files): one in front of a weight buffer, another one in front of an accumulation buffer, and another in front of an input activation buffer… Each of the vector multiply-accumulate units 1202 includes a weight collector 1320 buffer having a configurable depth (e.g., number of distinct registers or addresses in a register file used by the vector multiply-accumulate units 1202 during computations) of WD and a width V×N×WP (WP is also called the weight precision)… Some or all of WD, WP, IAP, AD, and AP may be configurable”); and
generating an inference model based at least on the training session of the neural network (e.g. in paragraphs 32 and 70, “inference on a trained model [i.e. generated] with less precise data representations to increase performance (improve throughput or latency per inference) and reduce computational energy expended per inference”),
but does not specifically teach an analog multiply-and-accumulation circuit and during a training session of the neural network: injecting noise into an output value generated by the node, the noise modeled at least on the characteristics specified by the configuration file, the injected noise emulating noise generated at an output of an analog-to-digital converter of the analog multiply-and-accumulation circuit; and the inference model associating a first weight parameter to the node that is based at least on a loss function that integrates the injected noise, the first weight parameter learned through training on a dataset with the loss function injected with an estimation of an intrinsic electrical noise of the analog multiply-and-accumulation circuit.
However, Bunandar teaches during a training session of a neural network: injecting noise into an output value generated by a node, the noise modeled at least on characteristics, the injected noise emulating noise generated at an output of an analog-to-digital converter of an analog circuit and an inference model associating a first weight parameter to the node that is based at least on a loss function that integrates the injected noise, the first weight parameter learned through training on a dataset with the loss function injected with an estimation of an intrinsic electrical noise of the analog circuit (e.g. in paragraphs 50, 53, 70, 106, 123, 143, 146, 160, and 167, “analog-to-digital converter (ADC)…used to convert analog signals output by the analog processor… mitigate the effect of noise by injecting noise representative of noise introduced as a result of using an analog processor into outputs of a machine learning model. For example, some embodiments inject noise representative of noise introduced as a result of using an analog processor into outputs of layers of a neural network during training… determine one or more gradients of a loss function with respect to parameters of the machine learning model 112 (e.g., neural network weights) based on the determined outputs… obtain a noise sample from a Gaussian distribution with…a standard deviation commensurate to noise expected from an analog processor. An expected standard deviation may be determined from previously collected outputs of the analog processor… use the gradient to update parameters (e.g., weights and/or biases) of the neural network”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Stevens to include the teachings of Bunandar because one of ordinary skill in the art would have recognized the benefit of optimizing a neural network,
but does not specifically teach wherein the analog circuit includes an analog multiply-and-accumulation circuit.
However, Cherupally teaches an analog circuit including an analog multiply-and- accumulation circuit (e.g. in paragraphs 3, 6-7, 9-10, 57, and 59, “multiply-and-accumulate (MAC)… (IMC)-based deep neural network (DNN) hardware is provided… performing the analog MAC computation… the injected noise data is directly based on measurements of actual IMC prototype chips… injected noise is directly from IMC chip measurement results on the quantized ADC outputs” and figure 2A). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of the combination to include the teachings of Cherupally because one of ordinary skill in the art would have recognized the benefit of incorporating well-known types of circuits (also amounts a simple substitution that yields predictable results [e.g. see KSR Int'l Co v. Teleflex Inc., 550 US 398,82 USPQ2d 1385,1396 (U.S. 2007) and MPEP § 2143(B)]),
As per claim 16, the rejection of claim 15 is incorporated and the combination further teaches wherein the particular layer comprises at least one of: a fully-connected layer; or a convolutional layer (e.g. Stevens, in paragraph 82, “convolutional and fully-connected layers”).
As per claim 17, the rejection of claim 15 is incorporated and the combination further teaches wherein the characteristics comprise at least one of: a bit width for input data provided as an input to the analog multiply-and-accumulation circuit; a bit width for a second weight parameter provided as an input to the analog multiply-and-accumulation circuit; a bit width for output data output by the analog-to-digital converter; an alpha parameter specifying a dominance level of the noise injected into the output value; or a vector size supported by the analog multiply-and-accumulation circuit (e.g. Stevens, in paragraphs 90-92, “weight collector 1320 buffer having a configurable depth (e.g., number of distinct registers or addresses in a register file used by the vector multiply-accumulate units 1202 during computations) of WD and a width V×N×WP... input activations have width IAP. Each of the vector multiply-accumulate units 1202 also includes an accumulation collector 1322 having a configurable operational depth AD and width N×AP… WP×N×V bits wide and is able to supply different weight vectors… values of V and N may be adjusted”).
As per claim 18, the rejection of claim 17 is incorporated and the combination further teaches wherein the noise injected into the output value is randomized in accordance with a distribution function (e.g. Bunandar, in paragraph 53, “injecting noise representative of noise introduced as a result of using an analog processor into outputs of a machine learning model”; Cherupally, in paragraphs 60, 65, and 78, “distributions… fitted Gaussian model… returns a random float in [0, 1]… random samplings of ADC quantization outputs were performed from each probability table for random inputs”).
As per claim 21, the rejection of claim 15 is incorporated and the combination further teaches applying a gradient descent optimization algorithm to the loss function during the training session to determine a second weight parameter (e.g. Bunandar, in paragraphs 17 and 70, “determine a gradient of a loss function; and updating parameters of the machine learning model using the gradient of the loss function… gradient descent training technique… parameters of the machine learning model 112 (e.g., neural network weights)”).
As per claim 22, the rejection of claim 15 is incorporated and the combination further teaches wherein said injecting the noise into the output value generated by the node comprises: during the training session, injecting the loss function with the estimation of the intrinsic electrical noise of the analog multiply-and-accumulation circuit (e.g. Bunandar, in paragraphs 50, 53, 70, 106, 123, 143, 160, and 167, “analog-to-digital converter (ADC)…used to convert analog signals output by the analog processor… mitigate the effect of noise by injecting noise representative of noise introduced as a result of using an analog processor into outputs of a machine learning model. For example, some embodiments inject noise representative of noise introduced as a result of using an analog processor into outputs of layers of a neural network during training… determine one or more gradients of a loss function with respect to parameters of the machine learning model 112 (e.g., neural network weights) based on the determined outputs… obtain a noise sample from a Gaussian distribution with…a standard deviation commensurate to noise expected from an analog processor. An expected standard deviation may be determined from previously collected outputs of the analog processor”; Cherupally, in paragraphs 3, 6-7, 9-10, 57, and 59, “multiply-and-accumulate (MAC)… (IMC)-based deep neural network (DNN) hardware is provided… performing the analog MAC computation… the injected noise data is directly based on measurements of actual IMC prototype chips… injected noise is directly from IMC chip measurement results on the quantized ADC outputs”)
Claims 8-10 are rejected under 35 U.S.C. 103 as being unpatentable over Stevens et al. (US 20220067513 A1) in view of Bunandar et al. (US 20220172052 A1) and Cherupally et al. (US 20220318628 A1), and further in view of Chai et al. (US 20200134461 A1), Lee et al. (US 20230077987 A1), and Park et al. (US 20230100036 A1).
Claims 8-10 are the method claims corresponding to system claims 1-3 and are rejected under the same reasons set forth and the combination further teaches generating an inference model based at least on the training session of the neural network with features causing output values generated by the multiply-and-accumulation circuits to have reduced precision (e.g. Stevens, in paragraphs 32 and 70, “inference on a trained model [i.e. generated] with less precise data representations to increase performance (improve throughput or latency per inference) and reduce computational energy expended per inference”),
but does not specifically teach determining an estimate of an amount of power consumed by the analog multiply-and-accumulation circuits during execution thereof based at least on a number of non-zero midterms generated by the nodes; and modifying a loss function of the neural network based at least on the estimate, the modified loss function comprising an equation that includes a value indicative of the estimate of the power consumed by the analog multiply-and-accumulation circuits and the modified loss function causing weight parameters of the inference model to have a sparse bit representation.
However, Chai teaches during a training session of a neural network: determining an estimate of an amount of power consumed by hardware during execution thereof and modifying a loss function of the neural network based at least on power, the modified loss function comprising an equation that includes a value (e.g. in paragraphs 37, 68, 72, 136, 145, and 147, “consume an estimated…Tflops/s… estimate of the parameters… map the neural network software architecture to appropriate processors in the system architecture. Machine learning system 104 may use a cost function to select the best mapping (e.g., a best fit algorithm can be use). The cost function can be one of size, weight, power and cost… the cost function may be the loss function as described previously as Equation (34), where λ1 and λ2 are set based to optimize…power…of the hardware architecture… In reference to the loss function selection the lambda parameters (λ.sub.1, λ.sub.2 and λ.sub.3) in equation (34) based in hardware parameters P, this disclosure describes hardware parameters P that may be updated during training to affect the selection of the annealing constraints… selection of hardware parameters P during training allows for a pareto-optimal selection of hardware resources with respect to…power”), wherein features include a modified loss function causing weight parameters of the inference model to have a sparse bit representation (e.g. in paragraphs 42 and 72, “accommodate low precision weights… sparsity may refer to the use of low-precision weights, which may only take a limited (sparse) number of values… where low-precision weights 116 are constrained…machine learning system 104 may use the loss function”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of the combination to include the teachings of Chai because one of ordinary skill in the art would have recognized the benefit of optimizing a neural network,
but does not specifically teach power based at least on a number of non-zero midterms generated by the nodes; and modifying a loss function based at least on the estimate, the modified loss function comprising an equation that includes a value indicative of the estimate of the power consumed by the analog multiply-and-accumulation circuits.
However, Lee teaches modify a loss function based at least on an estimate of power consumed by hardware, the modified loss function comprising an equation that includes a value indicative of the estimate of the power consumed by the hardware (e.g. in paragraphs 75, 81, 94-95, 98-99, and 124, “the cost estimation network may predict hardware metrics using the cost function, and the cost function may be defined as a linear combination of the latency, the area, and the energy consumption, or may be defined as the combination and the product between the latency, the area, and the energy consumption… execution of multiple MAC (Multiply-Accumulate) operations, which are the most common operations in recent CNNs… cost estimation network may generate as output the three cost metrics of interest (i.e., latency, area, and energy consumption) based on the ground truth generated by the evaluation software… controlling λ.sub.E, λ.sub.L, and λ.sub.A, conditions for how to measure the balance between each cost metric may be set… dynamic energy consumption may mainly depend on the number of MAC operations”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of the combination to include the teachings of Lee because one of ordinary skill in the art would have recognized the benefit of optimizing a neural network,
but does not specifically teach power based at least on a number of non-zero midterms generated by the nodes.
However, Park teaches power based at least on a number of non-zero midterms generated by nodes (e.g. in paragraphs 31, 54-57, and 80, “a feature tensor within a model (e.g., activation data output by one layer of a neural network and used as input to a subsequent layer… determining the predicted load can include determining the sparsity or density of each sub-tensor 315 (e.g., determining the number of elements with a value of zero, or determining the number of non-zero elements). Generally, if an element has a value of zero, then processing it using the MAC array will draw little or no power. In contrast, elements with non-zero values will require power draw…during the processing”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of the combination to include the teachings of Park because one of ordinary skill in the art would have recognized the benefit of determining power based on relevant factors.
Claims 4-6 and 11-13 are rejected under 35 U.S.C. 103 as being unpatentable over Stevens et al. (US 20220067513 A1) in view of Bunandar et al. (US 20220172052 A1) and Cherupally et al. (US 20220318628 A1) and further in view of Chai et al. (US 20200134461 A1), Lee et al. (US 20230077987 A1), Park et al. (US 20230100036 A1) and Mallinson (US 8766841 B2).
As per claim 4, the rejection of claim 1 is incorporated, but the combination does not specifically teach determine an estimate of the amount of power consumed by the analog multiply-and-accumulation circuits during execution thereof by: for each node of the nodes: determining a number of non-zero midterms generated by the node; determining a computational precision value of the node; combining the number of non-zero midterms generated by the node and the computational precision value of the node to generate a node estimate of an amount of power consumed by an analog multiply-and accumulation circuit of the analog multiply-and accumulation circuits corresponding to the node; and combining the node estimates to generate the estimate of the amount of power consumed by the analog multiply-and-accumulation circuits; and modify the loss function based at least on the estimate of the amount of power consumed by the analog multiply-and-accumulation circuits.
However, Chai teaches determining an estimate of an amount of power consumed by hardware during execution thereof to generate the estimate of the amount of power consumed by the hardware and modifying a loss function based at least on power (e.g. in paragraphs 37, 68, 72, 136, 145, and 147, “consume an estimated…Tflops/s… estimate of the parameters… map the neural network software architecture to appropriate processors in the system architecture. Machine learning system 104 may use a cost function to select the best mapping (e.g., a best fit algorithm can be use). The cost function can be one of size, weight, power and cost… the cost function may be the loss function as described previously as Equation (34), where λ1 and λ2 are set based to optimize…power…of the hardware architecture… In reference to the loss function selection the lambda parameters (λ.sub.1, λ.sub.2 and λ.sub.3) in equation (34) based in hardware parameters P, this disclosure describes hardware parameters P that may be updated during training to affect the selection of the annealing constraints… selection of hardware parameters P during training allows for a pareto-optimal selection of hardware resources with respect to…power”), wherein features include a modified loss function causing weight parameters of the inference model to have a sparse bit representation (e.g. in paragraphs 42 and 72, “accommodate low precision weights… sparsity may refer to the use of low-precision weights, which may only take a limited (sparse) number of values… where low-precision weights 116 are constrained…machine learning system 104 may use the loss function”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of the combination to include the teachings of Chai because one of ordinary skill in the art would have recognized the benefit of optimizing a neural network,
but does not specifically teach power thereof by: for each node of the nodes: determining a number of non-zero midterms generated by the node; determining a computational precision value of the node; combining the number of non-zero midterms generated by the node and the computational precision value of the node to generate a node estimate of an amount of power consumed by an analog multiply-and accumulation circuit of the analog multiply-and accumulation circuits corresponding to the node; and combining the node estimates to generate the estimate; and modify the loss function based at least on the estimate of the amount of power consumed by power based at least on a number of non-zero midterms generated by the nodes.
However, Lee teaches modify a loss function based at least on an estimate of power consumed by hardware, the modified loss function comprising an equation that includes a value indicative of the estimate of the power consumed by the hardware (e.g. in paragraphs 75, 81, 94-95, 98-99, and 124, “the cost estimation network may predict hardware metrics using the cost function, and the cost function may be defined as a linear combination of the latency, the area, and the energy consumption, or may be defined as the combination and the product between the latency, the area, and the energy consumption… execution of multiple MAC (Multiply-Accumulate) operations, which are the most common operations in recent CNNs… cost estimation network may generate as output the three cost metrics of interest (i.e., latency, area, and energy consumption) based on the ground truth generated by the evaluation software… controlling λ.sub.E, λ.sub.L, and λ.sub.A, conditions for how to measure the balance between each cost metric may be set… dynamic energy consumption may mainly depend on the number of MAC operations”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of the combination to include the teachings of Lee because one of ordinary skill in the art would have recognized the benefit of optimizing a neural network,
but does not specifically teach power thereof by: for each node of the nodes: determining a number of non-zero midterms generated by the node; determining a computational precision value of the node; combining the number of non-zero midterms generated by the node and the computational precision value of the node to generate a node estimate of an amount of power consumed by an analog multiply-and accumulation circuit of the analog multiply-and accumulation circuits corresponding to the node; and combining the node estimates to generate the estimate and power based at least on a number of non-zero midterms generated by the nodes.
However, Park teaches power thereof by for each node of nodes: determining a number of non-zero midterms generated by the node and power being based at least on a number of non-zero midterms generated by nodes associated with generating a node estimate of an amount of power consumed by a multiply-and accumulation circuit of multiply-and accumulation circuits corresponding to the node (e.g. in paragraphs 31, 54-57, and 80, “a feature tensor within a model (e.g., activation data output by one layer of a neural network and used as input to a subsequent layer… determining the predicted load can include determining the sparsity or density of each sub-tensor 315 (e.g., determining the number of elements with a value of zero, or determining the number of non-zero elements). Generally, if an element has a value of zero, then processing it using the MAC array will draw little or no power. In contrast, elements with non-zero values will require power…during the processing”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of the combination to include the teachings of Park because one of ordinary skill in the art would have recognized the benefit of determining power based on relevant factors.
but does not specifically teach determining a computational precision value of the node; combining the number of non-zero midterms generated by the node and the computational precision value of the node to generate a node estimate; and combining the node estimates to generate the estimate.
However, Mallinson teaches determining a number of non-zero midterms generated by a node (e.g. in column 1 lines 40-67 and column 4 lines 35-47, “0.8, 0.5, 0.7, 0.6, 0.9, 0.7, 0.8, 0.6, 0.7, 0.8 [i.e. non-zero midterms]… a resistor of value 2R is similarly connected (typically in parallel), as well as resistors of values 4R, 8R and so forth”), determining a computational precision value of the node (e.g. in column 5 lines 36-48, “code control bits of a given significance… sets of three uppermost bits in each branch of FIG. 3 code for the 3 most significant control bits… control signal”), combining the number of non-zero midterms generated by the node and the computational precision value of the node to generate a node feature (e.g. in column 1 lines 40-57, column 3 lines 47--65, and column 5 lines 36-48, “used only one time to merge each of the segments after they are summed--into the final output… voltage present at the node”), and combining the node features to generate a sum of features (e.g. in column 1 lines 40-57, “multiple MDAC's to form a sum-of-products… voltage on that output node”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of the combination to include the teachings of Mallinson because one of ordinary skill in the art would have recognized the benefit of facilitating relevant calculations.
As per claim 5, the rejection of claim 4 is incorporated and the combination further teaches wherein the computational precision value is based at least on a most significant bit of an output value generated by the node (e.g. Mallinson, in column 5 lines 36-48, “code control bits of a given significance… sets of three uppermost bits in each branch of FIG. 3 code for the 3 most significant control bits… control signal”).
As per claim 6, the rejection of claim 4 is incorporated and the combination further teaches wherein the neural network model trainer is further configured to: apply a gradient descent optimization algorithm to the modified loss function during the training session to determine the weight parameters (e.g. Bunandar, in paragraphs 17 and 70, “determine a gradient of a loss function; and updating parameters of the machine learning model using the gradient of the loss function… gradient descent training technique… parameters of the machine learning model 112 (e.g., neural network weights)”).
Claims 11-13 are the method claims corresponding to system claims 4-6 and are rejected under the same reasons set forth.
Claim 19 is rejected under 35 U.S.C. 103 as being unpatentable over Stevens et al. (US 20220067513 A1) in view of Bunandar et al. (US 20220172052 A1) and Cherupally et al. (US 20220318628 A1) and further in view of Aggarwal et al. (US 20090060095 A1).
As per claim 19, the rejection of claim 18 is incorporated and the combination further teaches wherein the distribution function is a normal distribution (e.g. Cherupally, in paragraphs 60, 65, and 78, “distributions… fitted Gaussian [i.e. normal] model… returns a random float in [0, 1]… random samplings of ADC quantization outputs were performed from each probability table for random inputs… injecting noise at the weight-level drawn from Gaussian distributions”), but does not specifically teach having a zero mean and a predetermined variance. However, Aggarwal teaches a distribution function having a zero mean and a predetermined variance (e.g. in paragraph 23, “example of such a distribution would be a gaussian kernel with kernel width h. This corresponds to the gaussian distribution with zero mean and variance h.sup.2”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of the combination to include the teachings of Aggarwal because one of ordinary skill in the art would have recognized the benefit of incorporating well-known types of distributions (also amounts a simple substitution that yields predictable results [e.g. see KSR Int'l Co v. Teleflex Inc., 550 US 398,82 USPQ2d 1385,1396 (U.S. 2007) and MPEP § 2143(B)]).
Claim 20 is rejected under 35 U.S.C. 103 as being unpatentable over Stevens et al. (US 20220067513 A1) in view of Bunandar et al. (US 20220172052 A1), Cherupally et al. (US 20220318628 A1), and Aggarwal et al. (US 20090060095 A1) and further in view of Chen et al. (US 20200097823 A1) and Sutherland et al. (US 5214745).
As per claim 20, the rejection of claim 19 is incorporated, but the combination does not specifically teach wherein the predetermined variance is based at least on the bit width for the output data that is outputted by the analog-to-digital converter and the alpha parameter.
However, the combination teaches output data that is outputted by the analog-to-digital converter (e.g. Bunandar, in paragraph 53, “incorporate noise (e.g., from the analog processor and/or an ADC)… mitigate the effect of noise by injecting noise representative of noise introduced as a result of using an analog processor into outputs of a machine learning model”; Cherupally, in paragraphs 6 and 52, “IMC performs MAC computation inside the on-chip memory (e.g., SRAM) by activating multiple/all rows of the memory array. The MAC result is represented by analog bitline voltage/current and subsequently digitized by an analog-to-digital converter (ADC) in the peripheral of the array… quantize the analog voltage/current into digital values”) and Chen teaches variance being based on a bit width for data (e.g. in abstract and paragraph 33, “variance values for the different master bit-widths that are evaluated for a current layer/channel”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of the combination to include the teachings of Chen because one of ordinary skill in the art would have recognized the benefit of incorporating relevant features associated with variance (also amounts a simple substitution that yields predictable results [e.g. see KSR Int'l Co v. Teleflex Inc., 550 US 398,82 USPQ2d 1385,1396 (U.S. 2007) and MPEP § 2143(B)]),
but does not specifically teach wherein the predetermined variance is based at least on the alpha parameter.
However, the combination also teaches an output associated with noise injected into the output value (e.g. Bunandar, in paragraph 53, “some embodiments inject noise representative of noise introduced as a result of using an analog processor into outputs of layers of a neural network during training”; Cherupally, in paragraph 10, “embodiments perform noise injection at the partial sum level”) and Sutherland teaches variance being based at least on an alpha parameter specifying a dominance level associated with an output (e.g. in column 14 lines 60-66, “displays a varying dominance or magnitude (.lambda..sub.p) which statistically is inversely proportional to the pattern variance (eq. 23). In other words, the encoded patterns most similar to the input stimulus pattern [S]* produce the more dominant contribution within the generated response output”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of the combination to include the teachings of Sutherland because one of ordinary skill in the art would have recognized the benefit of incorporating relevant features associated with variance (also amounts a simple substitution that yields predictable results [e.g. see KSR Int'l Co v. Teleflex Inc., 550 US 398,82 USPQ2d 1385,1396 (U.S. 2007) and MPEP § 2143(B)]).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
For example,
Deng et al. (US 20190180184 A1) teaches power based at least on a number of non-zero midterms generated by nodes (e.g. in paragraphs 25 and 36, “Optimally reducing the number of non-zero weights, in turn, optimally reduces the computational burden… consume less power because the DNN includes fewer multiply and accumulate (MAC) operations” and figure 1).
He et al. (US 20230196103 A1) teaches “the weight precision of each layer in the neural network is tried to be reduced… the resource utilization rate and the performance are improved, and the power consumption is reduced” (e.g. in paragraph 205).
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to WILLIAM WONG whose telephone number is (571)270-1399. The examiner can normally be reached Monday-Friday 9am-5pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, TAMARA KYLE can be reached at (571)272-4241. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/W.W/Examiner, Art Unit 2144 07/29/2026
/TAMARA T KYLE/Supervisory Patent Examiner, Art Unit 2144