Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 03/23/2025 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1, 3-6, 9, 14-15, and 17-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Cohen, J. P., Lo, H. Z., & Ding, W. (2016). RandomOut: Using a convolutional gradient norm to rescue convolutional filters. arXiv.Org. https://arxiv.org/abs/1602.05931\, hereinafter “Cohen” and further in light of U.S. Patent Application Publication 20200210807 “Lorraine”.
Claim 1:
Cohen teaches a method of training a neural network having network parameters and configured to process a network input to generate a network output for the network input, the neural network comprising: a plurality of neural network layers, the plurality of neural network layers comprising a first neural network layer having a plurality of neurons (i.e. pg. 2, Section 1. Introduction, “We propose the method called RANDOMOUT in §2 that scores filters and replaces them at training time if they have been abandoned by the network”, wherein it is noted that a convolution neural network has a layer with a plurality of neurons), and the network parameters comprising: for each of the neurons, a respective set of incoming weights associated with the neuron and a respective set of outgoing weights associated with the neuron (i.e. pg. 3, Section 2. RandomOut, Figure 3, “edges represent the output of inputs and each intermediate computation. Gradients ∂f / ∂w are shown inside the boxes of the inputs and weights”, wherein it is noted in Figure 3 that each neuron node has an input and output weight), the method comprising: at each of a plurality of training steps (i.e. pg. 2, Section 1. Introduction, “We propose the method called RANDOMOUT in §2 that scores filters and replaces them at training time if they have been abandoned by the network”, wherein it is noted the determination if a filter has been abandoned or becomes dormant occurs during the training time for the neural network): obtaining a set of training data for the training step (i.e. pg. 3, Section 3.Experimental Setup, “CraterCNN has two convolutional layers, followed by a fully connected layer, then softmax. The input is 15x15”, wherein it is noted that the set of training data encompasses the input data and its associated learning rate for the CNN) ; training the neural network on the training data to update the network parameters (i.e. pg. 4, Section 4. Experiments, “The resulting test error of the 28x28 Inception-V3 network in three experimental conditions are shown. The base network fails to converge to a satisfactory local minimum 26% of the time”, wherein it is noted that before utilizing the RandomOut algo to reset the abandoned layers, a CNN would have unsatisfactory training) determining whether a resetting criterion is satisfied at the training step (i.e. pg. 3, Section 2, RandomOut, “Formulating this into an algorithm, RANDOMOUT has two hyperparameters, a threshold τ and a “% of epochs active” P. During training each filter k is checked at regular intervals to see if CGN(k) < τ and, if so, filter k is re-initialized”, wherein the BRI for a resetting criterion encompasses the threshold τ that measures if a layer has been abandoned); in response to determining that the criterion is satisfied (i.e. pg. 1, Abstract, We use the gradient norm to evaluate the impact of a filter on error, and re-initialize filters when the gradient norm of its weights falls below a specific threshold): determining, for each of the plurality of neurons, an expected absolute value of an activation generated by the neuron during processing of a given network input (i.e. pg. 2, Section 2, “The RANDOMOUT approach is to reinitialize weights for abandoned parts of the network if their CGN is below a threshold τ near 0”, wherein it is noted that a filter is made up of individual neurons and that the BRI for an expected value of an activation encompasses that the convolutional filter weight should have a n expected convolutional gradient norm (CGN) above a threshold) ; determining, for each neuron and based on the expected absolute value for the neuron, whether to classify the neuron as a dormant neuron (i.e. pg. 2, Section 1, Introduction, “We call these filters “abandoned” by the network because they contribute little to minimizing the error”, wherein the BRI for a dormant neuron encompasses classifying the sum of each neuron in a layer as abandoned due to the sum of the CGN representing each neuron in a layer ais being below a threshold or close to zero.); and for any neuron that is classified as a dormant neuron, modifying the incoming weights associated with the dormant neuron and the outgoing weights associated with the dormant neuron (i.e. pg. 3, Section 2, “If the filters are randomized to a value that is used later in the network to reduce error its gradients will gradually increase and the section will slowly be introduced back into the network”, wherein the abandoned neuron layer would have its incoming modified to random values that are used later in the network which in turn modifies the outgoing weights as the abandoned neuron layer is slowed introduced back into the network).
While Cohen teaches a method of classifying and modifying neurons that are not contributing to a model because their output is below an expected threshold value. Cohen may not explicitly teach that the threshold is based on
Determining, an expected absolute value of an activation generated by the neuron
determining, for each neuron and based on the expected absolute value for the neuron, whether to classify the neuron as a dormant neuron.
However Lorraine teaches,
determining, for each of the plurality of neurons, an expected absolute value (i.e. para. [0099], computing a weighted sum h is achieved by accumulating the coefficient of the convolution kernel at each arrival of a spike on the corresponding input. The activation function of the neuron g may in this case be replaced by a threshold. When the absolute value of the sum h exceeds the threshold following the arrival of a spike on the input sub-matrix, the output neuron emits a spike of the sign of h) of an activation generated by the neuron (i.e. para. [0115], “Each neuron O.sub.j has its own synaptic weights W.sub.i,j with the corresponding inputs I.sub.i and performs the weighted sum h( ) f the input coefficients with the weights, which is then passed to the neuron activation function g( ) in order to obtain the output of the neuron t”, wherein the BRI for an expected absolute value of an activation generated by the neuron encompasses how the weighted sum h() is representative of an activation of incoming signals to the neuron)
determining, for each neuron and based on the expected absolute value for the neuron, whether to classify the neuron as a dormant neuron (i.e. para. [0115] “Each convolutional computation module 20 may be configured so as to compute the internal value (also called “integration value”) of the neurons of a convolution layer that have received a spike (input event). When this integration value exceeds a predefined threshold, the neurons are “triggered” or “activated” and emit a spike event at output”, wherein the BRI to classify the neuron as a dormant neuron encompasses how if absolute value of the sum h is below the threshold, then the neuron may be considered dormant and not eligible to for a reset to zero)
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to add determining, an expected absolute value of an activation generated by the neuron; determining, for each neuron and based on the expected absolute value for the neuron, whether to classify the neuron as a dormant neuron, to Cohen’s identification and reactivation of dormant neurons, with how the criteria determining to do something to a neuron is based on some sort of expected absolute value for the neuron, as taught by Lorraine. One would have been motivated to combine the dormant neuron resetting techniques of Cohen and a trigger that depends on an expected absolute value for the neuron of Lorraine, and would have had a reasonable expectation of success in order to reduce memory and reduce computational complexity and therefore the hardware resources necessary.
Claim 3:
Cohen and Lorraine teach the method of claim 1.
Cohen further teaches wherein modifying the incoming weights comprises: setting the weights to values that have been initialized using a parameter initialization technique (i.e. pg. 3, Section 2. RandomOut, “The RANDOMOUT approach is to reinitialize weights for abandoned parts of the network if their CGN is below a threshold τ near 0… If the filters are randomized to a value that is used later in the network to reduce error its gradients will gradually increase and the section will slowly be introduced back into the network”, wherein the BRI for a parameter initiation technique encompasses initazling the parameters to randomized values that would be used later in the network).
Claim 4:
Cohen and Lorraine teach the method of claim 3.
Cohen further teaches wherein initializing the weights using the parameter initialization technique comprises sampling values for the weights from a specified initialization distribution (i.e. pg. 3-4, Section 3. Experimental Setup, “The initial weights throughout the network are initialized using the Xavier initialization (Glorot & Bengio, 2010) scheme”, wherein the weights are sampled from the specified Xavier initialization distribution).
Claim 5:
Cohen and Lorraine teach the method of claim 1.
Lorraine further teaches
wherein modifying the incoming weights comprises: scaling each incoming weight to the dormant neuron using a mean of incoming weights for neurons that have not been classified as dormant neurons (i.e. para. [0070], Fig. 3, “A convolution layer comprises one or more output matrices comprising a set of output neurons, each output matrix being connected to an input matrix (the input matrix comprising a set of input neurons) by artificial synapses associated with a convolution matrix comprising the synaptic weight coefficients corresponding to output neurons of the output matrix (synaptic weights) (the synaptic weight coefficients are also called “synaptic weights” or “weight coefficient” or “convolutional coefficients” or “weightings”). The output value of each output neuron is determined from those input neurons of the input matrix to which the output neuron is connected and the synaptic weight coefficients of the convolution matrix associated with the output matrix”, wherein it is noted that the BRI for scaling each incoming weight to the dormant neurons encompasses any sort of multiplicative action to a weight, which encompasses how the matrix multiplication of the input and the synaptic weight matrices is effectively finding a mean of weights for neurons that have not been classified as dormant neurons and any neuron that has not had its weighted sum h reset to zero would be classified as not dormant. It is noted that once the fixed refractory period for a neuron elapses, the neuron would be free to eventually trigger and scale its output again)
Claim 6:
Cohen and Lorraine teach the method of claim 1.
Lorraine further teaches wherein modifying the outgoing weights comprises: setting the outgoing weights to zero (i.e. para. [0099], “When the absolute value of the sum h exceeds the threshold following the arrival of a spike on the input sub-matrix, the output neuron emits a spike of the sign of h and resets the weighted sum h to the value 0. The neuron then enters into what is called a “refractory” period during which it is no longer able to emit spikes for a fixed period”, wherein it is noted that setting the weighted sum h to zero effectively sets the outgoing weights to zero as the neuron is unable to produce an output for a period).
Claim 9:
Cohen and Lorraine teach the method of claim 1.
Cohen further teaches wherein determining whether to classify the neuron as a dormant neuron comprises: determining that the neuron is a dormant neuron when the expected absolute value is equal to zero (i.e. pg. 3, Section 2. RandomOut, “The motivation for the threshold is that the CGN is hardly ever 0, because learning rates are fractional so update rules only approach 0, but will become very close when the network has stopped learning a filter”, wherein it is noted that the expected value is a number at least approaching 0 and therefore equivalent to 0).
Lorraine further teaches looking at an expected absolute value (i.e. para. [0099], computing a weighted sum h is achieved by accumulating the coefficient of the convolution kernel at each arrival of a spike on the corresponding input. The activation function of the neuron g may in this case be replaced by a threshold. When the absolute value of the sum h exceeds the threshold following the arrival of a spike on the input sub-matrix, the output neuron emits a spike of the sign of h).
Claim 14:
Cohen and Lorraine method of claim 1.
Cohen further teaches wherein determining whether a resetting criterion is satisfied at the training step comprises: determining that the resetting criterion is satisfied when N training steps have elapsed after a preceding training step at which the resetting criterion was satisfied, wherein N is an integer greater than one (i.e. pg. 3, Section 2. RandomOut, “, RANDOMOUT has two hyperparameters, a threshold τ and a “% of epochs active” P. During training each filter k is checked at regu lar intervals to see if CGN(k) < τ and, if so, filter k is re-initialized… For our networks, τ = 10−8 yielded good results”, wherein the resetting criterion to reinitialize weight for abandoned filters is satisfied after any of epochs N. Wherein if at least one training epoch triggers a resetting criterion than the resetting criterion was satisfied at any number N that is greater than one).
Claim 15:
Claim 15 is the system claim reciting similar limitations to claim 1 and is rejected for similar reasons.
Claim 17:
Claim 17 is the system claim reciting similar limitations to claim 3 and is rejected for similar reasons.
Claim 18:
Claim 18 is the system claim reciting similar limitations to claim 4 and is rejected for similar reasons.
Claim 19:
Claim 19 is the system claim reciting similar limitations to claim 6 and is rejected for similar reasons.
Claim 20:
Claim 20 is the medium claim reciting similar limitations to claim 1 and is rejected for similar reasons.
Claim(s) 2, 11, and 16 is/are rejected under 35 U.S.C. 103 as being unpatentable over Cohen, J. P., Lo, H. Z., & Ding, W. (2016). RandomOut: Using a convolutional gradient norm to rescue convolutional filters. arXiv.Org. https://arxiv.org/abs/1602.05931\, hereinafter “Cohen” and further in light of U.S. Patent Application Publication 20200210807 “Lorraine”, as applied to Claim 1 above, and further in light of U.S. Patent Application Publication 20190258918 “Wang”.
Claim 2:
Cohen and Lorraine teach the method of claim 1.
Cohen and Lorraine may not explicitly teach
wherein the neural network is used to control an agent interacting with an environment and wherein training the neural network comprises training the neural network through reinforcement learning on trajectories sampled from a replay memory.
However, Wang teaches
wherein the neural network is used to control an agent interacting with an environment (i.e. para. [0006], This specification describes a reinforcement learning system implemented as computer programs on one or more computers in one or more locations that selects actions to be performed by an agent interacting with an environment) and wherein training the neural network comprises training the neural network through reinforcement learning on trajectories sampled from a replay memory (i.e. para. [0006] the system trains the action selection policy neural network on trajectories generated as a result of interactions of the agent with the environment. In particular, the system can sample trajectories from a replay memory and then adjust the current values of the parameters of the neuron network by training the neural network on the sampled transitions).
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to add wherein the neural network is used to control an agent interacting with an environment and wherein training the neural network comprises training the neural network through reinforcement learning on trajectories sampled from a replay memory, to Cohen-Lorraine’s identification and reactivation of dormant neurons, with wherein the neural network is used to control an agent interacting with an environment and wherein training the neural network comprises training the neural network through reinforcement learning on trajectories sampled from a replay memory, as taught by Wang. One would have been motivated to combine Wang and Cohen-Lorraine, and would have had a reasonable expectation of success in order in order to have faster and more accurate learning (Wang, para. [0010]).
Claim 11:
Cohen, Lorraine, and Kim teach the method claim 10.
Cohen, Lorraine, and Kim may not explicitly teach
wherein the neural network is used to control an agent interacting with an environment and wherein training the neural network comprises training the neural network through reinforcement learning on trajectories sampled from a replay memory, the method further comprising: sampling the batch from the replay memory.
However, Wang teaches
wherein the neural network is used to control an agent interacting with an environment (i.e. para. [0006], This specification describes a reinforcement learning system implemented as computer programs on one or more computers in one or more locations that selects actions to be performed by an agent interacting with an environment) and wherein training the neural network comprises training the neural network through reinforcement learning on trajectories sampled from a replay memory (i.e. para. [0006] the system trains the action selection policy neural network on trajectories generated as a result of interactions of the agent with the environment. In particular, the system can sample trajectories from a replay memory and then adjust the current values of the parameters of the neuron network by training the neural network on the sampled transitions), the method further comprising: sampling the batch from the replay memory (i.e. para. [0020], the method comprising may comprise receiving a batch of training data comprising a plurality of training examples. Then, for each of the plurality of training examples in the batch, the method may include: processing, using the main neural network in accordance with current values of the main network parameters, the training example to generate a main network output for the training example)
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to add sampling the batch from the replay memory, to Cohen-Lorraine’s identification and reactivation of dormant neurons, with sampling the batch from the replay memory, as taught by Wang. One would have been motivated to combine Wang and Cohen-Lorraine, and would have had a reasonable expectation of success in order in order to have faster and more accurate learning (Wang, para. [0010]).
Claim 16:
Claim 16 is the system claim reciting similar limitations to claim 2 and is rejected for similar reasons.
Claim(s) 10 and 12-13 is/are rejected under 35 U.S.C. 103 as being unpatentable over Cohen, J. P., Lo, H. Z., & Ding, W. (2016). RandomOut: Using a convolutional gradient norm to rescue convolutional filters. arXiv.Org. https://arxiv.org/abs/1602.05931\, hereinafter “Cohen” and further in light of U.S. Patent Application Publication 20200210807 “Lorraine”, as applied to Claim 1 above, and further in light of U.S. Patent Application Publication 20210383203 “Kim”.
Claim 10:
Cohen and Lorraine teach the method of claim 1.
Lorrain teaches wherein determining, for each of the plurality of neurons, an expected absolute value of an activation generated by the neuron during processing of a given network input neuron (i.e. para. [0115], “Each neuron O.sub.j has its own synaptic weights W.sub.i,j with the corresponding inputs I.sub.i and performs the weighted sum h( ) f the input coefficients with the weights, which is then passed to the neuron activation function g( ) in order to obtain the output of the neuron”, wherein the BRI for an expected absolute value of an activation generated by the neuron encompasses how the weighted sum h() is representative of an activation of incoming signals to the neuron).
Cohen and Lorrain may not explicitly teach the determining the expected absolute value comprises, for each neuron: determining an average of absolute values of activations generated by the neuron during processing of each network input in a batch of network inputs.
However, Kim teaches for each neuron: determining an average of absolute values of activations generated by the neuron during processing of each network input in a batch of network inputs (i.e. para. [0100], “when an average value of the absolute values of the 1024 initial weight value 610, which are 32-bit floating point numbers, is calculated for each neuron, the binary weight values 620 may be multiplied by a result of the calculation”, wherein an average value of the absolute value of the initial activation values are calculated for each neuron. Wherein the BRI for a batch of network inputs encompasses para. [0118], “input activations (e.g., initial input values or output values from a previous layer) by initial weight values (e.g., 32-bit floating point numbers)”).
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to add determining an average of absolute values of activations generated by the neuron during processing of each network input in a batch of network inputs, to Cohen-Lorraine’s identification and reactivation of dormant neurons, with determining an average of absolute values of activations generated by the neuron during processing of each network input in a batch of network inputs, as taught by Kim. One would have been motivated to combine normalization calculations of Kim and Cohen-Lorraine, and would have had a reasonable expectation of success in order in so that an operation count may be advantageously reduced, thereby reducing a memory used.
Claim 12:
Cohen, Lorraine, and Kim teach the method of claim 10.
Kim further teaches
wherein the batch of network inputs includes one more network inputs that are not in the set of training data for the training iteration (i.e. para. [0089], “During training and implementation such connections and connection weights may be selectively implemented, removed, and varied to generate or obtain a resultant neural network that is thereby trained and that may be correspondingly implemented for the trained objective, such as for any of the above example recognition objectives”, wherein it is noted that further inputs may be included past the initial input values ass training progresses).
Claim 13:
Cohen, Lorraine, and Kim teach the method of claim 10.
Kim further teaches wherein the batch of network inputs includes one more network inputs that are in the set of training data for the training iteration (i.e. para. [0089], “During training and implementation such connections and connection weights may be selectively implemented, removed, and varied to generate or obtain a resultant neural network that is thereby trained and that may be correspondingly implemented for the trained objective, such as for any of the above example recognition objectives”, wherein it is noted that the initial input values would be included as part of the initial training).
Allowable Subject Matter
Claim 7 is objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Claim 8 is objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
U.S. Patent Application Publication NO. 20210027166 “Gorokhov” teaches in para. [0026], to randomly removing neurons or groups of neurons from the network. The static techniques may also involve considering an absolute magnitude of weights and activations (e.g., importance of neurons) and removing the least of them in each network layer. In yet another example, the static techniques may consider an error of the network during the training time and attempt to learn parameters that represent the probability that a particular neuron or group of neurons may be dropped”
Any inquiry concerning this communication or earlier communications from the examiner should be directed to DAVID H TAN whose telephone number is (571)272-7433. The examiner can normally be reached M-F 7:30-4:30.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Cesar Paula can be reached at (571) 272-4128. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/D.T./Examiner, Art Unit 2145
/CESAR B PAULA/Supervisory Patent Examiner, Art Unit 2145