DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Continued Examination Under 37 CFR 1.114
A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 3/12/2026 has been entered.
DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Amendment
(Submitted on 3/12/2026)
The examiner submits that the applicant has made amendments to the independent claims 13, 21, and 23-24.
Regarding 101 rejections
- On Pages 8-12, the applicant argues that in the previous amendment "training the
artificial neural network in a second phase with the first gradient-based learning algorithm as a function of the optimized parameter set as a function of one second training task" itself reflects a technological improvement in the field of training of ANNs that integrates into a practical application the judicial exception purportedly recited in other claim limitations because this limitation improves neural network training by significantly reducing the training effort” and
cites [0002] and [0005] of the specification and testing whether the claims reflect this improvement by reciting a claim limitation that provides this improvement. The applicant further argues that the claims do indeed recite sensor-related limitations, and Prong Two does in fact require that any additional element asserted to integrate the judicial exception into a practical application be analyzed in view of the claims “ as a whole" under Prong 2 to establish the evidence that the limitations containing the judicial exception as well as the additional elements in the claim besides the judicial exception need to be evaluated together to determine whether the claim integrates the judicial exception into a practical application and argues that at a minimum, the examiner’s any rebuttal would have to engage with the actual applicant argument that that the "training" limitation referenced above is the additional element that under Prong Two integrates the supposed judicial exception into a practical application because this limitation provides the improvement of significantly reduced training effort. The rebuttal at
Examiner’s Response
The examiner respectfully “disagrees” with the applicant’s argument from Pages 8-12. First, the examiner recognizes that adding a second sensor to a sensor fusion system, when it improves the overall capabilities of the system are generally considered a technological improvement because sensor fusion is defined as the process of combining data from multiple sources to reduce uncertainty, improve accuracy, and increase reliability compared to any single sensor alone. In that context, the examiner had provided the responsive earlier that "Using multiple sensors is a sensor fusion which can be used to improve the accuracy, reliability, and robustness of a machine learning model and can be considered as merely applying generic function," "no special hardware is used in present context of the instance application specification," "If a hardware or software change to a second sensor improves the overall capabilities of a sensor fusion system, it may constitute a technological improvement as sensor fusion is fundamentally designed to integrate data from multiple sources to create a 'better' and more reliable output than any single sensor can provide on its own." On further review of the entire specification, the examiner reaffirms this observation that no such improvement to the accuracy, coverage, and robustness, reduced uncertainty, reliability (fault tolerance) is seen in the specification. The entire context of the invention is training using “sensor fusion” without presenting the “end solution” to support computer technology. Further within the context of any specific embedded application, the improvement can be demonstrated in terms of meaningful for the intended use (e.g., safety in autonomous driving, accuracy in robotics) Integration quality: The new sensor must be properly calibrated and integrated into the fusion algorithm, etc. Cost and complexity reduction could also be a technological improvement, practical adoption also depends on cost, size, power, and data processing requirements. No such improvements are demonstrated in the specification. To support this, the examiner cites the following specifications:
[0005] “application or independently of a specific application. Thus, for the specific application, a training may be carried out in a second training phase with only one second training task . This significantly reduces the training effort in an adaptation
[0005]” For deep neural networks, in particular, there is the possibility of easily adapting such an a priori optimized model for machine learning to a new training task . Fast in this case means, for example, using very few new characterized training data , in a short period of time and / or with little computing effort as opposed to the training that was necessary for the a priori optimization
[0007]” The first phase takes place, for example, with first training tasks , which originate from a generic application , in particular , offline . The second phase takes place , for example, for adaptation to a specific application with second training tasks , which originate from an operation of a specific application . The second training phase is carried out, for example , during operation of the application”
[0010 ] “ In accordance with an example embodiment of the present invention , a method is preferably provided for activating a computer - controlled machine , in particular , of an at least semi - autonomous robot , of a vehicle , of a home application , of a power tool , of a personal assistance system , of an access control system , training data for training tasks being generated as a function of digital sensor data , a device for machine learning , in particular , for regression and / or for classification , and / or another application that includes an artificial neural network , being trained with the aid of training tasks according to the described method , the computer - controlled machine being activated as a function of an output signal of the device thus trained . The training data are detected for the specific application and, in particular, used for training in the second training phase. This facilitates the adaptation of the artificial neural network”.
[0013] “It is preferably provided that the artificial neural network is trained with the first gradient - based learning algorithm as a function of the parameter set and as a function of a second training task. An adaptation to new training tasks may therefore be efficiently implemented”
[0088]” In the first phase , generic training data may be used for sensors of a particular sensor class , which includes , for example , sensor 106. Thus , when exchanging sensor 106 , artificial neural network may be easily adapted to a switch of a hardware or software generation through training in the second phase”
Regarding applicants’ argument on Page 11 “ For example, one improvement identified in the Specification is to "effectively learn new tasks in succession whilst protecting knowledge about previous tasks” interpreting that effectively learning new tasks in succession while protecting prior task knowledge is recognized as a technological improvement under § 101 when it improves the model’s operation and yields tangible benefits like reduced storage and complexity. The examiner submits that this a very “weak argument”. In comparison to a case as above, the examiner does not see any specific tangible improvement metrics claimed to support the comparison as provided in the comparison case. Regarding applicants’ argument on Page 12 that “examiners should not assert that an additional element (or combination of
elements) is well-understood, routine, or conventional unless the examiner finds, and expressly supports the rejection in writing”, the examiner “disagrees” with arguments and based on the above arguments reaffirms that the technique of “sensor fusion” is well-understood, routine and conventional for last two decades [See Meier, John, and Tirumale Ramesh. "Intelligent sensor fabric computing on a chip-a technology path for intelligent network computing." 2007 IEEE Aerospace Conference. IEEE, 2007]
In Conclusion, the examiner RETAINS the 101 rejections on independent claims 13, 18, 21, 23 and 24 and dependent claims 14-17, 19-20, and 22.
Regarding 103 rejections
- On Pages 12-13, the examiner recognizes the applicant has stated regarding claims 13 and 21-24 and argues that the prior arts Kabul, Liu or Chen do not teach the amended features.
Examiner’s Response
First, the examiner notes that the applicant has omitted to include claim 18 in the response which is also amended. However, the examiner recognizes that claim 18 is also amended for the purpose of the prosecution. The examiner has argued on Page 13 that the prior art does not teach the amended features and provides no other arguments. Applicants’ arguments with respect to the independent claims 13, 18, 21 and 23-24 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument.
As the applicant has amended as an RCE, the examiner has used new grounds of rejections to teach the amended claims and submits that independent claims 13, 18, 21, 23 and 24 and all dependent claims 14-17, 19-20 and 22 are rejected under 103 as NON-FINAL REJECTION.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 13-24 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1:
According to the first part of the analysis, in the instant case, claims 13, 18 and 21 are methods and thus claims 14-17, 19-20 and claim 22 are associated to a method of independent claims 13, 18 and 21. Claim 23 is directed to a device comprising one or more processors and memory and device and claim 24 is directed to a storage media with support for processor execution. Each of the claims falls within one of the four statutory categories (i.e. process, machine, manufacture, or composition of matter).
Regarding claim 13 (Currently Amended):
Step 2A Prong 1:
“determining a parameter set for an architecture and for weights of an artificial neural network
in a first phase with a first gradient-based learning algorithm” is a mental step of data determination.
“ and with a second gradient-based learning algorithm as a function of a plurality of first training
tasks from the distribution of training task” does not integrate the judicial exception into a practical application” are mental steps as a series of steps for the learning algorithm that does not require training.
“the second gradient-based learning algorithm being a meta-learning algorithm, which
ascertains an optimized parameter set as a function of the plurality of first training tasks and the
parameter set” are mental steps as a series of steps for the learning algorithm that does not require training.
“providing a plurality of training tasks from a distribution of training tasks, the training tasks characterizing processing of digital sensor data, each training task of the distribution of training tasks characterizing an assignment of the digital sensor data to a result of the processing **** same no change**
Additional Elements
Step 2A Prong 2:
“ A computer-implemented method for processing digital sensor data, the method
comprising the following steps:” recited in the preamble does not integrate the judicial exception into a practical application. This additional element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
“providing a plurality of training tasks from a distribution of training tasks” is data transmission (See MPEP 2106.05(d)) does not integrate the judicial exception into a practical application.
“the training tasks characterizing processing of digital sensor data” does not integrate the judicial exception into a practical application. This additional element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
“each training task of the distribution of training tasks characterizing an assignment of the digital sensor data to a result of the processing” does not integrate the judicial exception into a practical application. This additional element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
“the first gradient-based learning algorithm being a differentiable architecture search algorithm” does not integrate the judicial exception into a practical application. This additional element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
“training the artificial neural network in a second phase with the first gradient- based learning
algorithm and independently of the second gradient-based learning algorithm” does not
integrate the judicial exception into a practical application. This additional element is merely
using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
“ as a function of the optimized parameter set as a function of one second training
task from the plurality of training tasks from the distribution, the second training task being
different than the plurality of first training tasks” does not integrate the judicial exception into a
practical application. This additional element is merely using a computer as a tool to perform an
abstract idea. See MPEP 2106.05(h).
“and processing the digital sensor data as a function of the artificial neural network” does not integrate the judicial exception into a practical application. This additional element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
“ wherein the digital sensor data includes: first digital sensor data of the plurality of first training tasks, the first digital sensor data being acquired by a first sensor, and second digital sensor data of the second training task, the second digital sensor data being acquired by a second sensor” does not integrate the judicial exception into a practical application. This additional element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
does not integrate the judicial exception into a practical application. This additional element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
Step 2B:
“ A computer-implemented method for processing digital sensor data, the method
comprising the following steps:” recited in the preamble does not amount to significantly more than the judicial exception in the claim. This additional element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
“the training tasks characterizing processing of digital sensor data” does not amount to significantly more than the judicial exception in the claim. This additional element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
“each training task of the distribution of training tasks characterizing an assignment of the digital sensor data to a result of the processing” does not amount to significantly more than the judicial exception in the claim. This additional element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
“the first gradient-based learning algorithm being a differentiable architecture search algorithm” does not amount to significantly more than the judicial exception in the claim. This additional element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
“training the artificial neural network in a second phase with the first gradient- based learning
algorithm and independently of the second gradient-based learning algorithm” does not amount to significantly more than the judicial exception in the claim. This additional element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
“ as a function of the optimized parameter set as a function of one second training”
task from the plurality of training tasks from the distribution, the second training task being
different than the plurality of first training tasks” does not amount to significantly more than the judicial exception in the claim. This additional element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
“and processing the digital sensor data as a function of the artificial neural network” does not amount to significantly more than the judicial exception in the claim. This additional element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h
“ wherein the digital sensor data includes: first digital sensor data of the plurality of first training tasks, the first digital sensor data being acquired by a first sensor, and second digital sensor data of the second training task, the second digital sensor data being acquired by a second sensor” does not amount to significantly more than the judicial exception in the claim. This additional element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
Regarding claim 14 (Previously Presented):
Step 2A Prong 2:
“wherein the artificial neural network is defined by a plurality of layers, elements of the plurality of the layers including a shared input and defining a shared output, the architecture of the artificial neural network being defined, in addition to the weights for neurons in the elements, by parameters, each of the parameters characterizing a contribution of one of the elements of the plurality of layers to the output” does not integrate the judicial exception into a practical application. This additional element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
Step 2B:
“wherein the artificial neural network is defined by a plurality of layers, elements of the plurality of the layers including a shared input and defining a shared output, the architecture of the artificial neural network being defined, in addition to the weights for neurons in the elements, by parameters, each of the parameters characterizing a contribution of one of the elements of the plurality of layers to the output” does not amount to significantly more than the judicial exception in the claim. This additional element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
Regarding claim 15 (Previously Presented):
Step 2A Prong 2:
“ wherein the artificial neural network is trained in the second phase as a function of a second training task and as a function of the first gradient- based learning algorithm and independently of the second gradient-based learning algorithm” does not integrate the judicial exception into a practical application. This additional element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
Step 2B:
“wherein the artificial neural network is trained in the second phase as a function of a second training task and as a function of the first gradient- based learning algorithm and independently of the second gradient-based learning algorithm” does not amount to significantly more than the judicial exception in the claim. This additional element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
Regarding claim 16 (Previously Presented):
Step 2A Prong 2:
“wherein the artificial neural network is trained in the first phase as a function of the plurality of first training tasks, the artificial neural network being trained in the second phase as a function of a fraction of the training data from the second training task” does not integrate the judicial exception into a practical application. This element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
Step 2B:
“wherein the artificial neural network is trained in the first phase as a function of the plurality of first training tasks, the artificial neural network being trained in the second phase as a function of a fraction of the training data from the second training task” does not amount to significantly more than the judicial exception in the claim. This element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
Regarding claim 17 (Previously Presented):
Step 2A Prong 2:
“ wherein at least the parameters of the artificial neural network that define the architecture of the artificial neural network are trained with the second gradient-based learning algorithm”
does not integrate the judicial exception into a practical application. This element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
Step 2B:
“ wherein at least the parameters of the artificial neural network that define the architecture of the artificial neural network are trained with the second gradient-based learning algorithm”
does not amount to significantly more than the judicial exception in the claim. This element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
Regarding claim 18 (Previously Presented):
Step 2A Prong 1:
“determining a parameter set for an architecture and for weights of an artificial neural network
in a first phase with a first gradient-based learning algorithm” is a mental step of data determination.
“ and with a second gradient-based learning algorithm as a function of a plurality of first training
tasks from the distribution of training task” does not integrate the judicial exception into a practical application” is mental steps as a series of steps for the learning algorithm that does not require training.
“the second gradient-based learning algorithm being a meta-learning algorithm, which
ascertains an optimized parameter set as a function of the plurality of first training tasks and the
parameter set” is mental steps as a series of steps for the learning algorithm that does not require training.
Additional Elements
Step 2A Prong 2:
“A method for activating a computer-controlled machine, the method comprising the following steps: generating training data for training tasks as a function of digital sensor data; training a device which includes an artificial neural network by” recited in the preamble does not integrate the judicial exception into a practical application. This additional element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
“providing a plurality of training tasks from a distribution of training tasks” is data transmission (See MPEP 2106.05(d)) does not integrate the judicial exception into a practical application.
“the training tasks characterizing processing of digital sensor data” does not integrate the judicial exception into a practical application. This additional element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
“each training task of the distribution of training tasks characterizing an assignment of the digital sensor data to a result of the processing” does not integrate the judicial exception into a practical application. This additional element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
“the first gradient-based learning algorithm being a differentiable architecture search algorithm” does not integrate the judicial exception into a practical application. This additional element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
“training the artificial neural network in a second phase with the first gradient- based learning
algorithm and independently of the second gradient-based learning algorithm” does not
integrate the judicial exception into a practical application. This additional element is merely
using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
“as a function of the optimized parameter set as a function of one second training
task from the plurality of training tasks from the distribution, the second training task being
different than the plurality of first training tasks” does not integrate the judicial exception into a
practical application. This additional element is merely using a computer as a tool to perform an
abstract idea. See MPEP 2106.05(h).
“and processing the digital sensor data as a function of the artificial neural network” does not integrate the judicial exception into a practical application. This additional element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
“ wherein the digital sensor data includes: first digital sensor data of the plurality of first training tasks, the first digital sensor data being acquired by a first sensor, and second digital sensor data of the second training task, the second digital sensor data being acquired by a second sensor” does not integrate the judicial exception into a practical application. This additional element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
does not integrate the judicial exception into a practical application. This additional element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
Step 2B:
“A method for activating a computer-controlled machine, the method comprising the following steps: generating training data for training tasks as a function of digital sensor data; training a device which includes an artificial neural network by” recited in the preamble does not amount to significantly more than the judicial exception in the claim. This additional element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
“the training tasks characterizing processing of digital sensor data” does not amount to significantly more than the judicial exception in the claim. This additional element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
“each training task of the distribution of training tasks characterizing an assignment of the digital sensor data to a result of the processing” does not amount to significantly more than the judicial exception in the claim. This additional element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
“the first gradient-based learning algorithm being a differentiable architecture search algorithm” does not amount to significantly more than the judicial exception in the claim. This additional element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
“training the artificial neural network in a second phase with the first gradient- based learning
algorithm and independently of the second gradient-based learning algorithm” does not amount to significantly more than the judicial exception in the claim. This additional element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
“ as a function of the optimized parameter set as a function of one second training”
task from the plurality of training tasks from the distribution, the second training task being
different than the plurality of first training tasks” does not amount to significantly more than the judicial exception in the claim. This additional element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
“and processing the digital sensor data as a function of the artificial neural network” does not amount to significantly more than the judicial exception in the claim. This additional element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
“ wherein the digital sensor data includes: first digital sensor data of the plurality of first training tasks, the first digital sensor data being acquired by a first sensor, and second digital sensor data of the second training task, the second digital sensor data being acquired by a second sensor” does not amount to significantly more than the judicial exception in the claim. This additional element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
Regarding claim 20 (Previously Presented):
Step 2A Prong 2:
“ wherein the training data include image data, video data and/or digital sensor data of a sensor, from at least one camera and/or one infrared camera and/or one LIDAR sensor and/or one radar sensor and/or one acoustic sensor and/or one ultrasonic sensor and/or one receiver for a satellite navigation system and/or one rotational speed sensor and/or one torque sensor and/or one acceleration sensor and/or one position sensor” is directed to a field of use and the computer is used as a tool and does not integrate the judicial exception into a practical application. See MPEP 2106.05(h).
Step 2B:
“ wherein the training data include image data, video data and/or digital sensor data of a
sensor, from at least one camera and/or one infrared camera and/or one LIDAR sensor and/or
one radar sensor and/or one acoustic sensor and/or one ultrasonic sensor and/or one receiver
for a satellite navigation system and/or one rotational speed sensor and/or one torque sensor
and/or one acceleration sensor and/or one position sensor” is directed to a field of use and the computer is used as a tool and does not amount to significantly more than the judicial exception in the claim. See MPEP 2106.05(h).
Regarding claim 21(Currently Amended):
Step 2A Prong 1:
“determining a parameter set for an architecture and for weights of an artificial neural network
in a first phase with a first gradient-based learning algorithm” is a mental step of data determination.
“and a second gradient based learning algorithm as a function of a plurality of first training
tasks from the distribution of training task” are mental steps as a series of steps for the learning algorithm that does not require training.
“the second gradient-based learning algorithm being a meta-learning algorithm, which
ascertains an optimized parameter set as a function of the plurality of first training tasks and the
parameter set” are mental steps as a series of steps for the learning algorithm that does not require training.
Additional Elements
Step 2A Prong 2:
“ A computer-implemented method for training a device for machine learning, classification or activation of a computer-controlled machine, the method comprising the following steps;” recited in the preamble does not integrate the judicial exception into a practical application. This additional element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
“providing a plurality of training tasks from a distribution of training tasks” is data transmission (See MPEP 2106.05(d)) does not integrate the judicial exception into a practical application.
“the training tasks characterizing processing of digital sensor data” does not integrate the judicial exception into a practical application. This additional element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
“each training task of the distribution of training tasks characterizing an assignment of the digital sensor data to a result of the processing” does not integrate the judicial exception into a practical application. This additional element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
“the first gradient-based learning algorithm being a differentiable architecture search algorithm” does not integrate the judicial exception into a practical application. This additional element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
“training the artificial neural network in a second phase with the first gradient- based learning
algorithm and independently of the second gradient-based learning algorithm” does not
integrate the judicial exception into a practical application. This additional element is merely
using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
“ as a function of the optimized parameter set as a function of one second training
task from the plurality of training tasks from the distribution, the second training task being
different than the plurality of first training tasks” does not integrate the judicial exception into a
practical application. This additional element is merely using a computer as a tool to perform an
abstract idea. See MPEP 2106.05(h).
“and processing the digital sensor data as a function of the artificial neural network” does not integrate the judicial exception into a practical application. This additional element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
“ wherein the digital sensor data includes: first digital sensor data of the plurality of first training tasks, the first digital sensor data being acquired by a first sensor, and second digital sensor data of the second training task, the second digital sensor data being acquired by a second sensor” does not integrate the judicial exception into a practical application. This additional element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
does not integrate the judicial exception into a practical application. This additional element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
Step 2B:
“ A computer-implemented method for training a device for machine learning, classification or activation of a computer-controlled machine, the method comprising the following steps;”
recited in the preamble does not amount to significantly more than the judicial exception in the claim. This additional element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
“providing a plurality of training tasks from a distribution of training tasks” is data
transmission (See MPEP 2106.05(d)) does not amount to significantly more than the judicial exception in the claim. This additional element is merely using a computer as a tool to perform an abstract idea.
“the training tasks characterizing processing of digital sensor data” does not amount to significantly more than the judicial exception in the claim. This additional element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
“each training task of the distribution of training tasks characterizing an assignment of the digital sensor data to a result of the processing” does not amount to significantly more than the judicial exception in the claim. This additional element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
“the first gradient-based learning algorithm being a differentiable architecture search algorithm” does not amount to significantly more than the judicial exception in the claim. This additional element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
“training the artificial neural network in a second phase with the first gradient- based learning
algorithm and independently of the second gradient-based learning algorithm” does not amount to significantly more than the judicial exception in the claim. This additional element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h
“ as a function of the optimized parameter set as a function of one second training”
task from the plurality of training tasks from the distribution, the second training task being
different than the plurality of first training tasks” does not amount to significantly more than the judicial exception in the claim. This additional element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h.
“and processing the digital sensor data as a function of the artificial neural network” does not amount to significantly more than the judicial exception in the claim. This additional element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
“ wherein the digital sensor data includes: first digital sensor data of the plurality of first training tasks, the first digital sensor data being acquired by a first sensor, and second digital sensor data of the second training task, the second digital sensor data being acquired by a second sensor” does not amount to significantly more than the judicial exception in the claim. This additional element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
Regarding claim 22 (Previously Presented):
Step 2A Prong 2:
“wherein the artificial neural network is trained with the first gradient-based learning algorithm as a function of the parameter set and as a function of a second training task” does not integrate the judicial exception into a practical application. This element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
Step 2B:
“wherein the artificial neural network is trained with the first gradient-based learning algorithm as a function of the parameter set and as a function of a second training task” does not amount to significantly more than the judicial exception in the claim. This element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
Regarding claim 23 (Currently Amended):
Step 2A Prong 1:
“determine a parameter set for an architecture and for weights of an artificial neural network
in a first phase with a first gradient-based learning algorithm” is a mental step of data determination.
“and with a second gradient-based learning algorithm as a function of a plurality of first training
tasks from the distribution of training task” are mental step as a series of steps for the learning algorithm that does not require training.
“the second gradient-based learning algorithm being a meta-learning algorithm, which
ascertains an optimized parameter set as a function of the plurality of first training tasks and the
parameter set” are mental steps as a series of steps for the learning algorithm that does not require training.
Additional Elements
Step 2A Prong 2:
“A device for processing digital sensor data for machine learning, classification or activation of a computer-controlled machine, comprising: a processor; and memory for at least one artificial neural network; wherein the processor is configured to: “ recited in the preamble does not integrate the judicial exception into a practical application. This additional element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
“providing a plurality of training tasks from a distribution of training tasks” is data transmission (See MPEP 2106.05(d)) does not integrate the judicial exception into a practical application.
“the training tasks characterizing processing of digital sensor data” does not integrate the judicial exception into a practical application. This additional element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
“each training task of the distribution of training tasks characterizing an assignment of the digital sensor data to a result of the processing” does not integrate the judicial exception into a practical application. This additional element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
“the first gradient-based learning algorithm being a differentiable architecture search algorithm” does not integrate the judicial exception into a practical application. This additional element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
“training the artificial neural network in a second phase with the first gradient- based learning
algorithm and independently of the second gradient-based learning algorithm” does not
integrate the judicial exception into a practical application. This additional element is merely
using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
“ as a function of the optimized parameter set as a function of one second training
task from the plurality of training tasks from the distribution, the second training task being
different than the plurality of first training tasks” does not integrate the judicial exception into a
practical application. This additional element is merely using a computer as a tool to perform an
abstract idea. See MPEP 2106.05(h).
“and processing the digital sensor data as a function of the artificial neural network” does not integrate the judicial exception into a practical application. This additional element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
“ wherein the digital sensor data includes: first digital sensor data of the plurality of first training tasks, the first digital sensor data being acquired by a first sensor, and second digital sensor data of the second training task, the second digital sensor data being acquired by a second sensor” does not integrate the judicial exception into a practical application. This additional element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
Step 2B:
“A device for processing digital sensor data for machine learning, classification or activation of a computer-controlled machine, comprising: a processor; and memory for at least one artificial neural network; wherein the processor is configured to: “ recited in the preamble does not amount to significantly more than the judicial exception in the claim. This additional element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
“the training tasks characterizing processing of digital sensor data” does not amount to significantly more than the judicial exception in the claim. This additional element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
“each training task of the distribution of training tasks characterizing an assignment of the digital sensor data to a result of the processing” does not amount to significantly more than the judicial exception in the claim. This additional element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
“the first gradient-based learning algorithm being a differentiable architecture search algorithm” does not amount to significantly more than the judicial exception in the claim. This additional element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
“training the artificial neural network in a second phase with the first gradient- based learning
algorithm and independently of the second gradient-based learning algorithm” does not amount to significantly more than the judicial exception in the claim. This additional element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
“ as a function of the optimized parameter set as a function of one second training task from
the plurality of training tasks from the distribution, the second training task being
different than the plurality of first training tasks” does not amount to significantly more than the judicial exception in the claim. This additional element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
“and processing the digital sensor data as a function of the artificial neural network” does not amount to significantly more than the judicial exception in the claim. This additional element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
“ wherein the digital sensor data includes: first digital sensor data of the plurality of first training tasks, the first digital sensor data being acquired by a first sensor, and second digital sensor data of the second training task, the second digital sensor data being acquired by a second sensor” does not amount to significantly more than the judicial exception in the claim. This additional element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
Regarding claim 24 (Currently Amended):
Step 2A Prong 1:
“determining a parameter set for an architecture and for weights of an artificial neural network
in a first phase with a first gradient-based learning algorithm” is a mental step of data determination.
“and with a second gradient-based learning algorithm as a function of a plurality of first training
tasks from the distribution of training task” are mental steps as a series of steps for the learning algorithm that does not require training.
“the second gradient-based learning algorithm being a meta-learning algorithm, which
ascertains an optimized parameter set as a function of the plurality of first training tasks and the
parameter set” are mental steps as a series of steps for the learning algorithm that does not require training.
Additional Elements
Step 2A Prong 2:
“ A non-transitory machine-readable memory medium on which is stored a computer program for processing digital sensor data, the computer program, when executed by a computer, causing the computer to perform the following steps:” recited in the preamble does not integrate the judicial exception into a practical application. This additional element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
“providing a plurality of training tasks from a distribution of training tasks” is data transmission (See MPEP 2106.05(d)) does not integrate the judicial exception into a practical application.
“the training tasks characterizing processing of digital sensor data” does not integrate the judicial exception into a practical application. This additional element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
“each training task of the distribution of training tasks characterizing an assignment of the digital sensor data to a result of the processing” does not integrate the judicial exception into a practical application. This additional element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
“the first gradient-based learning algorithm being a differentiable architecture search algorithm” does not integrate the judicial exception into a practical application. This additional element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
“training the artificial neural network in a second phase with the first gradient- based learning
algorithm and independently of the second gradient-based learning algorithm” does not
integrate the judicial exception into a practical application. This additional element is merely
using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
“ as a function of the optimized parameter set as a function of one second training
task from the plurality of training tasks from the distribution, the second training task being
different than the plurality of first training tasks” does not integrate the judicial exception into a
practical application. This additional element is merely using a computer as a tool to perform an
abstract idea. See MPEP 2106.05(h).
“and processing the digital sensor data as a function of the artificial neural network” does not integrate the judicial exception into a practical application. This additional element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
“ wherein the digital sensor data includes: first digital sensor data of the plurality of first training tasks, the first digital sensor data being acquired by a first sensor, and second digital sensor data of the second training task, the second digital sensor data being acquired by a second sensor” does not integrate the judicial exception into a practical application. This additional element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
Step 2B:
“ A non-transitory machine-readable memory medium on which is stored a computer program for processing digital sensor data, the computer program, when executed by a computer, causing the computer to perform the following steps:” recited in the preamble does not amount to significantly more than the judicial exception in the claim. This additional element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
“the training tasks characterizing processing of digital sensor data” does not amount to significantly more than the judicial exception in the claim. This additional element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
“each training task of the distribution of training tasks characterizing an assignment of the digital sensor data to a result of the processing” does not amount to significantly more than the judicial exception in the claim. This additional element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
“the first gradient-based learning algorithm being a differentiable architecture search algorithm” does not amount to significantly more than the judicial exception in the claim. This additional element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
“training the artificial neural network in a second phase with the first gradient- based learning
algorithm and independently of the second gradient-based learning algorithm” does not amount to significantly more than the judicial exception in the claim. This additional element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
“ as a function of the optimized parameter set as a function of one second training”
task from the plurality of training tasks from the distribution, the second training task being
different than the plurality of first training tasks” does not amount to significantly more than the judicial exception in the claim. This additional element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
“and processing the digital sensor data as a function of the artificial neural network” does not amount to significantly more than the judicial exception in the claim. This additional element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
“ wherein the digital sensor data includes: first digital sensor data of the plurality of first training tasks, the first digital sensor data being acquired by a first sensor, and second digital sensor data of the second training task, the second digital sensor data being acquired by a second sensor” does not amount to significantly more than the judicial exception in the claim. This additional element is merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(h).
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 13-18, 21 and 23-24 are rejected under 35 U.S.C. 103 as being unpatentable over
Mustafa Kabul et.al. (hereinafter Kabul) US 2018/0307986 A1,
in view of Xin Chen et.al (hereinafter Chen2) Progressive Differentiable Architecture Search: Bridging the Depth Gap between Search and Evaluation, 2019 IEEE/CVF International Conference [Published in arXiv:1904.12760v1 [cs.CV] 29 Apr 2019].
Regarding claim 13 (Currently Amended):
Kabul discloses:
- A computer-implemented method for processing digital sensor data, the method
comprising the following steps:
[Abstract]:” A computing system provides distributed training of a neural network model”,
[0063]:” The data stored in input dataset 314 may be generated by and/or captured from a variety of sources including one or more sensors of the same or different type, one or more computing devices, etc.”.
- providing a plurality of training tasks from a distribution of training tasks, the training tasks
characterizing processing of digital sensor data; each training task of the distribution of training tasks characterizing an assignment of the digital sensor data to a result of the processing
[0064]:” Data stored in input dataset 314 may be sensor measurements or signal values captured by a sensor such as a camera, may be generated or captured in response to occurrence of an event or a transaction, generated by a device such as in response to an interaction by a user with the device, etc.
[0002] : “the computing device to provide distributed training of a neural network model. Explore phase options are distributed to each computing device of a plurality of computing devices. A subset of a training dataset is distributed to each computing device of a plurality of computing devices. A validation dataset is distributed to each computing device of the plurality of computing devices.
[BRI: In the context of assigning digital sensor data to processing results, each training
tasks in a distributed system is a specific computational unit that processes a defined subset of the data and produces a partial result, which is then combined with other results to form the final output. This is well known to a POSITA as a “distributed training” that assigns assigning a subset of the training dataset to each device and mapping that to a training task for data parallelism for matching which is the most common approach for large-scale training]
[0062]: “Input dataset 314 may be transposed. Input dataset 314 may include plurality of variables may define multiple dimensions for each observation vector. An observation vector x.sub.i may include a value for each of the plurality of variables associated with the observation i. Each variable of the plurality of variables may describe a characteristic of a physical object. For example, if input dataset 314 includes data related to operation of a vehicle, the variables may include an oil pressure, a speed, a gear indicator, a gas tank level, a tire pressure for each tire, an engine temperature, a radiator level, etc.
- determining a parameter set for an architecture and for weights of an artificial neural network in a first phase with a first gradient-based learning algorithm and with a second gradient-based learning algorithm as a function of a plurality of first training tasks from the distribution of training task
[Abstract]: “A computing system provides distributed training of a neural network model. Explore phase options, exploit phase options.
[0002]:” Explore phase options are distributed to each computing device of a plurality of computing devices”,
[0003]: ” Exploit phase options are distributed to each computing device of the plurality of computing devices”,
[0077]:” the goal of training is to determine a set of neural network weights that best predicts the targets in the training data”,
[0084]: “In the exploration phase, the worker nodes of worker system 106 perform local optimization on their copy of training dataset 414. Since their portion of input dataset 314 is a random sample of the entire set of training observations, its distribution is close to the distribution of the entire set of training observations such that the local optimization is valid”,
[0037]: “ Worker nodes compute the gradient on their local observations, and then enter into a collective aggregation operation, which computes a final gradient vector. The effective number of observations used to calculate the gradients is the total number of observations across the grid of worker nodes.
[BRI: within this context the first gradient and second gradient are associated with the first and second phases]
- training the artificial neural network in a second phase with the first gradient-based learning algorithm, and independently of the second gradient-based learning algorithm, as a function of the optimized parameter set as a function of one second training task from the plurality of training tasks from the distribution, the second training task being different than the plurality of first training tasks.
[Abstract]: “A computing system provides distributed training of a neural network model. Explore phase options, exploit phase options, a subset of a training dataset”,
[0028]: “ FIG. 21 illustrates an optimization problem improved by the neural network training system of FIG. 1”,
[0119]:” In an operation 638, next model configuration data to evaluate in a next iteration is computed based on the results received in operation 636 and the exploit phase options that include an optimization method and input parameters associated with the optimization method.
[0031]: “ Training a neural network is done by a data centric optimization algorithm.”,
[0197]: “ Prediction application 922, model training application 222, manager application 312, and/or model train/score application 412 may be the same or different applications that are integrated in various manners to execute a neural network model using input dataset 314 and/or second dataset 924.
[0176]:” Different parameter indices may be determined for each worker device 400 of worker system 106 for each iteration of operation 850, which is performed after operation 620 and before operation 622”,
[0146]” In an operation 734, a mini-batch number of observation vectors are randomly selected from training dataset 414 without replacement so that on successive iterations of operation 734 different observation vectors are selected. The mini-batch number of observation vectors may be defined from a mini-batch size parameter value included in the received exploit phase options. For illustration, the mini-batch number of observation vectors may be specified as an exploit phase option and may be greater than or equal to one. An illustrative default value may be five. The mini-batch number of observation vectors specified as an exploit phase option may be different than the mini-batch number of observation vectors specified as an explore phase option”,
[0202: “ Referring to FIG. 10, example operations of prediction application 922 are described. Additional, fewer, or different operations may be performed depending on the embodiment of prediction application 922. The order of presentation of the operations of FIG. 10 is not intended to be limiting. Although some of the operational flows are presented in sequence, the various operations may be performed in various repetitions, concurrently (in parallel, for example, using threads and/or a distributed computing system), and/or in other orders than those that are illustrated”,
[0085]:” The exploitation phase quickly progresses in the region defined by the exploration phase. A much larger mini-batch size may be used during the exploitation phase because the exploration phase identified a flat area where the solution generalizes well, reducing the variability in the gradient computation and resulting in a faster convergence”.
[0037]:”existing synchronous and asynchronous methods result in a generalization gap. Worker nodes compute the gradient on their local observations, and then enter into a collective aggregation operation, which computes a final gradient vector. The effective number of observations used to calculate the gradients is the total number of the observations across the grid of worker node”,
[0196]:” Dependent on the type of data stored in input dataset 314 and second dataset 924, prediction application 922 may identify anomalies as part of process control, for example, of a manufacturing process, for machine condition monitoring, for example, an electro-cardiogram device, for image classification, for intrusion detection, for fraud detection, etc. Some or all of the operations described herein may be embodied in prediction application 922”.
[BRI: in multi-embodiment prediction systems, different operations can be performed for a prediction application depending on the embodiment it represents, and this can include using a second training task (different prediction application operation) that is different from the plurality of first training tasks .This is implemented through embodiment-conditioned architectures, task-specific losses, and selective adaptation strategies, ensuring that the prediction model is both generalizable and tailored to the current embodiment’s needs
- and processing the digital sensor data as a function of the artificial neural network.
[0073]:” In an operation 508, a load of input dataset 314 may be requested. For example, user device 200 may request that input dataset 314 be loaded into a table that is ready for processing”,
[0143]:” the neural network model is re-initialized using the received weight parameter values for the links of the neural network, and processing continues in operation 710 to repeat the training process”,
- wherein the digital sensor data includes: first digital sensor data of the plurality of first training tasks, the first digital sensor data being acquired by a first sensor, and second digital sensor data of the second training task, the second digital sensor data being acquired by a second sensor,
[0063]:” The data stored in input dataset 314 may be generated by and/or captured from a variety of sources including one or more sensors of the same or different type”,
[0064]:” Data stored in input dataset 314 may be sensor measurements or signal values captured by a sensor such as a camera, may be generated or captured in response to occurrence of an event or a transaction, generated by a device such as in response to an interaction by a user with the device, etc. The data stored in input dataset 314 may include any type of content represented in any computer-readable format
[0064]” The content may include textual information, graphical information, image information, audio information, numeric information, etc. that further may be encoded using various encoding techniques as understood by a person of skill in the art. The data stored in input dataset 314 may be captured at different time points periodically, intermittently, when an event occurs, etc
[0065]: “input dataset 314 may include data captured at a high data rate such as 200 or more observations per second for one or more physical objects. For example, data stored in input dataset 314 may be generated as part of the Internet of Things (IoT), where things (e.g., machines, devices, phones, sensors) can be connected to networks and the data from these things collected and processed within the things and/or external to the things before being stored in input dataset 314. For example, the IoT can include sensors in many different devices and types of devices”,
[0062]:” Input dataset 314 may include labeled and/or unlabeled data. The plurality of variables may define multiple dimensions for each observation vector. An observation vector
x
i
may include a value for each of the plurality of variables associated with the observation i. Each variable of the plurality of variables may describe a characteristic of a physical object. For example, if input dataset 314 includes data related to operation of a vehicle, the variables may include an oil pressure, a speed, a gear indicator, a gas tank level, a tire pressure for each tire, an engine temperature, a radiator level, etc”,
[0062]: “ Input dataset 314 may include data captured as a function of time for one or more physical objects. As another example, input dataset 314 may include data related to images, where each row includes the pixels that define a single image. The images may be of any item for which image recognition or classification may be performed including, but not limited to, faces, objects”.
[0064]: “The data stored in input dataset 314 may be captured at different time points periodically, intermittently, when an event occurs, etc.”.
[BRI: Purpose of second digital sensor data is to provide additional information such as additional vehicle information that, when combined with the first dataset, improves the accuracy and completeness of the fused location and distance data for the first vehicle. The plurality of sensors represents first and second sensors. In the context of vehicle variables such as oil pressure, etc., represents locations of first and second sensors that monitors each of the variables acquiring the data for classification (tasks)]
- and wherein one of the second sensor corresponds to the first sensor with a software update, or the second sensor represents a hardware or software generation change compared to the first sensor.
[0063]: “The data stored in input dataset 314 may be generated by and/or captured from a variety of sources including one or more sensors of the same or different type, one or more computing devices, etc. The data stored in input dataset 314 may be received directly or indirectly from the source and may or may not be pre-processed in some manner.
[0063]: “The data may be organized using delimited fields, such as comma or space separated fields, fixed width fields, using a SAS® dataset, etc. The SAS dataset may be a SAS® file stored in a SAS® library that a SAS® software tool creates and processes. The SAS dataset contains data values that are organized as a table of observations (rows) and variables (columns) that can be processed by one or more SAS software tools”,
[0194]:” Referring to FIG. 9, a block diagram of a prediction device 900 is shown in accordance with an illustrative embodiment. Prediction device 900 may include a fourth input interface 902, a fourth output interface 904, a fourth communication interface 906, a fourth non-transitory computer-readable medium 908, a fourth processor 910, a prediction application 922”,
PNG
media_image1.png
562
821
media_image1.png
Greyscale
[0032]:” To make neural network predictions more accurate, the values of its parameters are updated iteratively. Before doing this, how much each value changes is computed. An efficient way of computing these values is done using a back-propagation method that computes a rate of change by going through each computation unit or node in the neural network, by calculating partial derivatives of the network error with respect to each parameter linking pairs of nodes, and by applying the chain rule to do the same thing to the next computation unit. With one backward pass, the gradient (rate of change) values for each of the parameters of the neural network is computed. These gradient values are input to the optimization method to make the actual parameter update. There are a large variety of different optimization methods, almost all of which depend on knowledge of the values of the gradient for each parameter in the neural network”,
[0065]: “ data stored in input dataset 314 may be generated as part of the Internet of Things (IoT), where things (e.g., machines, devices, phones, sensors) can be connected to networks and the data from these things collected and processed within the things and/or external to the things before being stored in input dataset 314. For example, the IoT can include sensors in many different devices and types of devices, and high value analytics can be applied to identify hidden relationships and drive increased efficiencies. This can apply to both big data analytics and real-time analytics. Some of these devices may be referred to as edge devices, and may involve edge computing circuitry. These devices may provide a variety of stored or generated data, such as network data or data specific to the network devices themselves”.
[BRI: Perhaps known to a POSITA that edge computing is already a core enabler for automotive multi-sensor systems, enabling real-time perception, fusion, and control without relying solely on cloud infrastructure. A software tool stores sensor data in a shared library, processes it through threads, and the user changes the data or mapping, the update can affect how the “second sensor” corresponds to the “first sensor.” The sensor data stored in a software library that is processed through threads under user control, and a software update provides changes to how a second sensor corresponds to or depends on a first sensor, depending on the system’s sensor-dependency architecture]
Kabul does not explicitly disclose:
- the first gradient-based learning algorithm being a differentiable architecture search algorithm
- the second gradient-based learning algorithm being a meta-learning algorithm,
which ascertains an optimized parameter set as a function of the plurality of first training tasks and the parameter set.
However, Chen2 discloses:
- the first gradient-based learning algorithm being a differentiable architecture search algorithm
[1, Page 1]:
PNG
media_image2.png
8
424
media_image2.png
Greyscale
Figure 1: Difference between DARTS and P-DARTS (our approach), with the former searching architectures in a shallow setting and evaluating them in a deep one, and the latter progressively increasing the searching depth, so as to bridge the depth gap between search and evaluation. Green and blue indicate search and evaluation, respectively”.
[3.2, Page 3]: “ we find that when searching on a deeper architecture, the differentiable approaches tend to bias towards the skip-connect operation, because it accelerates forward/backward propagation and often leads to the fastest way of gradient descent”.
[BRI: in DARTS, each stage of the search process is itself a gradient-based optimization step, so multi-stage search can be seen as multiple sequential gradient learning phases] P-DARTS uses the same bi-level optimization as DARTS, which requires computing two coupled gradients—one for network weights and one for architecture parameters at every search stage. P-DARTS does not change the gradient structure. Instead, it modifies: Search depth: progressively increases the number of cells during search to reduce the depth gap between search and evaluation]
- the second gradient-based learning algorithm being a meta-learning algorithm,
which ascertains an optimized parameter set as a function of the plurality of first training tasks and the parameter set;
[1. Page 1]: “Early works on NAS focused on the optimal configuration of layer type, filter size and number, activation function, etc., to construct a complete network”,
[1, Page 1]: “follow-up works started to explore the possibility of searching for network building blocks or so called cells with reinforcement learning (RL) [35, 37] and evolutionary algorithm”,
[1. Page 1]: “ The discovered cells are then stacked orderly to construct the network for specific tasks. However, those RL-based and EA-based approaches share a common pipeline to sample and evaluate (from scratch) numerous architectures in the search space”,
[3.2, Page 3]: “we find that when searching on a deeper architecture, the differentiable approaches tend to bias towards the skip-connect operation, because it accelerates forward/backward propagation and often leads to the fastest way of gradient descent.
[2, Page 3]: ” EA-based [22] and RL based [37] NAS approaches achieved state-of-the-art performance in image recognition, where architectures were sampled and evaluated from the search space under the guidance of an EA-based or RL-based meta-controller”.
[BRI: in the broad machine learning sense, the described method is a form of meta-learning, because it learns a controller (via EA or RL) that can adaptively select and evaluate architectures from a search space, with the goal of improving optimization performance. the “meta-learning” here is in the architecture search and controller learning, while the accelerated forward/backward methods are the inner optimization engine that the meta-learner exploits.
It would be obvious to one of ordinary skills in the art before the effective filing date of the present application to combine Kabul and Chen2.
Kabul teaches digital sensor, characterization and processing and teaches two phases of training and teaches gradient based meta learning.
Chen2 teaches gradient learning with differential architecture search and teaches the meta-learning algorithm.
One of ordinary skills would be motivated to combine Kabul and Chen2 that can achieve remarkable performance and efficiency improvement (Chen2 [1, Page 1])
Regarding claim 14 (Previously Presented):
Kabul does not explicitly disclose:
- wherein the artificial neural network is defined by a plurality of layers, elements of the plurality of the layers including a shared input and defining a shared output, the architecture of the artificial neural network being defined, in addition to the weights for neurons in the elements, by parameters, each of the parameters characterizing a contribution of one of the elements of the plurality of layers to the output.
However, Chen2 discloses:
- wherein the artificial neural network is defined by a plurality of layers, elements of the plurality of the layers including a shared input and defining a shared output, the architecture of the artificial neural network being defined, in addition to the weights for neurons in the elements, by parameters, each of the parameters characterizing a contribution of one of the elements of the plurality of layers to the output.
[3.1, Page 2]: “In this work, we leverage DARTS [18] as our baseline framework. Our goal is to search for a robust cell and apply it to a network of L cells. A cell is defined as a directed acyclic graph (DAG) of N nodes, {
x
0
,
x
1
··· ,
x
N
-
1
}, where each node is a network layer, i.e., performing a specific mathematical function”,
[4.4.1, Page 7]:” The architecture generated by stage 3 achieves the low est test error among others, which validates the effective ness of our scheme. From Figure 3 we can observe that”,
PNG
media_image3.png
595
1284
media_image3.png
Greyscale
Figure 3: Normal cells discovered by different search stages of P-DARTS and second order DARTS (DARTS V2). The depths of search networks are 5, 11 and 17 cells for stage 1, 2 and 3 of P-DARTS and 8 for DARTS V2. When the depth of the search network increases, more deep connections are preserved. Note that the operation on edge E (0,1) of stage 1 is a parameter-free skip connect, thus it is strictly not a deep connection.
[4.4.1, Page 8]:” these architectures share some common edges, for example sep_ conv 3×3 at edge
E
(
C
k
-
2
,
2
)
for stage 1, 2 and 3 and at edge
E
(
C
k
-
1
,
0
)
for stage 2, 3 and DARTS V2.
[BRI: Perhaps known to a POSITA that DARTS architectures share common edges in the search space defined over reusable computation cells creating shared input–output structures across layers. This is a natural outcome of the differentiable, cell-based search process]
[1, Page 2]: “the algorithm can be biased heavily towards skip-connect as it often leads to rapidest error decay during optimization,
[3.2, Page 3]: “Second, we find that when searching on a deeper architecture, the differentiable approaches tend to bias towards the skip-connect operation, because it accelerates forward/backward propagation and often leads to the fastest way of gradient descent”.
[BRI: In a neural network, each parameter is a number that characterizes how one neuron’s input contributes to another neuron’s output. These parameters are weights and biases. Certain optimization algorithms can be biased toward skip connections because these connections directly improve gradient flow and can accelerate error decay, which can make the optimizer more likely to favor them during training. Bias an additional parameter for each neuron that is added to the weighted sum of inputs before applying the activation function. In neural networks "sharing a layer" (also called parameter sharing or weight sharing))
It would be obvious to one of ordinary skills in the art before the effective filing date of the present application to combine Kabul and Chen2.
Kabul teaches digital sensor, characterization and processing and teaches two phases of training and teaches gradient based meta learning.
Chen2 teaches gradient learning with differential architecture search and teaches the meta-learning algorithm.
One of ordinary skills would be motivated to combine Kabul and Chen2 that can achieve remarkable performance and efficiency improvement (Chen2 [1, Page 1])
Regarding claim 15 (Previously Presented)
Kabul discloses:
- wherein the artificial neural network is trained in the second phase as a function of a second
training task and as a function of the first gradient- based learning algorithm and independently of the second gradient-based learning algorithm
[Abstract]: “ A computing system provides distributed training of a neural network model. Explore phase options, exploit phase options, a subset of a training dataset”,
[0028]: “ FIG. 21 illustrates an optimization problem improved by the neural network training system of FIG. 1”,
[0119]:” In an operation 638, next model configuration data to evaluate in a next iteration is computed based on the results received in operation 636 and the exploit phase options that include an optimization method and input parameters associated with the optimization method.
[0031]: “ Training a neural network is done by a data centric optimization algorithm.”,
[0197]: “ Prediction application 922, model training application 222, manager application 312, and/or model train/score application 412 may be the same or different applications that are integrated in various manners to execute a neural network model using input dataset 314 and/or second dataset 924.
[0176]:” Different parameter indices may be determined for each worker device 400 of worker system 106 for each iteration of operation 850, which is performed after operation 620 and before operation 622”,
[0146]” In an operation 734, a mini-batch number of observation vectors are randomly selected from training dataset 414 without replacement so that on successive iterations of operation 734 different observation vectors are selected. The mini-batch number of observation vectors may be defined from a mini-batch size parameter value included in the received exploit phase options. For illustration, the mini-batch number of observation vectors may be specified as an exploit phase option and may be greater than or equal to one. An illustrative default value may be five. The mini-batch number of observation vectors specified as an exploit phase option may be different than the mini-batch number of observation vectors specified as an explore phase option”,
[0202: “ Referring to FIG. 10, example operations of prediction application 922 are described. Additional, fewer, or different operations may be performed depending on the embodiment of prediction application 922. The order of presentation of the operations of FIG. 10 is not intended to be limiting. Although some of the operational flows are presented in sequence, the various operations may be performed in various repetitions, concurrently (in parallel, for example, using threads and/or a distributed computing system), and/or in other orders than those that are illustrated”,
[0085]:” The exploitation phase quickly progresses in the region defined by the exploration phase. A much larger mini-batch size may be used during the exploitation phase because the exploration phase identified a flat area where the solution generalizes well, reducing the variability in the gradient computation and resulting in a faster convergence”.
[0037]:”existing synchronous and asynchronous methods result in a generalization gap. Worker nodes compute the gradient on their local observations, and then enter into a collective aggregation operation, which computes a final gradient vector. The effective number of observations used to calculate the gradients is the total number of the observations across the grid of worker node”,
[0196]:” Dependent on the type of data stored in input dataset 314 and second dataset 924, prediction application 922 may identify anomalies as part of process control, for example, of a manufacturing process, for machine condition monitoring, for example, an electro-cardiogram device, for image classification, for intrusion detection, for fraud detection, etc. Some or all of the operations described herein may be embodied in prediction application 922”.
[BRI: in multi-embodiment prediction systems, different operations can be performed for a prediction application depending on the embodiment it represents, and this can include using a second training task (different prediction application operation) that is different from the plurality of first training tasks .This is implemented through embodiment-conditioned architectures, task-specific losses, and selective adaptation strategies, ensuring that the prediction model is both generalizable and tailored to the current embodiment’s needs]
[0085]:” The exploitation phase quickly progresses in the region defined by the exploration phase. A much larger mini-batch size may be used during the exploitation phase because the exploration phase identified a flat area where the solution generalizes well, reducing the variability in the gradient computation and resulting in a faster convergence.
[BRI: the second phase is a “exploitation phase”, the exploration phase (first phase) is a “first phase” that has associated “gradient learning (first gradient based). As the exploitation phase defined by the exploration phase, the second gradient is not utilized)
Regarding claim 16 (Previously Presented):
Kabul discloses:
- wherein the artificial neural network is trained in the first phase as a function of the plurality of first training tasks,
[0105]:” In an operation 614, the explore phase options may be received from user device 200 or directly from the user of user device 200”,
[0114]:”the determination that the exploration phase is complete may be based on a number of iterations of operation 630 in comparison to a maximum number of iterations specified as an explore phase option”,
[0076]: “ In an operation 514, a distribution of training dataset 414 to each worker device 400 of worker system 106 by or coordinated by controller device 104 is requested”,
[0076]:” Each distributed training dataset 414 distributed to worker device 400 may be different from every other such that selections are made without replacement”,
[0076]:”A size of each training dataset 414 may be approximately equal as in include approximately the same number of observations. For example, an entirety of a remainder of input dataset 314 after extracting validation dataset 416 is approximately equally distributed among the W number of nodes of worker system 106”,
[0079]:” The initial model weights may be randomly generated. The variable list defines the variables defined for each observation vector of training dataset 414 and validation dataset 416 to include in training the neural network model.
- the artificial neural network being trained in the second phase as a function of a fraction of the training data from the second training task.
[0003]:” Exploit phase options are distributed to each computing device of the plurality of computing devices. (e) Execution of the neural network model by the plurality of computing devices using the subset of the training dataset stored at each computing device of the plurality of computing devices is requested.
Regarding claim 17 (Previously Presented):
Kabul discloses:
- wherein at least the parameters of the artificial neural network that define the architecture of the artificial neural network are trained with the second gradient-based learning algorithm
[Abstract]: “A computing system provides distributed training of a neural network model. Explore phase options, exploit phase options”,
[0002]:” Explore phase options are distributed to each computing device of a plurality of computing devices”,
[0003]: ” Exploit phase options are distributed to each computing device of the plurality of computing devices”,
[0077]:” the goal of training is to determine a set of neural network weights that best predicts the targets in the training data”,
[0084]: “In the exploration phase, the worker nodes of worker system 106 perform local optimization on their copy of training dataset 414. Since their portion of input dataset 314 is a random sample of the entire set of training observations, its distribution is close to the distribution of the entire set of training observations such that the local optimization is valid”,
[0037]: “ Worker nodes compute the gradient on their local observations, and then enter into a collective aggregation operation, which computes a final gradient vector. The effective number of observations used to calculate the gradients is the total number of the observations across the grid of worker nodes.
[BRI: within this context the first gradient and second gradient are associated with the first and second phases]
[0085]: “The exploitation phase quickly progresses in the region defined by the exploration phase. A much larger mini-batch size may be used during the exploitation phase because the exploration phase identified a flat area where the solution generalizes well, reducing the variability in the gradient computation and resulting in a faster convergence.
Regarding claim 18 (Currently Amended):
Kabul discloses:
- A method for activating a computer-controlled machine, the method comprising the
following steps:
[0062]:
Each variable of the plurality of variables may describe a characteristic of a physical object. For example, if input dataset 314 includes data related to operation of a vehicle, the variables may include an oil pressure, a speed, a gear indicator, a gas tank level, a tire pressure for each tire, an engine temperature, a radiator level, etc.
[0037]: “ Worker nodes compute the gradient on their local observations, and then enter into a collective aggregation operation”,
[0085]: “parameter merging described below consolidates the information gained from each worker node's local optimization. The exploitation phase quickly progresses in the region defined by the exploitation phase”.
[BRI: in principle, a “worker node” can be any compute host that runs workloads, including a vehicle equipped with the necessary hardware, software, and network connectivity]
- generating training data for training tasks as a function of digital sensor data;
[0063]:” The data stored in input dataset 314 may be generated by and/or captured from a variety of sources including one or more sensors of the same or different type, one or more computing devices, etc.”,
[0061]: “ Model training application 222, manager application 312, and model train/score application 412 may be the same or different applications that are integrated in various manners to train a neural network model using input dataset 314 distributed across a plurality of computing devices that may include controller device 104 and/or worker system 106.
[BRI: the sensor data stored as an input dataset 314 is used as training data by the training application (training data)]
- training a device which includes an artificial neural network by: providing a plurality of training tasks from a distribution of training tasks, the training tasks
characterizing processing of digital sensor data; each training task of the distribution of training tasks characterizing an assignment of the digital sensor data to a result of the processing
[0064]:” Data stored in input dataset 314 may be sensor measurements or signal values captured by a sensor such as a camera, may be generated or captured in response to occurrence of an event or a transaction, generated by a device such as in response to an interaction by a user with the device, etc.
[0002] : “the computing device to provide distributed training of a neural network model. Explore phase options are distributed to each computing device of a plurality of computing devices. A subset of a training dataset is distributed to each computing device of a plurality of computing devices. A validation dataset is distributed to each computing device of the plurality of computing devices.
[BRI: In the context of assigning digital sensor data to processing results, each training
tasks in a distributed system is a specific computational unit that processes a defined subset of the data and produces a partial result, which is then combined with other results to form the final output. This is well known to a POSITA as a “distributed training” that assigns assigning a subset of the training dataset to each device and mapping that to a training task for data parallelism for matching which is the most common approach for large-scale training]
[0062]: “Input dataset 314 may be transposed. Input dataset 314 may include plurality of variables may define multiple dimensions for each observation vector. An observation vector
x
i
may include a value for each of the plurality of variables associated with the observation i. Each variable of the plurality of variables may describe a characteristic of a physical object. For example, if input dataset 314 includes data related to operation of a vehicle, the variables may include an oil pressure, a speed, a gear indicator, a gas tank level, a tire pressure for each tire, an engine temperature, a radiator level, etc.
- determining a parameter set for an architecture and for weights of an artificial neural network in a first phase with a first gradient-based learning algorithm and with a second gradient-based learning algorithm as a function of a plurality of first training tasks from the distribution of training task
[Abstract]: “A computing system provides distributed training of a neural network model. Explore phase options, exploit phase options.
[0002]:” Explore phase options are distributed to each computing device of a plurality of computing devices”,
[0003]: ” Exploit phase options are distributed to each computing device of the plurality of computing devices”,
[0077]:” the goal of training is to determine a set of neural network weights that best predicts the targets in the training data”,
[0084]: “ In the exploration phase, the worker nodes of worker system 106 perform local optimization on their copy of training dataset 414. Since their portion of input dataset 314 is a random sample of the entire set of training observations, its distribution is close to the distribution of the entire set of training observations such that the local optimization is valid”,
[0037]: “ Worker nodes compute the gradient on their local observations, and then enter into a collective aggregation operation, which computes a final gradient vector. The effective number of observations used to calculate the gradients is the total number of the observations across the grid of worker nodes.
[BRI: within this context the first gradient and second gradient are associated with the first and second phases]
- training the artificial neural network in a second phase with the first gradient-based learning algorithm, and independently of the second gradient-based learning algorithm, as a function of the optimized parameter set as a function of one second training task from the plurality of training tasks from the distribution, the second training task being different than the plurality of first training tasks.
[Abstract]: “A computing system provides distributed training of a neural network model. Explore phase options, exploit phase options, a subset of a training dataset”,
[0028]: “ FIG. 21 illustrates an optimization problem improved by the neural network training system of FIG. 1”,
[0119]:” In an operation 638, next model configuration data to evaluate in a next iteration is computed based on the results received in operation 636 and the exploit phase options that include an optimization method and input parameters associated with the optimization method.
[0031]: “ Training a neural network is done by a data centric optimization algorithm.”,
[0197]: “ Prediction application 922, model training application 222, manager application 312, and/or model train/score application 412 may be the same or different applications that are integrated in various manners to execute a neural network model using input dataset 314 and/or second dataset 924.
[0176]:” Different parameter indices may be determined for each worker device 400 of worker system 106 for each iteration of operation 850, which is performed after operation 620 and before operation 622”,
[0146]” In an operation 734, a mini-batch number of observation vectors are randomly selected from training dataset 414 without replacement so that on successive iterations of operation 734 different observation vectors are selected. The mini-batch number of observation vectors may be defined from a mini-batch size parameter value included in the received exploit phase options. For illustration, the mini-batch number of observation vectors may be specified as an exploit phase option and may be greater than or equal to one. An illustrative default value may be five. The mini-batch number of observation vectors specified as an exploit phase option may be different than the mini-batch number of observation vectors specified as an explore phase option”,
[0202: “ Referring to FIG. 10, example operations of prediction application 922 are described. Additional, fewer, or different operations may be performed depending on the embodiment of prediction application 922. The order of presentation of the operations of FIG. 10 is not intended to be limiting. Although some of the operational flows are presented in sequence, the various operations may be performed in various repetitions, concurrently (in parallel, for example, using threads and/or a distributed computing system), and/or in other orders than those that are illustrated”,
[0085]:” The exploitation phase quickly progresses in the region defined by the exploration phase. A much larger mini-batch size may be used during the exploitation phase because the exploration phase identified a flat area where the solution generalizes well, reducing the variability in the gradient computation and resulting in a faster convergence”.
[0037]:”existing synchronous and asynchronous methods result in a generalization gap. Worker nodes compute the gradient on their local observations, and then enter into a collective aggregation operation, which computes a final gradient vector. The effective number of observations used to calculate the gradients is the total number of the observations across the grid of worker node”,
[0196]:” Dependent on the type of data stored in input dataset 314 and second dataset 924, prediction application 922 may identify anomalies as part of process control, for example, of a manufacturing process, for machine condition monitoring, for example, an electro-cardiogram device, for image classification, for intrusion detection, for fraud detection, etc. Some or all of the operations described herein may be embodied in prediction application 922”.
[BRI: in multi-embodiment prediction systems, different operations can be performed for a prediction application depending on the embodiment it represents, and this can include using a second training task (different prediction application operation) that is different from the plurality of first training tasks .This is implemented through embodiment-conditioned architectures, task-specific losses, and selective adaptation strategies, ensuring that the prediction model is both generalizable and tailored to the current embodiment’s needs
- and processing the digital sensor data as a function of the artificial neural network.
[0073]:” In an operation 508, a load of input dataset 314 may be requested. For example, user device 200 may request that input dataset 314 be loaded into a table that is ready for processing”,
[0143]:” the neural network model is re-initialized using the received weight parameter values for the links of the neural network, and processing continues in operation 710 to repeat the training process”,
- wherein the digital sensor data includes: first digital sensor data of the plurality of first training tasks, the first digital sensor data being acquired by a first sensor, and second digital sensor data of the second training task, the second digital sensor data being acquired by a second sensor,
[0063]:” The data stored in input dataset 314 may be generated by and/or captured from a variety of sources including one or more sensors of the same or different type”,
[0064]:” Data stored in input dataset 314 may be sensor measurements or signal values captured by a sensor such as a camera, may be generated or captured in response to occurrence of an event or a transaction, generated by a device such as in response to an interaction by a user with the device, etc. The data stored in input dataset 314 may include any type of content represented in any computer-readable format
[0064]” The content may include textual information, graphical information, image information, audio information, numeric information, etc. that further may be encoded using various encoding techniques as understood by a person of skill in the art. The data stored in input dataset 314 may be captured at different time points periodically, intermittently, when an event occurs, etc
[0065]: “input dataset 314 may include data captured at a high data rate such as 200 or more observations per second for one or more physical objects. For example, data stored in input dataset 314 may be generated as part of the Internet of Things (IoT), where things (e.g., machines, devices, phones, sensors) can be connected to networks and the data from these things collected and processed within the things and/or external to the things before being stored in input dataset 314. For example, the IoT can include sensors in many different devices and types of devices”,
[0062]:” Input dataset 314 may include labeled and/or unlabeled data. The plurality of variables may define multiple dimensions for each observation vector. An observation vector
x
i
may include a value for each of the plurality of variables associated with the observation i. Each variable of the plurality of variables may describe a characteristic of a physical object. For example, if input dataset 314 includes data related to operation of a vehicle, the variables may include an oil pressure, a speed, a gear indicator, a gas tank level, a tire pressure for each tire, an engine temperature, a radiator level, etc”,
[0062]: “ Input dataset 314 may include data captured as a function of time for one or more physical objects. As another example, input dataset 314 may include data related to images, where each row includes the pixels that define a single image. The images may be of any item for which image recognition or classification may be performed including, but not limited to, faces, objects”.
[0064]: “The data stored in input dataset 314 may be captured at different time points periodically, intermittently, when an event occurs, etc.”.
[BRI: Purpose of second digital sensor data is to provide additional information such as additional vehicle information that, when combined with the first dataset, improves the accuracy and completeness of the fused location and distance data for the first vehicle. The plurality of sensors represents first and second sensors. In the context of vehicle variables such as oil pressure, etc., represents locations of first and second sensors that monitors each of the variables acquiring the data for classification (tasks)]
- and wherein one of the second sensor corresponds to the first sensor with a software update, or the second sensor represents a hardware or software generation change compared to the first sensor.
[0063]: “The data stored in input dataset 314 may be generated by and/or captured from a variety of sources including one or more sensors of the same or different type, one or more computing devices, etc. The data stored in input dataset 314 may be received directly or indirectly from the source and may or may not be pre-processed in some manner.
[0063]: “The data may be organized using delimited fields, such as comma or space separated fields, fixed width fields, using a SAS® dataset, etc. The SAS dataset may be a SAS® file stored in a SAS® library that a SAS® software tool creates and processes. The SAS dataset contains data values that are organized as a table of observations (rows) and variables (columns) that can be processed by one or more SAS software tools”,
[0194]:” Referring to FIG. 9, a block diagram of a prediction device 900 is shown in accordance with an illustrative embodiment. Prediction device 900 may include a fourth input interface 902, a fourth output interface 904, a fourth communication interface 906, a fourth non-transitory computer-readable medium 908, a fourth processor 910, a prediction application 922”,
PNG
media_image1.png
562
821
media_image1.png
Greyscale
[0032]:” To make neural network predictions more accurate, the values of its parameters are updated iteratively. Before doing this, how much each value changes is computed. An efficient way of computing these values is done using a back-propagation method that computes a rate of change by going through each computation unit or node in the neural network, by calculating partial derivatives of the network error with respect to each parameter linking pairs of nodes, and by applying the chain rule to do the same thing to the next computation unit. With one backward pass, the gradient (rate of change) values for each of the parameters of the neural network is computed. These gradient values are input to the optimization method to make the actual parameter update. There are a large variety of different optimization methods, almost all of which depend on knowledge of the values of the gradient for each parameter in the neural network”,
[0065]: “ data stored in input dataset 314 may be generated as part of the Internet of Things (IoT), where things (e.g., machines, devices, phones, sensors) can be connected to networks and the data from these things collected and processed within the things and/or external to the things before being stored in input dataset 314. For example, the IoT can include sensors in many different devices and types of devices, and high value analytics can be applied to identify hidden relationships and drive increased efficiencies. This can apply to both big data analytics and real-time analytics. Some of these devices may be referred to as edge devices, and may involve edge computing circuitry. These devices may provide a variety of stored or generated data, such as network data or data specific to the network devices themselves”.
[BRI: Perhaps known to a POSITA that edge computing is already a core enabler for automotive multi-sensor systems, enabling real-time perception, fusion, and control without relying solely on cloud infrastructure. A software tool stores sensor data in a shared library, processes it through threads, and the user changes the data or mapping, the update can affect how the “second sensor” corresponds to the “first sensor.” The sensor data stored in a software library that is processed through threads under user control, and a software update provides changes to how a second sensor corresponds to or depends on a first sensor, depending on the system’s sensor-dependency architecture]
Kabul does not explicitly disclose:
- the first gradient-based learning algorithm being a differentiable architecture search algorithm
- the second gradient-based learning algorithm being a meta-learning algorithm,
which ascertains an optimized parameter set as a function of the plurality of first training tasks and the parameter set.
However, Chen2 discloses:
- the first gradient-based learning algorithm being a differentiable architecture search algorithm
[1, Page 1]:
PNG
media_image2.png
8
424
media_image2.png
Greyscale
Figure 1: Difference between DARTS and P-DARTS (our approach), with the former searching architectures in a shallow setting and evaluating them in a deep one, and the latter progressively increasing the searching depth, so as to bridge the depth gap between search and evaluation. Green and blue indicate search and evaluation, respectively”.
[3.2, Page 3]: “ we find that when searching on a deeper architecture, the differentiable approaches tend to bias towards the skip-connect operation, because it accelerates forward/backward propagation and often leads to the fastest way of gradient descent”.
[BRI: in DARTS, each stage of the search process is itself a gradient-based optimization step, so multi-stage search can be seen as multiple sequential gradient learning phases] P-DARTS uses the same bi-level optimization as DARTS, which requires computing two coupled gradients—one for network weights and one for architecture parameters at every search stage. P-DARTS does not change the gradient structure. Instead, it modifies: Search depth: progressively increases the number of cells during search to reduce the depth gap between search and evaluation]
- the second gradient-based learning algorithm being a meta-learning algorithm,
which ascertains an optimized parameter set as a function of the plurality of first training tasks and the parameter set;
[1. Page 1]: “Early works on NAS focused on the optimal configuration of layer type, filter size and number, activation function, etc., to construct a complete network”,
[1, Page 1]: “follow-up works started to explore the possibility of searching for network building blocks or so called cells with reinforcement learning (RL) [35, 37] and evolutionary algorithm”,
[1. Page 1]: “ The discovered cells are then stacked orderly to construct the network for specific tasks. However, those RL-based and EA-based approaches share a common pipeline to sample and evaluate (from scratch) numerous architectures in the search space”,
[3.2, Page 3]: “we find that when searching on a deeper architecture, the differentiable approaches tend to bias towards the skip-connect operation, because it accelerates forward/backward propagation and often leads to the fastest way of gradient descent.
[2, Page 3]: ” EA-based [22] and RL based [37] NAS approaches achieved state-of-the-art performance in image recognition, where architectures were sampled and evaluated from the search space under the guidance of an EA-based or RL-based meta-controller”.
[BRI: in the broad machine learning sense, the described method is a form of meta-learning, because it learns a controller (via EA or RL) that can adaptively select and evaluate architectures from a search space, with the goal of improving optimization performance. the “meta-learning” here is in the architecture search and controller learning, while the accelerated forward/backward methods are the inner optimization engine that the meta-learner exploits.
It would be obvious to one of ordinary skills in the art before the effective filing date of the present application to combine Kabul and Chen2.
Kabul teaches digital sensor, characterization and processing and teaches two phases of training and teaches gradient based meta learning.
Chen2 teaches gradient learning with differential architecture search and teaches the meta-learning algorithm.
One of ordinary skills would be motivated to combine Kabul and Chen2 that can achieve remarkable performance and efficiency improvement (Chen2 [1, Page 1])
Regarding claim 21 (Currently Amended):
Kabul discloses:
- A computer-implemented method for training a device for machine learning, classification or activation of a computer-controlled machine, the method comprising the following steps:
[0062] :” input dataset 314 may include data related to images, where each row includes the pixels that define a single image. The images may be of any item for which image recognition or classification may be performed including, but not limited to, faces, objects, alphanumeric letters, terrain, plants, animals, etc.
- providing a plurality of training tasks from a distribution of training tasks, the training tasks
characterizing processing of digital sensor data; each training task of the distribution of training tasks characterizing an assignment of the digital sensor data to a result of the processing
[0064]:” Data stored in input dataset 314 may be sensor measurements or signal values captured by a sensor such as a camera, may be generated or captured in response to occurrence of an event or a transaction, generated by a device such as in response to an interaction by a user with the device, etc.
[0002] : “the computing device to provide distributed training of a neural network model. Explore phase options are distributed to each computing device of a plurality of computing devices. A subset of a training dataset is distributed to each computing device of a plurality of computing devices. A validation dataset is distributed to each computing device of the plurality of computing devices.
[BRI: In the context of assigning digital sensor data to processing results, each training
tasks in a distributed system is a specific computational unit that processes a defined subset of the data and produces a partial result, which is then combined with other results to form the final output. This is well known to a POSITA as a “distributed training” that assigns assigning a subset of the training dataset to each device and mapping that to a training task for data parallelism for matching which is the most common approach for large-scale training]
[0062]: “Input dataset 314 may be transposed. Input dataset 314 may include plurality of variables may define multiple dimensions for each observation vector. An observation vector x.sub.i may include a value for each of the plurality of variables associated with the observation i. Each variable of the plurality of variables may describe a characteristic of a physical object. For example, if input dataset 314 includes data related to operation of a vehicle, the variables may include an oil pressure, a speed, a gear indicator, a gas tank level, a tire pressure for each tire, an engine temperature, a radiator level, etc.
- determining a parameter set for an architecture and for weights of an artificial neural network in a first phase with a first gradient-based learning algorithm and with a second gradient-based learning algorithm as a function of a plurality of first training tasks from the distribution of training task
[Abstract]: “A computing system provides distributed training of a neural network model. Explore phase options, exploit phase options.
[0002]:” Explore phase options are distributed to each computing device of a plurality of computing devices”,
[0003]: ” Exploit phase options are distributed to each computing device of the plurality of computing devices”,
[0077]:” the goal of training is to determine a set of neural network weights that best predicts the targets in the training data”,
[0084]: “ In the exploration phase, the worker nodes of worker system 106 perform local optimization on their copy of training dataset 414. Since their portion of input dataset 314 is a random sample of the entire set of training observations, its distribution is close to the distribution of the entire set of training observations such that the local optimization is valid”,
[0037]: “ Worker nodes compute the gradient on their local observations, and then enter into a collective aggregation operation, which computes a final gradient vector. The effective number of observations used to calculate the gradients is the total number of the observations across the grid of worker nodes.
[BRI: within this context the first gradient and second gradient are associated with the first and second phases]
- training the artificial neural network in a second phase with the first gradient-based learning algorithm, and independently of the second gradient-based learning algorithm, as a function of the optimized parameter set as a function of one second training task from the plurality of training tasks from the distribution, the second training task being different than the plurality of first training tasks.
[Abstract]: “A computing system provides distributed training of a neural network model. Explore phase options, exploit phase options, a subset of a training dataset”,
[0028]: “ FIG. 21 illustrates an optimization problem improved by the neural network training system of FIG. 1”,
[0119]:” In an operation 638, next model configuration data to evaluate in a next iteration is computed based on the results received in operation 636 and the exploit phase options that include an optimization method and input parameters associated with the optimization method.
[0031]: “ Training a neural network is done by a data centric optimization algorithm.”,
[0197]: “ Prediction application 922, model training application 222, manager application 312, and/or model train/score application 412 may be the same or different applications that are integrated in various manners to execute a neural network model using input dataset 314 and/or second dataset 924.
[0176]:” Different parameter indices may be determined for each worker device 400 of worker system 106 for each iteration of operation 850, which is performed after operation 620 and before operation 622”,
[0146]” In an operation 734, a mini-batch number of observation vectors are randomly selected from training dataset 414 without replacement so that on successive iterations of operation 734 different observation vectors are selected. The mini-batch number of observation vectors may be defined from a mini-batch size parameter value included in the received exploit phase options. For illustration, the mini-batch number of observation vectors may be specified as an exploit phase option and may be greater than or equal to one. An illustrative default value may be five. The mini-batch number of observation vectors specified as an exploit phase option may be different than the mini-batch number of observation vectors specified as an explore phase option”,
[0202: “ Referring to FIG. 10, example operations of prediction application 922 are described. Additional, fewer, or different operations may be performed depending on the embodiment of prediction application 922. The order of presentation of the operations of FIG. 10 is not intended to be limiting. Although some of the operational flows are presented in sequence, the various operations may be performed in various repetitions, concurrently (in parallel, for example, using threads and/or a distributed computing system), and/or in other orders than those that are illustrated”,
[0085]:” The exploitation phase quickly progresses in the region defined by the exploration phase. A much larger mini-batch size may be used during the exploitation phase because the exploration phase identified a flat area where the solution generalizes well, reducing the variability in the gradient computation and resulting in a faster convergence”.
[0037]:”existing synchronous and asynchronous methods result in a generalization gap. Worker nodes compute the gradient on their local observations, and then enter into a collective aggregation operation, which computes a final gradient vector. The effective number of observations used to calculate the gradients is the total number of the observations across the grid of worker node”,
[0196]:” Dependent on the type of data stored in input dataset 314 and second dataset 924, prediction application 922 may identify anomalies as part of process control, for example, of a manufacturing process, for machine condition monitoring, for example, an electro-cardiogram device, for image classification, for intrusion detection, for fraud detection, etc. Some or all of the operations described herein may be embodied in prediction application 922”.
[BRI: in multi-embodiment prediction systems, different operations can be performed for a prediction application depending on the embodiment it represents, and this can include using a second training task (different prediction application operation) that is different from the plurality of first training tasks .This is implemented through embodiment-conditioned architectures, task-specific losses, and selective adaptation strategies, ensuring that the prediction model is both generalizable and tailored to the current embodiment’s needs
- and processing the digital sensor data as a function of the artificial neural network.
[0073]:” In an operation 508, a load of input dataset 314 may be requested. For example, user device 200 may request that input dataset 314 be loaded into a table that is ready for processing”,
[0143]:” the neural network model is re-initialized using the received weight parameter values for the links of the neural network, and processing continues in operation 710 to repeat the training process”,
- wherein the digital sensor data includes: first digital sensor data of the plurality of first training tasks, the first digital sensor data being acquired by a first sensor, and second digital sensor data of the second training task, the second digital sensor data being acquired by a second sensor,
[0063]:” The data stored in input dataset 314 may be generated by and/or captured from a variety of sources including one or more sensors of the same or different type”,
[0064]:” Data stored in input dataset 314 may be sensor measurements or signal values captured by a sensor such as a camera, may be generated or captured in response to occurrence of an event or a transaction, generated by a device such as in response to an interaction by a user with the device, etc. The data stored in input dataset 314 may include any type of content represented in any computer-readable format
[0064]” The content may include textual information, graphical information, image information, audio information, numeric information, etc. that further may be encoded using various encoding techniques as understood by a person of skill in the art. The data stored in input dataset 314 may be captured at different time points periodically, intermittently, when an event occurs, etc
[0065]: “input dataset 314 may include data captured at a high data rate such as 200 or more observations per second for one or more physical objects. For example, data stored in input dataset 314 may be generated as part of the Internet of Things (IoT), where things (e.g., machines, devices, phones, sensors) can be connected to networks and the data from these things collected and processed within the things and/or external to the things before being stored in input dataset 314. For example, the IoT can include sensors in many different devices and types of devices”,
[0062]:” Input dataset 314 may include labeled and/or unlabeled data. The plurality of variables may define multiple dimensions for each observation vector. An observation vector
x
i
may include a value for each of the plurality of variables associated with the observation i. Each variable of the plurality of variables may describe a characteristic of a physical object. For example, if input dataset 314 includes data related to operation of a vehicle, the variables may include an oil pressure, a speed, a gear indicator, a gas tank level, a tire pressure for each tire, an engine temperature, a radiator level, etc”,
[0062]: “ Input dataset 314 may include data captured as a function of time for one or more physical objects. As another example, input dataset 314 may include data related to images, where each row includes the pixels that define a single image. The images may be of any item for which image recognition or classification may be performed including, but not limited to, faces, objects”.
[0064]: “The data stored in input dataset 314 may be captured at different time points periodically, intermittently, when an event occurs, etc.”.
[BRI: Purpose of second digital sensor data is to provide additional information such as additional vehicle information that, when combined with the first dataset, improves the accuracy and completeness of the fused location and distance data for the first vehicle. The plurality of sensors represents first and second sensors. In the context of vehicle variables such as oil pressure, etc., represents locations of first and second sensors that monitors each of the variables acquiring the data for classification (tasks)]
- and wherein one of the second sensor corresponds to the first sensor with a software update, or the second sensor represents a hardware or software generation change compared to the first sensor.
[0063]: “The data stored in input dataset 314 may be generated by and/or captured from a variety of sources including one or more sensors of the same or different type, one or more computing devices, etc. The data stored in input dataset 314 may be received directly or indirectly from the source and may or may not be pre-processed in some manner.
[0063]: “The data may be organized using delimited fields, such as comma or space separated fields, fixed width fields, using a SAS® dataset, etc. The SAS dataset may be a SAS® file stored in a SAS® library that a SAS® software tool creates and processes. The SAS dataset contains data values that are organized as a table of observations (rows) and variables (columns) that can be processed by one or more SAS software tools”,
[0194]:” Referring to FIG. 9, a block diagram of a prediction device 900 is shown in accordance with an illustrative embodiment. Prediction device 900 may include a fourth input interface 902, a fourth output interface 904, a fourth communication interface 906, a fourth non-transitory computer-readable medium 908, a fourth processor 910, a prediction application 922”,
PNG
media_image1.png
562
821
media_image1.png
Greyscale
[0032]:” To make neural network predictions more accurate, the values of its parameters are updated iteratively. Before doing this, how much each value changes is computed. An efficient way of computing these values is done using a back-propagation method that computes a rate of change by going through each computation unit or node in the neural network, by calculating partial derivatives of the network error with respect to each parameter linking pairs of nodes, and by applying the chain rule to do the same thing to the next computation unit. With one backward pass, the gradient (rate of change) values for each of the parameters of the neural network is computed. These gradient values are input to the optimization method to make the actual parameter update. There are a large variety of different optimization methods, almost all of which depend on knowledge of the values of the gradient for each parameter in the neural network”,
[0065]: “ data stored in input dataset 314 may be generated as part of the Internet of Things (IoT), where things (e.g., machines, devices, phones, sensors) can be connected to networks and the data from these things collected and processed within the things and/or external to the things before being stored in input dataset 314. For example, the IoT can include sensors in many different devices and types of devices, and high value analytics can be applied to identify hidden relationships and drive increased efficiencies. This can apply to both big data analytics and real-time analytics. Some of these devices may be referred to as edge devices, and may involve edge computing circuitry. These devices may provide a variety of stored or generated data, such as network data or data specific to the network devices themselves”.
[BRI: Perhaps known to a POSITA that edge computing is already a core enabler for automotive multi-sensor systems, enabling real-time perception, fusion, and control without relying solely on cloud infrastructure. A software tool stores sensor data in a shared library, processes it through threads, and the user changes the data or mapping, the update can affect how the “second sensor” corresponds to the “first sensor.” The sensor data stored in a software library that is processed through threads under user control, and a software update provides changes to how a second sensor corresponds to or depends on a first sensor, depending on the system’s sensor-dependency architecture]
Kabul does not explicitly disclose:
- the first gradient-based learning algorithm being a differentiable architecture search algorithm
- the second gradient-based learning algorithm being a meta-learning algorithm,
which ascertains an optimized parameter set as a function of the plurality of first training tasks and the parameter set.
However, Chen2 discloses:
- the first gradient-based learning algorithm being a differentiable architecture search algorithm
[1, Page 1]:
PNG
media_image2.png
8
424
media_image2.png
Greyscale
Figure 1: Difference between DARTS and P-DARTS (our approach), with the former searching architectures in a shallow setting and evaluating them in a deep one, and the latter progressively increasing the searching depth, so as to bridge the depth gap between search and evaluation. Green and blue indicate search and evaluation, respectively”.
[3.2, Page 3]: “ we find that when searching on a deeper architecture, the differentiable approaches tend to bias towards the skip-connect operation, because it accelerates forward/backward propagation and often leads to the fastest way of gradient descent”.
[BRI: in DARTS, each stage of the search process is itself a gradient-based optimization step, so multi-stage search can be seen as multiple sequential gradient learning phases] P-DARTS uses the same bi-level optimization as DARTS, which requires computing two coupled gradients—one for network weights and one for architecture parameters at every search stage. P-DARTS does not change the gradient structure. Instead, it modifies: Search depth: progressively increases the number of cells during search to reduce the depth gap between search and evaluation]
- the second gradient-based learning algorithm being a meta-learning algorithm,
which ascertains an optimized parameter set as a function of the plurality of first training tasks and the parameter set;
[1. Page 1]: “Early works on NAS focused on the optimal configuration of layer type, filter size and number, activation function, etc., to construct a complete network”,
[1, Page 1]: “follow-up works started to explore the possibility of searching for network building blocks or so called cells with reinforcement learning (RL) [35, 37] and evolutionary algorithm”,
[1. Page 1]: “ The discovered cells are then stacked orderly to construct the network for specific tasks. However, those RL-based and EA-based approaches share a common pipeline to sample and evaluate (from scratch) numerous architectures in the search space”,
[3.2, Page 3]: “we find that when searching on a deeper architecture, the differentiable approaches tend to bias towards the skip-connect operation, because it accelerates forward/backward propagation and often leads to the fastest way of gradient descent.
[2, Page 3]: ” EA-based [22] and RL based [37] NAS approaches achieved state-of-the-art performance in image recognition, where architectures were sampled and evaluated from the search space under the guidance of an EA-based or RL-based meta-controller”.
[BRI: in the broad machine learning sense, the described method is a form of meta-learning, because it learns a controller (via EA or RL) that can adaptively select and evaluate architectures from a search space, with the goal of improving optimization performance. the “meta-learning” here is in the architecture search and controller learning, while the accelerated forward/backward methods are the inner optimization engine that the meta-learner exploits.
It would be obvious to one of ordinary skills in the art before the effective filing date of the present application to combine Kabul and Chen2.
Kabul teaches digital sensor, characterization and processing and teaches two phases of training and teaches gradient based meta learning.
Chen2 teaches gradient learning with differential architecture search and teaches the meta-learning algorithm.
One of ordinary skills would be motivated to combine Kabul and Chen2 that can achieve remarkable performance and efficiency improvement (Chen2 [1, Page 1])
Regarding claim 23 (Currently Amended):
Kabul discloses:
- A device for processing digital sensor data for machine learning, classification or activation of a computer-controlled machine, comprising: a processor; and a memory for at least one artificial neural network; wherein the processor is configured to:
[0008]” FIG. 1 depicts a block diagram of a neural network training system in accordance with an illustrative embodiment”,
[0011]” FIG. 4 depicts a block diagram of a worker device of the neural network training system of FIG. 1 in accordance with an illustrative embodiment”,
[0065]” data stored in input dataset 314 may be generated as part of the Internet of Things (IoT)”,
[0065]” the IoT can include sensors in many different devices and types of devices”.
[0062]:” input dataset 314 may include data related to images, where each row includes the pixels that define a single image. The images may be of any item for which image recognition or classification may be performed including, but not limited to, faces, objects”,
[0070]; [0051]
- providing a plurality of training tasks from a distribution of training tasks, the training tasks
characterizing processing of digital sensor data; each training task of the distribution of training tasks characterizing an assignment of the digital sensor data to a result of the processing
[0064]:” Data stored in input dataset 314 may be sensor measurements or signal values captured by a sensor such as a camera, may be generated or captured in response to occurrence of an event or a transaction, generated by a device such as in response to an interaction by a user with the device, etc.
[0002] : “the computing device to provide distributed training of a neural network model. Explore phase options are distributed to each computing device of a plurality of computing devices. A subset of a training dataset is distributed to each computing device of a plurality of computing devices. A validation dataset is distributed to each computing device of the plurality of computing devices.
[0062]: “Input dataset 314 may be transposed. Input dataset 314 may include plurality of variables may define multiple dimensions for each observation vector. An observation vector x.sub.i may include a value for each of the plurality of variables associated with the observation i. Each variable of the plurality of variables may describe a characteristic of a physical object. For example, if input dataset 314 includes data related to operation of a vehicle, the variables may include an oil pressure, a speed, a gear indicator, a gas tank level, a tire pressure for each tire, an engine temperature, a radiator level, etc.
- determining a parameter set for an architecture and for weights of an artificial neural network in a first phase with a first gradient-based learning algorithm and with a second gradient-based learning algorithm as a function of a plurality of first training tasks from the distribution of training task
[Abstract]: “A computing system provides distributed training of a neural network model. Explore phase options, exploit phase options.
[0002]:” Explore phase options are distributed to each computing device of a plurality of computing devices”,
[0003]: ” Exploit phase options are distributed to each computing device of the plurality of computing devices”,
[0077]:” the goal of training is to determine a set of neural network weights that best predicts the targets in the training data”,
[0084]: “ In the exploration phase, the worker nodes of worker system 106 perform local optimization on their copy of training dataset 414. Since their portion of input dataset 314 is a random sample of the entire set of training observations, its distribution is close to the distribution of the entire set of training observations such that the local optimization is valid”,
[0037]: “ Worker nodes compute the gradient on their local observations, and then enter into a collective aggregation operation, which computes a final gradient vector. The effective number of observations used to calculate the gradients is the total number of the observations across the grid of worker nodes.
- training the artificial neural network in a second phase with the first gradient-based learning algorithm, and independently of the second gradient-based learning algorithm, as a function of the optimized parameter set as a function of one second training task from the plurality of training tasks from the distribution, the second training task being different than the plurality of first training tasks.
[Abstract]: “A computing system provides distributed training of a neural network model. Explore phase options, exploit phase options, a subset of a training dataset”,
[0028]: “ FIG. 21 illustrates an optimization problem improved by the neural network training system of FIG. 1”,
[0119]:” In an operation 638, next model configuration data to evaluate in a next iteration is computed based on the results received in operation 636 and the exploit phase options that include an optimization method and input parameters associated with the optimization method.
[0031]: “ Training a neural network is done by a data centric optimization algorithm.”,
[0197]: “ Prediction application 922, model training application 222, manager application 312, and/or model train/score application 412 may be the same or different applications that are integrated in various manners to execute a neural network model using input dataset 314 and/or second dataset 924.
[0176]:” Different parameter indices may be determined for each worker device 400 of worker system 106 for each iteration of operation 850, which is performed after operation 620 and before operation 622”,
[0146]:” In an operation 734, a mini-batch number of observation vectors are randomly selected from training dataset 414 without replacement so that on successive iterations of operation 734 different observation vectors are selected. The mini-batch number of observation vectors may be defined from a mini-batch size parameter value included in the received exploit phase options. For illustration, the mini-batch number of observation vectors may be specified as an exploit phase option and may be greater than or equal to one. An illustrative default value may be five. The mini-batch number of observation vectors specified as an exploit phase option may be different than the mini-batch number of observation vectors specified as an explore phase option”,
[0202]: “ Referring to FIG. 10, example operations of prediction application 922 are described. Additional, fewer, or different operations may be performed depending on the embodiment of prediction application 922. The order of presentation of the operations of FIG. 10 is not intended to be limiting. Although some of the operational flows are presented in sequence, the various operations may be performed in various repetitions, concurrently (in parallel, for example, using threads and/or a distributed computing system), and/or in other orders than those that are illustrated”,
[0085]:” The exploitation phase quickly progresses in the region defined by the exploration phase. A much larger mini-batch size may be used during the exploitation phase because the exploration phase identified a flat area where the solution generalizes well, reducing the variability in the gradient computation and resulting in a faster convergence”.
[0037]:”existing synchronous and asynchronous methods result in a generalization gap. Worker nodes compute the gradient on their local observations, and then enter into a collective aggregation operation, which computes a final gradient vector. The effective number of observations used to calculate the gradients is the total number of the observations across the grid of worker node”,
[0196]:” Dependent on the type of data stored in input dataset 314 and second dataset 924, prediction application 922 may identify anomalies as part of process control, for example, of a manufacturing process, for machine condition monitoring, for example, an electro-cardiogram device, for image classification, for intrusion detection, for fraud detection, etc. Some or all of the operations described herein may be embodied in prediction application 922”.
- and processing the digital sensor data as a function of the artificial neural network.
[0073]:” In an operation 508, a load of input dataset 314 may be requested. For example, user device 200 may request that input dataset 314 be loaded into a table that is ready for processing”,
[0143]:” the neural network model is re-initialized using the received weight parameter values for the links of the neural network, and processing continues in operation 710 to repeat the training process”,
- wherein the digital sensor data includes: first digital sensor data of the plurality of first training tasks, the first digital sensor data being acquired by a first sensor, and second digital sensor data of the second training task, the second digital sensor data being acquired by a second sensor,
[0063]:” The data stored in input dataset 314 may be generated by and/or captured from a variety of sources including one or more sensors of the same or different type”,
[0064]:” Data stored in input dataset 314 may be sensor measurements or signal values captured by a sensor such as a camera, may be generated or captured in response to occurrence of an event or a transaction, generated by a device such as in response to an interaction by a user with the device, etc. The data stored in input dataset 314 may include any type of content represented in any computer-readable format
[0064]” The content may include textual information, graphical information, image information, audio information, numeric information, etc. that further may be encoded using various encoding techniques as understood by a person of skill in the art. The data stored in input dataset 314 may be captured at different time points periodically, intermittently, when an event occurs, etc
[0065]: “input dataset 314 may include data captured at a high data rate such as 200 or more observations per second for one or more physical objects. For example, data stored in input dataset 314 may be generated as part of the Internet of Things (IoT), where things (e.g., machines, devices, phones, sensors) can be connected to networks and the data from these things collected and processed within the things and/or external to the things before being stored in input dataset 314. For example, the IoT can include sensors in many different devices and types of devices”,
[0062]:” Input dataset 314 may include labeled and/or unlabeled data. The plurality of variables may define multiple dimensions for each observation vector. An observation vector
x
i
may include a value for each of the plurality of variables associated with the observation i. Each variable of the plurality of variables may describe a characteristic of a physical object. For example, if input dataset 314 includes data related to operation of a vehicle, the variables may include an oil pressure, a speed, a gear indicator, a gas tank level, a tire pressure for each tire, an engine temperature, a radiator level, etc”,
[0062]: “ Input dataset 314 may include data captured as a function of time for one or more physical objects. As another example, input dataset 314 may include data related to images, where each row includes the pixels that define a single image. The images may be of any item for which image recognition or classification may be performed including, but not limited to, faces, objects”.
[0064]: “The data stored in input dataset 314 may be captured at different time points periodically, intermittently, when an event occurs, etc.”.
- and wherein one of the second sensor corresponds to the first sensor with a software update, or the second sensor represents a hardware or software generation change compared to the first sensor.
[0063]: “The data stored in input dataset 314 may be generated by and/or captured from a variety of sources including one or more sensors of the same or different type, one or more computing devices, etc. The data stored in input dataset 314 may be received directly or indirectly from the source and may or may not be pre-processed in some manner.
[0063]: “The data may be organized using delimited fields, such as comma or space separated fields, fixed width fields, using a SAS® dataset, etc. The SAS dataset may be a SAS® file stored in a SAS® library that a SAS® software tool creates and processes. The SAS dataset contains data values that are organized as a table of observations (rows) and variables (columns) that can be processed by one or more SAS software tools”,
[0194]:” Referring to FIG. 9, a block diagram of a prediction device 900 is shown in accordance with an illustrative embodiment. Prediction device 900 may include a fourth input interface 902, a fourth output interface 904, a fourth communication interface 906, a fourth non-transitory computer-readable medium 908, a fourth processor 910, a prediction application 922”,
PNG
media_image1.png
562
821
media_image1.png
Greyscale
[0032]:” To make neural network predictions more accurate, the values of its parameters are updated iteratively. Before doing this, how much each value changes is computed. An efficient way of computing these values is done using a back-propagation method that computes a rate of change by going through each computation unit or node in the neural network, by calculating partial derivatives of the network error with respect to each parameter linking pairs of nodes, and by applying the chain rule to do the same thing to the next computation unit. With one backward pass, the gradient (rate of change) values for each of the parameters of the neural network is computed. These gradient values are input to the optimization method to make the actual parameter update. There are a large variety of different optimization methods, almost all of which depend on knowledge of the values of the gradient for each parameter in the neural network”,
[0065]: “ data stored in input dataset 314 may be generated as part of the Internet of Things (IoT), where things (e.g., machines, devices, phones, sensors) can be connected to networks and the data from these things collected and processed within the things and/or external to the things before being stored in input dataset 314. For example, the IoT can include sensors in many different devices and types of devices, and high value analytics can be applied to identify hidden relationships and drive increased efficiencies. This can apply to both big data analytics and real-time analytics. Some of these devices may be referred to as edge devices, and may involve edge computing circuitry. These devices may provide a variety of stored or generated data, such as network data or data specific to the network devices themselves”.
Kabul does not explicitly disclose:
- the first gradient-based learning algorithm being a differentiable architecture search algorithm
- the second gradient-based learning algorithm being a meta-learning algorithm,
which ascertains an optimized parameter set as a function of the plurality of first training tasks and the parameter set.
However, Chen2 discloses:
- the first gradient-based learning algorithm being a differentiable architecture search algorithm
[1, Page 1]:
PNG
media_image2.png
8
424
media_image2.png
Greyscale
Figure 1: Difference between DARTS and P-DARTS (our approach), with the former searching architectures in a shallow setting and evaluating them in a deep one, and the latter progressively increasing the searching depth, so as to bridge the depth gap between search and evaluation. Green and blue indicate search and evaluation, respectively”.
[3.2, Page 3]: “ we find that when searching on a deeper architecture, the differentiable approaches tend to bias towards the skip-connect operation, because it accelerates forward/backward propagation and often leads to the fastest way of gradient descent”.
- the second gradient-based learning algorithm being a meta-learning algorithm,
which ascertains an optimized parameter set as a function of the plurality of first training tasks and the parameter set;
[1. Page 1]: “Early works on NAS focused on the optimal configuration of layer type, filter size and number, activation function, etc., to construct a complete network”,
[1, Page 1]: “follow-up works started to explore the possibility of searching for network building blocks or so called cells with reinforcement learning (RL) [35, 37] and evolutionary algorithm”,
[1. Page 1]: “The discovered cells are then stacked orderly to construct the network for specific tasks. However, those RL-based and EA-based approaches share a common pipeline to sample and evaluate (from scratch) numerous architectures in the search space”,
[3.2, Page 3]: “we find that when searching on a deeper architecture, the differentiable approaches tend to bias towards the skip-connect operation, because it accelerates forward/backward propagation and often leads to the fastest way of gradient descent.
[2, Page 3]: ” EA-based [22] and RL based [37] NAS approaches achieved state-of-the-art performance in image recognition, where architectures were sampled and evaluated from the search space under the guidance of an EA-based or RL-based meta-controller”.
It would be obvious to one of ordinary skills in the art before the effective filing date of the present application to combine Kabul and Chen2.
Kabul teaches digital sensor, characterization and processing and teaches two phases of training and teaches gradient based meta learning.
Chen2 teaches gradient learning with differential architecture search and teaches the meta-learning algorithm.
One of ordinary skills would be motivated to combine Kabul and Chen2 that can achieve remarkable performance and efficiency improvement (Chen2 [1, Page 1])
Regarding claim 24 (Currently Amended):
Kabul discloses:
- A non-transitory machine-readable memory medium on which is stored a computer program for processing digital sensor data, the computer program, when executed by a computer, causing the computer to perform the following steps:
[0194]: “Referring to FIG. 9, a block diagram of a prediction device 900 is shown in accordance with an illustrative embodiment. Prediction device 900 may include a fourth input interface 902, a fourth output interface 904, a fourth communication interface 906, a fourth non-transitory computer-readable medium 908”,
[0197]:” Referring to the example embodiment of FIG. 9, prediction application 922 is implemented in software (comprised of computer-readable and/or computer-executable instructions) stored in fourth computer-readable medium 908 and accessible by fourth processor 910 for execution of the instructions that embody the operations of prediction application 922. Prediction application 922 may be written using one or more programming languages”,
[0050]: “ For example, data stored in input dataset 314 may be generated as part of the Internet of Things (IoT), where things (e.g., machines, devices, phones, sensors) can be connected to networks and the data from these things collected and processed within the things and/or external to the things before being stored in input dataset 314.
- providing a plurality of training tasks from a distribution of training tasks, the training tasks
characterizing processing of digital sensor data; each training task of the distribution of training tasks characterizing an assignment of the digital sensor data to a result of the processing
[0064]:” Data stored in input dataset 314 may be sensor measurements or signal values captured by a sensor such as a camera, may be generated or captured in response to occurrence of an event or a transaction, generated by a device such as in response to an interaction by a user with the device, etc.
[0002] : “the computing device to provide distributed training of a neural network model. Explore phase options are distributed to each computing device of a plurality of computing devices. A subset of a training dataset is distributed to each computing device of a plurality of computing devices. A validation dataset is distributed to each computing device of the plurality of computing devices.
[0062]: “Input dataset 314 may be transposed. Input dataset 314 may include plurality of variables may define multiple dimensions for each observation vector. An observation vector x.sub.i may include a value for each of the plurality of variables associated with the observation i. Each variable of the plurality of variables may describe a characteristic of a physical object. For example, if input dataset 314 includes data related to operation of a vehicle, the variables may include an oil pressure, a speed, a gear indicator, a gas tank level, a tire pressure for each tire, an engine temperature, a radiator level, etc.
- determining a parameter set for an architecture and for weights of an artificial neural network in a first phase with a first gradient-based learning algorithm and with a second gradient-based learning algorithm as a function of a plurality of first training tasks from the distribution of training task
[Abstract]: “A computing system provides distributed training of a neural network model. Explore phase options, exploit phase options.
[0002]:” Explore phase options are distributed to each computing device of a plurality of computing devices”,
[0003]: ” Exploit phase options are distributed to each computing device of the plurality of computing devices”,
[0077]:” the goal of training is to determine a set of neural network weights that best predicts the targets in the training data”,
[0084]: “ In the exploration phase, the worker nodes of worker system 106 perform local optimization on their copy of training dataset 414. Since their portion of input dataset 314 is a random sample of the entire set of training observations, its distribution is close to the distribution of the entire set of training observations such that the local optimization is valid”,
[0037]: “ Worker nodes compute the gradient on their local observations, and then enter into a collective aggregation operation, which computes a final gradient vector. The effective number of observations used to calculate the gradients is the total number of the observations across the grid of worker nodes.
- training the artificial neural network in a second phase with the first gradient-based learning algorithm, and independently of the second gradient-based learning algorithm, as a function of the optimized parameter set as a function of one second training task from the plurality of training tasks from the distribution, the second training task being different than the plurality of first training tasks.
[Abstract]: “A computing system provides distributed training of a neural network model. Explore phase options, exploit phase options, a subset of a training dataset”,
[0028]: “ FIG. 21 illustrates an optimization problem improved by the neural network training system of FIG. 1”,
[0119]:” In an operation 638, next model configuration data to evaluate in a next iteration is computed based on the results received in operation 636 and the exploit phase options that include an optimization method and input parameters associated with the optimization method.
[0031]: “ Training a neural network is done by a data centric optimization algorithm.”,
[0197]: “ Prediction application 922, model training application 222, manager application 312, and/or model train/score application 412 may be the same or different applications that are integrated in various manners to execute a neural network model using input dataset 314 and/or second dataset 924.
[0176]:” Different parameter indices may be determined for each worker device 400 of worker system 106 for each iteration of operation 850, which is performed after operation 620 and before operation 622”,
[0146]” In an operation 734, a mini-batch number of observation vectors are randomly selected from training dataset 414 without replacement so that on successive iterations of operation 734 different observation vectors are selected. The mini-batch number of observation vectors may be defined from a mini-batch size parameter value included in the received exploit phase options. For illustration, the mini-batch number of observation vectors may be specified as an exploit phase option and may be greater than or equal to one. An illustrative default value may be five. The mini-batch number of observation vectors specified as an exploit phase option may be different than the mini-batch number of observation vectors specified as an explore phase option”,
[0202: “ Referring to FIG. 10, example operations of prediction application 922 are described. Additional, fewer, or different operations may be performed depending on the embodiment of prediction application 922. The order of presentation of the operations of FIG. 10 is not intended to be limiting. Although some of the operational flows are presented in sequence, the various operations may be performed in various repetitions, concurrently (in parallel, for example, using threads and/or a distributed computing system), and/or in other orders than those that are illustrated”,
[0085]:” The exploitation phase quickly progresses in the region defined by the exploration phase. A much larger mini-batch size may be used during the exploitation phase because the exploration phase identified a flat area where the solution generalizes well, reducing the variability in the gradient computation and resulting in a faster convergence”.
[0037]:”existing synchronous and asynchronous methods result in a generalization gap. Worker nodes compute the gradient on their local observations, and then enter into a collective aggregation operation, which computes a final gradient vector. The effective number of observations used to calculate the gradients is the total number of the observations across the grid of worker node”,
[0196]:” Dependent on the type of data stored in input dataset 314 and second dataset 924, prediction application 922 may identify anomalies as part of process control, for example, of a manufacturing process, for machine condition monitoring, for example, an electro-cardiogram device, for image classification, for intrusion detection, for fraud detection, etc. Some or all of the operations described herein may be embodied in prediction application 922”.
- and processing the digital sensor data as a function of the artificial neural network.
[0073]:” In an operation 508, a load of input dataset 314 may be requested. For example, user device 200 may request that input dataset 314 be loaded into a table that is ready for processing”,
[0143]:” the neural network model is re-initialized using the received weight parameter values for the links of the neural network, and processing continues in operation 710 to repeat the training process”,
- wherein the digital sensor data includes: first digital sensor data of the plurality of first training tasks, the first digital sensor data being acquired by a first sensor, and second digital sensor data of the second training task, the second digital sensor data being acquired by a second sensor,
[0063]:” The data stored in input dataset 314 may be generated by and/or captured from a variety of sources including one or more sensors of the same or different type”,
[0064]:” Data stored in input dataset 314 may be sensor measurements or signal values captured by a sensor such as a camera, may be generated or captured in response to occurrence of an event or a transaction, generated by a device such as in response to an interaction by a user with the device, etc. The data stored in input dataset 314 may include any type of content represented in any computer-readable format
[0064]” The content may include textual information, graphical information, image information, audio information, numeric information, etc. that further may be encoded using various encoding techniques as understood by a person of skill in the art. The data stored in input dataset 314 may be captured at different time points periodically, intermittently, when an event occurs, etc
[0065]: “input dataset 314 may include data captured at a high data rate such as 200 or more observations per second for one or more physical objects. For example, data stored in input dataset 314 may be generated as part of the Internet of Things (IoT), where things (e.g., machines, devices, phones, sensors) can be connected to networks and the data from these things collected and processed within the things and/or external to the things before being stored in input dataset 314. For example, the IoT can include sensors in many different devices and types of devices”,
[0062]:” Input dataset 314 may include labeled and/or unlabeled data. The plurality of variables may define multiple dimensions for each observation vector. An observation vector
x
i
may include a value for each of the plurality of variables associated with the observation i. Each variable of the plurality of variables may describe a characteristic of a physical object. For example, if input dataset 314 includes data related to operation of a vehicle, the variables may include an oil pressure, a speed, a gear indicator, a gas tank level, a tire pressure for each tire, an engine temperature, a radiator level, etc”,
[0062]: “ Input dataset 314 may include data captured as a function of time for one or more physical objects. As another example, input dataset 314 may include data related to images, where each row includes the pixels that define a single image. The images may be of any item for which image recognition or classification may be performed including, but not limited to, faces, objects”.
[0064]: “The data stored in input dataset 314 may be captured at different time points periodically, intermittently, when an event occurs, etc.”.
- and wherein one of the second sensor corresponds to the first sensor with a software update, or the second sensor represents a hardware or software generation change compared to the first sensor.
[0063]: “The data stored in input dataset 314 may be generated by and/or captured from a variety of sources including one or more sensors of the same or different type, one or more computing devices, etc. The data stored in input dataset 314 may be received directly or indirectly from the source and may or may not be pre-processed in some manner.
[0063]: “The data may be organized using delimited fields, such as comma or space separated fields, fixed width fields, using a SAS® dataset, etc. The SAS dataset may be a SAS® file stored in a SAS® library that a SAS® software tool creates and processes. The SAS dataset contains data values that are organized as a table of observations (rows) and variables (columns) that can be processed by one or more SAS software tools”,
[0194]:” Referring to FIG. 9, a block diagram of a prediction device 900 is shown in accordance with an illustrative embodiment. Prediction device 900 may include a fourth input interface 902, a fourth output interface 904, a fourth communication interface 906, a fourth non-transitory computer-readable medium 908, a fourth processor 910, a prediction application 922”,
PNG
media_image1.png
562
821
media_image1.png
Greyscale
[0032]:” To make neural network predictions more accurate, the values of its parameters are updated iteratively. Before doing this, how much each value changes is computed. An efficient way of computing these values is done using a back-propagation method that computes a rate of change by going through each computation unit or node in the neural network, by calculating partial derivatives of the network error with respect to each parameter linking pairs of nodes, and by applying the chain rule to do the same thing to the next computation unit. With one backward pass, the gradient (rate of change) values for each of the parameters of the neural network is computed. These gradient values are input to the optimization method to make the actual parameter update. There are a large variety of different optimization methods, almost all of which depend on knowledge of the values of the gradient for each parameter in the neural network”,
[0065]: “ data stored in input dataset 314 may be generated as part of the Internet of Things (IoT), where things (e.g., machines, devices, phones, sensors) can be connected to networks and the data from these things collected and processed within the things and/or external to the things before being stored in input dataset 314. For example, the IoT can include sensors in many different devices and types of devices, and high value analytics can be applied to identify hidden relationships and drive increased efficiencies. This can apply to both big data analytics and real-time analytics. Some of these devices may be referred to as edge devices, and may involve edge computing circuitry. These devices may provide a variety of stored or generated data, such as network data or data specific to the network devices themselves”.
Kabul does not explicitly disclose:
- the first gradient-based learning algorithm being a differentiable architecture search algorithm
- the second gradient-based learning algorithm being a meta-learning algorithm,
which ascertains an optimized parameter set as a function of the plurality of first training tasks and the parameter set.
However, Chen2 discloses:
- the first gradient-based learning algorithm being a differentiable architecture search algorithm
[1, Page 1]:
PNG
media_image2.png
8
424
media_image2.png
Greyscale
Figure 1: Difference between DARTS and P-DARTS (our approach), with the former searching architectures in a shallow setting and evaluating them in a deep one, and the latter progressively increasing the searching depth, so as to bridge the depth gap between search and evaluation. Green and blue indicate search and evaluation, respectively”.
[3.2, Page 3]: “ we find that when searching on a deeper architecture, the differentiable approaches tend to bias towards the skip-connect operation, because it accelerates forward/backward propagation and often leads to the fastest way of gradient descent”.
- the second gradient-based learning algorithm being a meta-learning algorithm,
which ascertains an optimized parameter set as a function of the plurality of first training tasks and the parameter set;
[1. Page 1]: “Early works on NAS focused on the optimal configuration of layer type, filter size and number, activation function, etc., to construct a complete network”,
[1, Page 1]: “follow-up works started to explore the possibility of searching for network building blocks or so called cells with reinforcement learning (RL) [35, 37] and evolutionary algorithm”,
[1. Page 1]: “ The discovered cells are then stacked orderly to construct the network for specific tasks. However, those RL-based and EA-based approaches share a common pipeline to sample and evaluate (from scratch) numerous architectures in the search space”,
[3.2, Page 3]: “we find that when searching on a deeper architecture, the differentiable approaches tend to bias towards the skip-connect operation, because it accelerates forward/backward propagation and often leads to the fastest way of gradient descent.
[2, Page 3]: ” EA-based [22] and RL based [37] NAS approaches achieved state-of-the-art performance in image recognition, where architectures were sampled and evaluated from the search space under the guidance of an EA-based or RL-based meta-controller”.
It would be obvious to one of ordinary skills in the art before the effective filing date of the present application to combine Kabul and Chen2.
Kabul teaches digital sensor, characterization and processing and teaches two phases of training and teaches gradient based meta learning.
Chen2 teaches gradient learning with differential architecture search and teaches the meta-learning algorithm.
One of ordinary skills would be motivated to combine Kabul and Chen2 that can achieve remarkable performance and efficiency improvement (Chen2 [1, Page 1])
claims 19-20 are rejected under 35 U.S.C. 103 as being unpatentable over
Mustafa Kabul et.al. (hereinafter Kabul) US 2018/0307986 A1,
in view of Xin Chen et.al (hereinafter Chen2) Progressive Differentiable Architecture Search: Bridging the Depth Gap between Search and Evaluation, 2019 IEEE/CVF International Conference [Published in arXiv:1904.12760v1 [cs.CV] 29 Apr 2019],
further in view of Antonino Mondello et.al (hereinafter Mondello) US 2019/0205744 A1.
Regarding claim 19 (Previously Presented):
Kabul does not explicitly disclose:
- wherein the computer-controlled machine is an at least semi-autonomous robot, or a vehicle, or a home application, or a power tool, or a personal assistance system, or an access control system.
However, Mondello discloses:
- wherein the computer-controlled machine is an at least semi-autonomous robot, or a vehicle, or a home application, or a power tool, or a personal assistance system, or an access control system.
[0035]:” A function of the vehicles (111, . . . , 113) for autonomous driving and/or advanced driver assistance may process such an unknown item according to a pre-programmed policy.
It would be obvious to one of ordinary skills in the art before the effective filing date of the present application to combine Kabul, Chen2 and Mondello.
Kabul teaches digital sensor, characterization and processing and teaches two phases of training and teaches gradient based meta learning.
Chen2 teaches gradient learning with differential architecture search and teaches the meta-learning algorithm.
Mondello teaches machine-controlled output actions (vehicle).
One of ordinary skills would be motivated to combine Kabul, Chen2 and Mondello that can provide an improved identification or classification of an object (Mondello [0055]).
Regarding claim 20 (Previously Presented):
Kabul and Chen2 do not explicitly disclose:
- wherein the training data include image data, video data and/or digital sensor data of a sensor, from at least one camera and/or one infrared camera and/or one LIDAR sensor and/or one radar sensor and/or one acoustic sensor and/or one ultrasonic sensor and/or one receiver for a satellite navigation system and/or one rotational speed sensor and/or one torque sensor and/or one acceleration sensor and/or one position sensor.
However, Mondello discloses:
- wherein the training data include image data, video data and/or digital sensor data of a sensor, from at least one camera and/or one infrared camera and/or one LIDAR sensor and/or one radar sensor and/or one acoustic sensor and/or one ultrasonic sensor and/or one receiver for a satellite navigation system and/or one rotational speed sensor and/or one torque sensor and/or one acceleration sensor and/or one position sensor.
[0052]: “ The one or more sensors (137) may include a visible light camera, an infrared camera, a LIDAR, RADAR, or sonar system, and/or peripheral sensors, which are configured to provide sensor input to the computer (131). A module of the firmware (or software) (127) executed in the processor(s) (133) applies the sensor input to an ANN defined by the model (119) to generate an output that identifies or classifies an event or object captured in the sensor input, such as an image or video clip.
It would be obvious to one of ordinary skills in the art before the effective filing date of the present application to combine Kabul, Chen2 and Mondello.
Kabul teaches digital sensor, characterization and processing and teaches two phases of training and teaches gradient based meta learning.
Chen2 teaches gradient learning with differential architecture search and teaches the meta-learning algorithm.
Mondello teaches types of training data.
One of ordinary skill would be motivated to combine Kabul and Mondello that can provide an improved identification or classification of an object (Mondello [0055]).
Claim 22 is rejected under 35 U.S.C. 103 as being unpatentable over
Mustafa Kabul et.al. (hereinafter Kabul) US 2018/0307986 A1,
in view of Xin Chen et.al (hereinafter Chen2) Progressive Differentiable Architecture Search: Bridging the Depth Gap between Search and Evaluation, 2019 IEEE/CVF International Conference [Published in arXiv:1904.12760v1 [cs.CV] 29 Apr 2019]
further in view of Zhao Chen et.al (hereinafter Chen3) US 2019/01230275 A1.
Regarding claim 22 (Previously Presented)
Kabul and Chen2 do not explicitly disclose:
- wherein the artificial neural network is trained with the first gradient-based learning algorithm as a function of the parameter set and as a function of a second training task.
However, Chen3 discloses:
- wherein the artificial neural network is trained with the first gradient-based learning algorithm as a function of the parameter set and as a function of a second training task.
[0099]: “ The local processing and data module 924 may comprise a hardware processor, as well as non-transitory digital memory, such as non-volatile memory e.g., flash memory, both of which may be utilized to assist in the processing, caching, and storage of data. The data include data (a) captured from sensors (which may be, e.g., operatively coupled to the frame 912 or otherwise attached to the wearer 904), such as image capture devices (such as cameras)”,
[0146]: “receive a sensor datum captured by the sensor; determine a task output for each task of the plurality of tasks using the multitask network with the sensor datum as input
[0011]:” FIG. 1B is an example schematic illustration of balanced gradient norms across task when training a multitask network”,
[0029]: “Clustering methods have shown success beyond deep models, while constructs such as deep relationship networks and cross-stich networks give deep networks the capacity to search for meaningful relationships between tasks and to learn which features to share between them. Groupings amongst labels can be used to search through possible architectures for learning”.
[BRI: Within the context of a multi-task network this represents a form of Differential Neural Architecture Search (NAS), where the “grouping” of DNNs by meaningful task relationships is part of defining the search space and strategy for finding optimal architectures and may be extension of NAS beyond single-task design]
[0025]:” Multitask networks can difficult to train: different tasks need to be properly balanced so network parameters converge to robust shared features that are useful across all tasks. In some methods, methods in multitask learning can find this balance by manipulating the forward pass of the network (e.g., through constructing explicit statistical relationships between features or optimizing multitask network architectures.
[0023]:” gradient normalization (GradNorm) methods that automatically balance training in deep multitask models by dynamically tuning gradient magnitudes.
[0023]:” GradNorm can improve accuracy and/or reduce overfitting across multiple tasks”,
[0026]:” In some embodiments, the multitask loss function is a weighted linear combination of the single task losses
L
i
,
∑
i
w
i
L
i
where the sum runs over all T tasks. An adaptive method is disclosed herein to vary
w
i
at one or more training steps or iterations (e.g., each training step t:
w
i
=
w
i
(t)). This linear form of the loss function can be convenient for implementing gradient balancing, as
w
i
directly and linearly couples to the backpropagated gradient magnitudes from each task. The gradient normalization methods disclosed herein can find a good value (e.g., the best value) for each
w
i
at each training step t that balances the contribution of each task for improved (e.g., optimal) model training”.
[0174]:” determining the weights for each of the single task loss functions comprises penalizing the multitask neural network when backpropagated gradients from a first task of the plurality of tasks are substantially different from backpropagated gradients from a second task of the plurality of tasks.
[0175]: “determining the weights for each of the single task loss functions comprises decreasing a first weight for a first task of the plurality of tasks relative to a second weight for a second task of the plurality of tasks when a first training rate for the first task exceeds a second training rate for the second task”.
[0101]:” the parameters of the CNN (e.g., weights, bias terms, subsampling factors for pooling layers, number and size of kernels in different layers, number of feature maps, etc.) can be stored in data modules 924 and/or 932”,
[0053]:” All model parameters were shared amongst all tasks until the final layer”.
[BRI: using different weights for different tasks in MTL does represents trained with a gradient-based algorithm as a function of the parameter set and an auxiliary training task, provided the weights are themselves learned (or updated) during training and is a known and active approach in MTL research for mitigating negative interference and balancing task contributions. In this context, the decrease of weight for each of the tasks is weights are themselves learned]
It would be obvious to one of ordinary skills in the art before the effective filing date of the present application to combine Kabul, Chen2 and Chen3.
Kabul teaches digital sensor, characterization and processing and teaches two phases of training and teaches gradient based meta learning.
Chen2 teaches gradient learning with differential architecture search and teaches the meta-learning algorithm.
Chen3 teaches the first gradient-based learning algorithm as a function of the parameter set and as a function of a second training task.
One of ordinary skills would be to combine Kabul, Chen2 and Chen3 that can provide improved task balance results (Chen3[0046]).
Conclusion
Any inquiry concerning this communication or earlier communications from the
examiner should be directed to TIRUMALE KRISHNASWAMY RAMESH whose telephone number is (571)272-4605. The examiner can normally be reached by phone.
Examiner interviews are available via telephone, in-person, and video conferencing
using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at
http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Li B Zhen can be reached on phone (571-272-3768). The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be
obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit:
https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for
information about filing in DOCX format. For additional questions, contact the Electronic
Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO
Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/TIRUMALE K RAMESH/Examiner, Art Unit 2121
/Li B. Zhen/Supervisory Patent Examiner, Art Unit 2121