DETAILED ACTION
This action is in response to the Applicant Response filed 07 February 2024 for application 18/435,374 filed 07 February 2024.
Claim(s) 1-20 is/are pending.
Claim(s) 1-20 is/are rejected.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claim 20 is rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor, or for pre-AIA the applicant regards as the invention.
Claim 20 recites the limitation to at least one other computing device. However, the claim fails to provide proper antecedent basis of an initial computing device from which to refer “one other” computing device. Further, claim 20 recites from each computing device of the plurality of computing devices while failing to provide a proper antecedent basis. It is suggested that the term the plurality of computing devices be amended to recite “a plurality of computing devices.” Correction or clarification is required.
Examiner’s Note: For the purposes of examination, Examiner will interpret the phrase to at least one other computing device to recite “to at least one computing device.”
Double Patenting
The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969).
A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on nonstatutory double patenting provided the reference application or patent either is shown to be commonly owned with the examined application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. See MPEP § 717.02 for applications subject to examination under the first inventor to file provisions of the AIA as explained in MPEP § 2159. See MPEP § 2146 et seq. for applications not subject to examination under the first inventor to file provisions of the AIA . A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b).
The filing of a terminal disclaimer by itself is not a complete reply to a nonstatutory double patenting (NSDP) rejection. A complete reply requires that the terminal disclaimer be accompanied by a reply requesting reconsideration of the prior Office action. Even where the NSDP rejection is provisional the reply must be complete. See MPEP § 804, subsection I.B.1. For a reply to a non-final Office action, see 37 CFR 1.111(a). For a reply to final Office action, see 37 CFR 1.113(c). A request for reconsideration while not provided for in 37 CFR 1.113(c) may be filed after final for consideration. See MPEP §§ 706.07(e) and 714.13.
The USPTO Internet website contains terminal disclaimer forms which may be used. Please visit www.uspto.gov/patent/patents-forms. The actual filing date of the application in which the form is filed determines what form (e.g., PTO/SB/25, PTO/SB/26, PTO/AIA /25, or PTO/AIA /26) should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto-processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to www.uspto.gov/patents/apply/applying-online/eterminal-disclaimer.
Claims 1-20 are rejected on the ground of nonstatutory double patenting as being unpatentable over claims 1-9, 12-18, 20 of U.S. Patent No. 11,928,571 (16/950,129) in view of Tomioka et al. (US 2018/0336458 A1 – Asynchronous Neural Network, hereinafter referred to as “Tomioka”). Although the claims at issue are not identical, they are not patentably distinct from each other, because, as noted in the table below, claims 1-20 of the instant application have similar limitations as recited in U.S. Patent No. 11,928,571 (claims 1-9, 12-18, 20).
Application 18/435,374
U.S. Pat. No. 11,928,571 (16/950,129)
Claim 1
Claim 1 (Claim 7, claim 8)
A computer-implemented method, comprising:
A computer-implemented method for training a distributed machine learning model, comprising:
initializing a distributed machine learning model on a plurality of computing devices, the distributed machine learning model comprising a plurality of computational nodes, each computing device of the plurality of computing devices comprising a respective subset of the plurality of computational nodes, each computational node comprising at least one parameter;
initializing a distributed machine learning model on a plurality of computing devices, the distributed machine learning model comprising a plurality of computational nodes, each computing device of the plurality of computing devices comprising a respective subset of the plurality of computational nodes, each computational node comprising at least one parameter;
receiving training data associated with a plurality of samples at a first computing device of the plurality of computing devices;
receiving training data associated with a plurality of samples at a first computing device of the plurality of computing devices;
forward propagating each sample of the plurality of samples through the distributed machine learning model to generate an output for each sample of the plurality of samples, wherein forward propagating comprises communicating, from each computing device of the plurality of computing devices other than a last computing device, an intermediate output of the respective subset of the plurality of computational nodes to a next computing device of the plurality of computing devices via a first message queue;
Claim 1
forward propagating each sample of the plurality of samples through the distributed machine learning model to generate an output for each sample of the plurality of samples;
Claim 7
wherein forward propagating comprises communicating, from each computing device of the plurality of computing devices other than a last computing device, an intermediate output of the respective subset of the plurality of computational nodes to a next computing device of the plurality of computing devices via a first message queue.
determining a loss for each sample of the plurality of samples based on the output;
determining a loss for each sample of the plurality of samples based on the output;
backward propagating the loss for each sample of the plurality of samples to each computing device of the plurality of computing devices, wherein backward propagating comprises communicating, from each computing device of the plurality of computing devices other than the first computing device, a gradient value associated with the loss associated with the respective subset of the plurality of computational nodes to a previous computing device of the plurality of computing devices via a second message queue;
Claim 1
backward propagating the loss for each sample of the plurality of samples to each computing device of the plurality of computing devices;
Claim 8
wherein backward propagating comprises communicating, from each computing device of the plurality of computing devices other than the first computing device, a gradient value associated with the loss associated with the respective subset of the plurality of computational nodes to a previous computing device of the plurality of computing devices via a second message queue.
asynchronously updating the at least one parameter of each computational node based on the loss for each sample as the loss for each sample is backward propagated while at least one of the plurality of samples is forward propagating through the distributed machine learning model;
asynchronously updating the at least one parameter of each computational node based on the loss for each sample as the loss for each sample is backward propagated while at least one of the plurality of samples is forward propagating through the distributed machine learning model;
storing, at each computing device of the plurality of computing devices, the at least one parameter of each computational node of the respective subset of the plurality of computational nodes as updated; and
storing, at each computing device of the plurality of computing devices, the at least one parameter of each computational node of the respective subset of the plurality of computational nodes as updated;
communicating, from each computing device of the plurality of computing devices to all other computing devices of the plurality of computing devices, data associated with the at least one parameter of each computational node of the respective subset of the plurality of computational nodes as updated, each of the other computing devices of the plurality of computing devices storing the at least one parameter of each computational node as updated.
communicating, from each computing device of the plurality of computing devices to all other computing devices of the plurality of computing devices, data associated with the at least one parameter of each computational node of the respective subset of the plurality of computational nodes as updated, each of the other computing devices of the plurality of computing devices storing the at least one parameter of each computational node as updated;
Claim 2
Claim 2
wherein the distributed machine learning model comprises a deep neural network and each computational node of the plurality of computational nodes comprises a neuron.
wherein the distributed machine learning model comprises a deep neural network and each computational node of the plurality of computational nodes comprises a neuron.
Claim 3
Claim 3
wherein the plurality of computational nodes are divided into a plurality of layers, each subset of computational nodes being associated with at least one layer of the plurality of layers, each layer being implemented on at least one of the plurality of computing devices.
wherein the plurality of computational nodes are divided into a plurality of layers, each subset of computational nodes being associated with at least one layer of the plurality of layers, each layer being implemented on at least one of the plurality of computing devices.
Claim 4
Claim 4
wherein the at least one parameter of each computational node comprises at least one of a weight parameter, a bias parameter, or any combination thereof.
wherein the at least one parameter of each computational node comprises at least one of a weight parameter, a bias parameter, or any combination thereof.
Claim 5
Claim 5
wherein storing the at least one parameter of each computational node comprises storing the at least one parameter of each computational node in at least one cache memory of each computing device of the plurality of computing devices.
wherein storing the at least one parameter of each computational node comprises storing the at least one parameter of each computational node in at least one cache memory of each computing device of the plurality of computing devices.
Claim 6
Claim 6
storing, in a backup storage, the at least one parameter of each computational node of the plurality of computational nodes.
storing, in a backup storage, the at least one parameter of each computational node of the plurality of computational nodes.
Claim 7
wherein the first message queue comprises a first publisher-subscriber model queue, and wherein the second message queue comprises a second publisher-subscriber model queue.
Claim 8
wherein communicating the intermediate output via the first message queue comprises publishing, by each computing device of the plurality of computing devices other than the last computing device, the intermediate output of the respective subset of the plurality of computational nodes to the first publisher-subscriber model queue, wherein the next computing device of the plurality of computing devices comprises a subscriber to the first publisher-subscriber model queue,
wherein communicating the gradient value via the second message queue comprises publishing, by each computing device of the plurality of computing devices other than the first computing device, the gradient value associated with the loss associated with the respective subset of the plurality of computational nodes to the second publisher-subscriber model queue, wherein the previous computing device of the plurality of computing devices comprises a subscriber to the second publisher-subscriber model queue.
Claim 9
Claim 9
storing, in a database, a label for each sample of the plurality of samples, wherein determining the loss comprises determining the loss for each sample of the plurality of samples based on the output and the label.
storing, in a database, a label for each sample of the plurality of samples, wherein determining the loss comprises determining the loss for each sample of the plurality of samples based on the output and the label.
Claim 10
Claim 1
determining a variance in the at least one parameter for each computational node;
in response to determining the loss for the first sample satisfies the threshold, determining a variance in the at least one parameter for each computational node; and
generating at least one new computational node on at least one computing device of the plurality of computing devices based on the variance of at least one computational node of the respective subset of the plurality of computational nodes.
generating at least one new computational node on at least one computing device of the plurality of computing devices based on the variance of at least one computational node of the respective subset of the plurality of computational nodes.
Claim 11
Claim 1
determining the loss for a first sample of the plurality of samples satisfies a threshold associated with at least one computing device of the plurality of computing devices becoming unavailable before determining the variance in the at least one parameter for each computational node.
determining the loss for a first sample of the plurality of samples satisfies a threshold associated with at least one computing device of the plurality of computing devices becoming unavailable;
Claim 12
Claim 12 (Claim 16, Claim 17)
A system, comprising:
A system for training a distributed machine learning model, comprising:
a first message queue;
Claim 16
a first message queue, wherein forward propagating comprises communicating, from each computing device of the plurality of computing devices other than a last computing device, an intermediate output of the respective subset of the plurality of computational nodes to a next computing device of the plurality of computing devices via the first message queue.
a second message queue;
Claim 17
a second message queue, wherein backward propagating comprises communicating, from each computing device of the plurality of computing devices other than a first computing device, a gradient value associated with the loss associated with the respective subset of the plurality of computational nodes to a previous computing device of the plurality of computing devices via the second message queue.
a plurality of computing devices, each computing device of the plurality of computing devices comprising a respective subset of a plurality of computational nodes of a distributed machine learning model, each computational node comprising at least one parameter, each computing device comprising at least one processor and at least one non-transitory computer-readable medium including one or more instructions that, when executed by the at least one processor, cause the at least one processor to:
a plurality of computing devices, each computing device of the plurality of computing devices comprising a respective subset of a plurality of computational nodes of a distributed machine learning model, each computational node comprising at least one parameter, each computing device comprising at least one processor and at least one non-transitory computer-readable medium including one or more instructions that, when executed by the at least one processor, cause the at least one processor to:
forward propagate each sample of a plurality of samples of training data through the distributed machine learning model to generate an output for each sample of the plurality of samples, wherein forward propagating comprises communicating, from each computing device of the plurality of computing devices other than a last computing device, an intermediate output of the respective subset of the plurality of computational nodes to a next computing device of the plurality of computing devices via the first message queue;
Claim 12
forward propagate each sample of a plurality of samples of training data through the distributed machine learning model to generate an output for each sample of the plurality of samples;
Claim 16
a first message queue, wherein forward propagating comprises communicating, from each computing device of the plurality of computing devices other than a last computing device, an intermediate output of the respective subset of the plurality of computational nodes to a next computing device of the plurality of computing devices via the first message queue.
determine a loss for each sample of the plurality of samples based on the output;
determine a loss for each sample of the plurality of samples based on the output;
backward propagate the loss for each sample of the plurality of samples to each computing device of the plurality of computing devices, wherein backward propagating comprises communicating, from each computing device of the plurality of computing devices other than a first computing device, a gradient value associated with the loss associated with the respective subset of the plurality of computational nodes to a previous computing device of the plurality of computing devices via the second message queue;
Claim 12
backward propagate the loss for each sample of the plurality of samples to each computing device of the plurality of computing devices;
Claim 17
a second message queue, wherein backward propagating comprises communicating, from each computing device of the plurality of computing devices other than a first computing device, a gradient value associated with the loss associated with the respective subset of the plurality of computational nodes to a previous computing device of the plurality of computing devices via the second message queue.
asynchronously update the at least one parameter of each computational node based on the loss for each sample as the loss for each sample is backward propagated while at least one of the plurality of samples is forward propagating through the distributed machine learning model;
asynchronously update the at least one parameter of each computational node based on the loss for each sample as the loss for each sample is backward propagated while at least one of the plurality of samples is forward propagating through the distributed machine learning model;
store, at each computing device of the plurality of computing devices, the at least one parameter of each computational node of the respective subset of the plurality of computational nodes as updated; and
store, at each computing device of the plurality of computing devices, the at least one parameter of each computational node of the respective subset of the plurality of computational nodes as updated;
communicate, from each computing device of the plurality of computing devices to all other computing devices of the plurality of computing devices, data associated with the at least one parameter of each computational node of the respective subset of the plurality of computational nodes as updated, each of the other computing devices of the plurality of computing devices storing the at least one parameter of each computational node as updated.
communicate, from each computing device of the plurality of computing devices to all other computing devices of the plurality of computing devices, data associated with the at least one parameter of each computational node of the respective subset of the plurality of computational nodes as updated, each of the other computing devices of the plurality of computing devices storing the at least one parameter of each computational node as updated;
Claim 13
Claim 13
wherein the distributed machine learning model comprises a deep neural network and each computational node of the plurality of computational nodes comprises a neuron, and
wherein the distributed machine learning model comprises a deep neural network and each computational node of the plurality of computational nodes comprises a neuron, and
wherein the plurality of computational nodes are divided into a plurality of layers, each subset of computational nodes being associated with at least one layer of the plurality of layers, each layer being implemented on at least one of the plurality of computing devices.
wherein the plurality of computational nodes are divided into a plurality of layers, each subset of computational nodes being associated with at least one layer of the plurality of layers, each layer being implemented on at least one of the plurality of computing devices.
Claim 14
Claim 14
wherein each computing device of the plurality of computing devices comprises at least one cache memory, and
wherein each computing device of the plurality of computing devices comprises at least one cache memory, and
wherein storing the at least one parameter of each computational node comprises storing the at least one parameter of each computational node in the at least one cache memory of each computing device of the plurality of computing devices.
wherein storing the at least one parameter of each computational node comprises storing the at least one parameter of each computational node in the at least one cache memory of each computing device of the plurality of computing devices.
Claim 15
Claim 15
a backup storage, wherein the backup storage is configured to store the at least one parameter of each computational node of the plurality of computational nodes.
a backup storage, wherein the backup storage is configured to store the at least one parameter of each computational node of the plurality of computational nodes.
Claim 16
wherein the first message queue comprises a first publisher-subscriber model queue, and wherein the second message queue comprises a second publisher-subscriber model queue.
Claim 17
wherein communicating the intermediate output via the first message queue comprises publishing, by each computing device of the plurality of computing devices other than the last computing device, the intermediate output of the respective subset of the plurality of computational nodes to the first publisher-subscriber model queue, wherein the next computing device of the plurality of computing devices comprises a subscriber to the first publisher-subscriber model queue, and
wherein communicating the gradient value via the second message queue comprises publishing, by each computing device of the plurality of computing devices other than the first computing device, the gradient value associated with the loss associated with the respective subset of the plurality of computational nodes to the second publisher-subscriber model queue, wherein the previous computing device of the plurality of computing devices comprises a subscriber to the second publisher-subscriber model queue.
Claim 18
Claim 18
a database configured to store a label for each sample of the plurality of samples, wherein determining the loss comprises determining the loss for each sample of the plurality of samples based on the output and the label.
a database configured to store a label for each sample of the plurality of samples, wherein determining the loss comprises determining the loss for each sample of the plurality of samples based on the output and the label.
Claim 19
Claim 12
determine a variance in the at least one parameter for each computational node; and
in response to determining the loss for the first sample satisfies the threshold, determine a variance in the at least one parameter for each computational node; and
generate at least one new computational node on at least one computing device of the plurality of computing devices based on the variance of at least one computational node of the respective subset of the plurality of computational nodes.
generate at least one new computational node on at least one computing device of the plurality of computing devices based on the variance of at least one computational node of the respective subset of the plurality of computational nodes.
Claim 20
Claim 20
A computer program product comprising at least one non-transitory computer-readable medium including one or more instructions that, when executed by at least one processor, cause the at least one processor to:
A computer program product for training a distributed machine learning model, the computer program product comprising at least one non-transitory computer-readable medium including one or more instructions that, when executed by at least one processor, cause the at least one processor to:
initialize a respective subset of a plurality of computational nodes of a distributed machine learning model, each computational node comprising at least one parameter;
initialize a respective subset of a plurality of computational nodes of a distributed machine learning model, each computational node comprising at least one parameter;
receive training data associated with a plurality of samples;
receive training data associated with a plurality of samples;
forward propagate each sample of the plurality of samples through the respective subset of the plurality of computational nodes of the distributed machine learning model to generate an intermediate output for each sample of the plurality of samples, wherein forward propagating comprises communicating, from each computing device of the plurality of computing devices other than a last computing device, an intermediate output of the respective subset of the plurality of computational nodes to a next computing device of the plurality of computing devices via a first message queue;
forward propagate each sample of the plurality of samples through the respective subset of the plurality of computational nodes of the distributed machine learning model to generate an intermediate output for each sample of the plurality of samples;
determine a loss for each sample of the plurality of samples based on the output;
determine a gradient value associated with a loss associated with the respective subset of the plurality of computational nodes for each sample of the plurality of samples;
backward propagate the loss for each sample of the plurality of samples to each computing device of the plurality of computing devices, wherein backward propagating comprises communicating, from each computing device of the plurality of computing devices other than a first computing device, a gradient value associated with the loss associated with the respective subset of the plurality of computational nodes to a previous computing device of the plurality of computing devices via a second message queue;
backward propagate the gradient value for each sample of the plurality of samples through the respective subset of the plurality of computational nodes of the distributed machine learning model;
asynchronously update the at least one parameter of each computational node based on the gradient value associated with the loss for each sample as the gradient value associated with the loss for each sample is backward propagated while at least one of the plurality of samples is forward propagating through the distributed machine learning model;
asynchronously update the at least one parameter of each computational node based on the gradient value associated with the loss for each sample as the gradient value associated with the loss for each sample is backward propagated while at least one of the plurality of samples is forward propagating through the distributed machine learning model;
store the at least one parameter of each computational node of the respective subset of the plurality of computational nodes as updated; and
store the at least one parameter of each computational node of the respective subset of the plurality of computational nodes as updated;
communicate, to at least one other computing device, data associated with the at least one parameter of each computational node of the respective subset of the plurality of computational nodes as updated to cause the at least one other computing device to store the at least one parameter of each computational node as updated.
communicate, to at least one computing device, data associated with the at least one parameter of each computational node of the respective subset of the plurality of computational nodes as updated to cause the at least one computing device to store the at least one parameter of each computational node as updated;
Regarding claim 7 (similarly claim 16), U.S. Patent No. 11,928,571 teaches all of the limitations of claim 1 (claim 12), as stated in the chart. However, U.S. Patent No. 11,928,571 does not explicitly teach wherein the first message queue comprises a first publisher-subscriber model queue, and wherein the second message queue comprises a second publisher-subscriber model queue.
Tomioka teaches wherein the first message queue comprises a first publisher-subscriber model queue, and wherein the second message queue comprises a second publisher-subscriber model queue (Tomioka, [0047]-[0050] - teaches message passing using message queues for communicating data both forward and backward; see also Tomioka, [0033]).
It would have been obvious to one of ordinary skill in the art before the filing date of the claimed invention to modify U.S. Patent No. 11,928,571 with the teachings of Tomioka in order to increase training efficiency in the field of distributed machine learning (Tomioka, [0027] – “In various examples described herein model parallelism is combined with asynchronous updates of the neural network subgraph parameters at the individual worker nodes. This scheme is found to give extremely good efficiency (as explained with reference to FIG. 9 below) and is found empirically to work well in practice despite the fact that the conventional theoretical convergence guarantee for stochastic gradient descent would not apply to the asynchronous updates of the subgraph parameters.”).
Regarding claim 8 (similarly claim 17), U.S. Patent No. 11,928,571 in view of Tomioka teaches all of the limitations of claim 7 (claim 16), as stated in the chart. Tomioka further teaches
wherein communicating the intermediate output via the first message queue comprises publishing, by each computing device of the plurality of computing devices other than the last computing device, the intermediate output of the respective subset of the plurality of computational nodes to the first publisher-subscriber model queue, wherein the next computing device of the plurality of computing devices comprises a subscriber to the first publisher-subscriber model queue (Tomioka, [0047]-[0050] - teaches message passing using message queues for communicating data both forward and backward; see also Tomioka, [0033]), and
wherein communicating the gradient value via the second message queue comprises publishing, by each computing device of the plurality of computing devices other than the first computing device, the gradient value associated with the loss associated with the respective subset of the plurality of computational nodes to the second publisher-subscriber model queue, wherein the previous computing device of the plurality of computing devices comprises a subscriber to the second publisher-subscriber model queue (Tomioka, [0047]-[0050] - teaches message passing using message queues for communicating data both forward and backward; see also Tomioka, [0033]).
It would have been obvious to one of ordinary skill in the art before the filing date of the claimed invention to combine the teachings of U.S. Patent No. 11,928,571 and Tomioka in order to communicate between computing devices to increase training efficiency (Tomioka, [0027]).
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claim(s) 1-20 is/are rejected under 35 U.S.C. 101, because the claim(s) is/are directed to an abstract idea, and because the claim elements, whether considered individually or in combination, do not amount to significantly more than the abstract idea, see Alice Corporation Pty. Ltd. V. CLS Bank International et al., 573 US 208 (2014).
Regarding claim 1, the claim is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1 Analysis: Claim 1 is directed to a method, which is directed to a process, one of the statutory categories.
Step 2A Prong One Analysis: The claim recites a(n) computer-implemented method.
The limitation of initializing a distributed machine learning model on a plurality of computing devices …, as drafted, is a process that, under its broadest reasonable interpretation, covers a mental process. The limitation is directed to observation, evaluation, judgment and opinion and is a process capable of being performed by a human mentally or using pen and paper.
The limitation of determining a loss for each sample of the plurality of samples based on the output, as drafted, is a process that, under its broadest reasonable interpretation, covers a mathematical concept. The limitation encompasses calculating a loss.
The limitation of asynchronously updating the at least one parameter of each computational node based on the loss for each sample as the loss for each sample is backward propagated while at least one of the plurality of samples is forward propagating through the distributed machine learning model, as drafted, is a process that, under its broadest reasonable interpretation, covers a mathematical concept. The limitation encompasses calculating parameter changes.
If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind, then it falls within the "Mental Processes" grouping. If a claim limitation, under its broadest reasonable interpretation, covers performance of mathematical concepts, then it falls within the "Mathematical Concepts" grouping. Accordingly, the claim recites an abstract idea.
Step 2A Prong Two Analysis: With respect to the abstract idea, the judicial exception is not integrated into a practical application.
The claim recites additional element(s) – computer-implemented, plurality of computing devices, first message queue, second message queue. The additional element(s) is/are recited at a high-level of generality (i.e., as generic computer components performing generic computer functions of executing instructions on the computers) such that it amounts to no more than mere instructions to apply the exception using generic computer components (MPEP 2106.05(b)).
The claim recites additional element(s) – distributed machine learning model. The additional element(s) is/are recited at a high-level of generality such that it amounts to no more than indicating a field of use or technological environment in which to apply the judicial exception (MPEP 2106.05(h)).
The claim recites forward propagating each sample of the plurality of samples through the distributed machine learning model to generate an output for each sample of the plurality of samples ... which is simply applying the model recited at a high level of generality and amounts to the recitation of the words “apply it” (or an equivalent) or amounts to no more than mere instructions to implement an abstract idea or other exception on a computer (MPEP 2106.05(f)).
The claim recites receiving training data associated with a plurality of samples at a first computing device of the plurality of computing devices; ... wherein forward propagating comprises communicating, from each computing device of the plurality of computing devices other than a last computing device, an intermediate output of the respective subset of the plurality of computational nodes to a next computing device of the plurality of computing devices via a first message queue; backward propagating the loss for each sample of the plurality of samples to each computing device of the plurality of computing devices, wherein backward propagating comprises communicating, from each computing device of the plurality of computing devices other than the first computing device, a gradient value associated with the loss associated with the respective subset of the plurality of computational nodes to a previous computing device of the plurality of computing devices via a second message queue; storing, at each computing device of the plurality of computing devices, the at least one parameter of each computational node of the respective subset of the plurality of computational nodes as updated; communicating, from each computing device of the plurality of computing devices to all other computing devices of the plurality of computing devices, data associated with the at least one parameter of each computational node of the respective subset of the plurality of computational nodes as updated, each of the other computing devices of the plurality of computing devices storing the at least one parameter of each computational node as updated, which is simply acquiring data, transmitting data and storing data recited at a high level of generality. This is nothing more than insignificant extra-solution activity (MPEP 2106.05(g)).
The claim recites ... the distributed machine learning model comprising a plurality of computational nodes, each computing device of the plurality of computing devices comprising a respective subset of the plurality of computational nodes, each computational node comprising at least one parameter which is simply additional information regarding the model, and the element(s) do(es) not apply the exception in a meaningful way (MPEP 2106.05(e)).
Accordingly, the additional element(s) do(es) not integrate the abstract idea into a practical application because the additional element(s) do(es) not impose any meaningful limits on practicing the abstract idea, and, therefore, the claim is directed to an abstract idea.
Step 2B Analysis: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to the integration of the abstract idea into a practical application, the additional element(s) of:
computer-implemented, plurality of computing devices, first message queue, second message queue amount(s) to no more than mere instructions to apply the exception using generic computer components (MPEP 2106.05(b))
applying the model amount(s) to no more than mere instructions to apply the exception (MPEP 2106.05(f))
acquiring data, transmitting data and storing data amount(s) to no more than insignificant extra-solution activity (MPEP 2106.05(g)), wherein the insignificant extra-solution activity is the well-understood routine and conventional activit(y/ies) of receiving or transmitting data over a network and/or storing and retrieving information in memory (MPEP 2016.05(d))
distributed machine learning model amount(s) to no more than indicating a field of use or technological environment in which to apply the judicial exception (MPEP 2106.05(h))
additional information regarding the model do(es) not apply the exception in a meaningful way (MPEP 2106.05(e))
The additional element(s) do(es) not provide an inventive concept, and, therefore, the claim is not patent eligible.
Regarding claim 2, the claim is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1 Analysis: Claim 2 is directed to a method, which is directed to a process, one of the statutory categories.
Step 2A Prong One Analysis: The claim recites a(n) computer-implemented method. The Step 2A Prong One Analysis for claim 1 is applicable here since claim 2 carries out the method of claim 1 but for the recitation of additional element(s) of wherein the distributed machine learning model comprises a deep neural network and each computational node of the plurality of computational nodes comprises a neuron.
Step 2A Prong Two Analysis: With respect to the abstract idea, the judicial exception is not integrated into a practical application.
The claim recites wherein the distributed machine learning model comprises a deep neural network and each computational node of the plurality of computational nodes comprises a neuron which is simply additional information regarding the model, and the element(s) do(es) not apply the exception in a meaningful way (MPEP 2106.05(e)).
The claim recites additional element(s) – deep neural network. The additional element(s) is/are recited at a high-level of generality such that it amounts to no more than indicating a field of use or technological environment in which to apply the judicial exception (MPEP 2106.05(h)).
Accordingly, the additional element(s) do(es) not integrate the abstract idea into a practical application because the additional element(s) do(es) not impose any meaningful limits on practicing the abstract idea, and, therefore, the claim is directed to an abstract idea.
Step 2B Analysis: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to the integration of the abstract idea into a practical application, the additional element(s) of:
deep neural network amount(s) to no more than indicating a field of use or technological environment in which to apply the judicial exception (MPEP 2106.05(h))
additional information regarding the model do(es) not apply the exception in a meaningful way (MPEP 2106.05(e))
The additional element(s) do(es) not provide an inventive concept, and, therefore, the claim is not patent eligible.
Regarding claim 3, the claim is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1 Analysis: Claim 3 is directed to a method, which is directed to a process, one of the statutory categories.
Step 2A Prong One Analysis: The claim recites a(n) computer-implemented method. The Step 2A Prong One Analysis for claim 1 is applicable here since claim 3 carries out the method of claim 1 but for the recitation of additional element(s) of wherein the plurality of computational nodes are divided into a plurality of layers, each subset of computational nodes being associated with at least one layer of the plurality of layers, each layer being implemented on at least one of the plurality of computing devices.
Step 2A Prong Two Analysis: With respect to the abstract idea, the judicial exception is not integrated into a practical application. In particular, the claim recites additional information regarding the model and the element(s) do(es) not apply the exception in a meaningful way (MPEP 2106.05(e)). Accordingly, the additional element(s) do(es) not integrate the abstract idea into a practical application because the additional element(s) do(es) not impose any meaningful limits on practicing the abstract idea, and, therefore, the claim is directed to an abstract idea.
Step 2B Analysis: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to the integration of the abstract idea into a practical application, the additional element(s) of additional information regarding the model do(es) not apply the exception in a meaningful way (MPEP 2106.05(e)). Not applying the exception in a meaningful way does not provide an inventive concept, and, therefore, the claim is not patent eligible.
Regarding claim 4, the claim is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1 Analysis: Claim 4 is directed to a method, which is directed to a process, one of the statutory categories.
Step 2A Prong One Analysis: The claim recites a(n) computer-implemented method. The Step 2A Prong One Analysis for claim 1 is applicable here since claim 4 carries out the method of claim 1 but for the recitation of additional element(s) of wherein the at least one parameter of each computational node comprises at least one of a weight parameter, a bias parameter, or any combination thereof.
Step 2A Prong Two Analysis: With respect to the abstract idea, the judicial exception is not integrated into a practical application. In particular, the claim recites additional information regarding the parameters and the element(s) do(es) not apply the exception in a meaningful way (MPEP 2106.05(e)). Accordingly, the additional element(s) do(es) not integrate the abstract idea into a practical application because the additional element(s) do(es) not impose any meaningful limits on practicing the abstract idea, and, therefore, the claim is directed to an abstract idea.
Step 2B Analysis: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to the integration of the abstract idea into a practical application, the additional element(s) of additional information regarding the parameters do(es) not apply the exception in a meaningful way (MPEP 2106.05(e)). Not applying the exception in a meaningful way does not provide an inventive concept, and, therefore, the claim is not patent eligible.
Regarding claim 5, the claim is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1 Analysis: Claim 5 is directed to a method, which is directed to a process, one of the statutory categories.
Step 2A Prong One Analysis: The claim recites a(n) computer-implemented method. The Step 2A Prong One Analysis for claim 1 is applicable here since claim 5 carries out the method of claim 1 but for the recitation of additional element(s) of wherein storing the at least one parameter of each computational node comprises storing the at least one parameter of each computational node in at least one cache memory of each computing device of the plurality of computing devices.
Step 2A Prong Two Analysis: With respect to the abstract idea, the judicial exception is not integrated into a practical application.
The claim recites additional element(s) – at least one cache memory. The additional element(s) is/are recited at a high-level of generality (i.e., as generic computer components performing generic computer functions of executing instructions on the computers) such that it amounts to no more than mere instructions to apply the exception using generic computer components (MPEP 2106.05(b)).
The claim recites wherein storing the at least one parameter of each computational node comprises storing the at least one parameter of each computational node in at least one cache memory of each computing device of the plurality of computing devices, which is simply storing data recited at a high level of generality. This is nothing more than insignificant extra-solution activity (MPEP 2106.05(g)).
Accordingly, the additional element(s) do(es) not integrate the abstract idea into a practical application because the additional element(s) do(es) not impose any meaningful limits on practicing the abstract idea, and, therefore, the claim is directed to an abstract idea.
Step 2B Analysis: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to the integration of the abstract idea into a practical application, the additional element(s) of:
at least one cache memory amount(s) to no more than mere instructions to apply the exception using generic computer components (MPEP 2106.05(b))
storing data amount(s) to no more than insignificant extra-solution activity (MPEP 2106.05(g)), wherein the insignificant extra-solution activity is the well-understood routine and conventional activit(y/ies) of storing and retrieving information in memory (MPEP 2016.05(d))
The additional element(s) do(es) not provide an inventive concept, and, therefore, the claim is not patent eligible.
Regarding claim 6, the claim is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1 Analysis: Claim 6 is directed to a method, which is directed to a process, one of the statutory categories.
Step 2A Prong One Analysis: The claim recites a(n) computer-implemented method. The Step 2A Prong One Analysis for claim 5 is applicable here since claim 6 carries out the method of claim 5 but for the recitation of additional element(s) of storing, in a backup storage, the at least one parameter of each computational node of the plurality of computational nodes.
Step 2A Prong Two Analysis: With respect to the abstract idea, the judicial exception is not integrated into a practical application.
The claim recites additional element(s) – backup storage. The additional element(s) is/are recited at a high-level of generality (i.e., as generic computer components performing generic computer functions of executing instructions on the computers) such that it amounts to no more than mere instructions to apply the exception using generic computer components (MPEP 2106.05(b)).
The claim recites storing, in a backup storage, the at least one parameter of each computational node of the plurality of computational nodes, which is simply storing data recited at a high level of generality. This is nothing more than insignificant extra-solution activity (MPEP 2106.05(g)).
Accordingly, the additional element(s) do(es) not integrate the abstract idea into a practical application because the additional element(s) do(es) not impose any meaningful limits on practicing the abstract idea, and, therefore, the claim is directed to an abstract idea.
Step 2B Analysis: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to the integration of the abstract idea into a practical application, the additional element(s) of:
backup storage amount(s) to no more than mere instructions to apply the exception using generic computer components (MPEP 2106.05(b))
storing data amount(s) to no more than insignificant extra-solution activity (MPEP 2106.05(g)), wherein the insignificant extra-solution activity is the well-understood routine and conventional activit(y/ies) of storing and retrieving information in memory (MPEP 2016.05(d))
The additional element(s) do(es) not provide an inventive concept, and, therefore, the claim is not patent eligible.
Regarding claim 7, the claim is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1 Analysis: Claim 7 is directed to a method, which is directed to a process, one of the statutory categories.
Step 2A Prong One Analysis: The claim recites a(n) computer-implemented method. The Step 2A Prong One Analysis for claim 1 is applicable here since claim 7 carries out the method of claim 1 but for the recitation of additional element(s) of wherein the first message queue comprises a first publisher-subscriber model queue, and wherein the second message queue comprises a second publisher-subscriber model queue.
Step 2A Prong Two Analysis: With respect to the abstract idea, the judicial exception is not integrated into a practical application.
The claim recites wherein the first message queue comprises a first publisher-subscriber model queue, and wherein the second message queue comprises a second publisher-subscriber model queue which is simply additional information regarding the computer components, and the element(s) do(es) not apply the exception in a meaningful way (MPEP 2106.05(e)).
The claim recites additional element(s) – first publisher-subscriber model queue, second publisher-subscriber model queue. The additional element(s) is/are recited at a high-level of generality (i.e., as generic computer components performing generic computer functions of executing instructions on the computers) such that it amounts to no more than mere instructions to apply the exception using generic computer components (MPEP 2106.05(b)).
Accordingly, the additional element(s) do(es) not integrate the abstract idea into a practical application because the additional element(s) do(es) not impose any meaningful limits on practicing the abstract idea, and, therefore, the claim is directed to an abstract idea.
Step 2B Analysis: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to the integration of the abstract idea into a practical application, the additional element(s) of:
first publisher-subscriber model queue, second publisher-subscriber model queue amount(s) to no more than mere instructions to apply the exception using generic computer components (MPEP 2106.05(b))
additional information regarding the computer components do(es) not apply the exception in a meaningful way (MPEP 2106.05(e))
The additional element(s) do(es) not provide an inventive concept, and, therefore, the claim is not patent eligible.
Regarding claim 8, the claim is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1 Analysis: Claim 8 is directed to a method, which is directed to a process, one of the statutory categories.
Step 2A Prong One Analysis: The claim recites a(n) computer-implemented method. The Step 2A Prong One Analysis for claim 7 is applicable here since claim 8 carries out the method of claim 7 but for the recitation of additional element(s) of wherein communicating the intermediate output via the first message queue comprises publishing, by each computing device of the plurality of computing devices other than the last computing device, the intermediate output of the respective subset of the plurality of computational nodes to the first publisher-subscriber model queue, wherein the next computing device of the plurality of computing devices comprises a subscriber to the first publisher-subscriber model queue, and wherein communicating the gradient value via the second message queue comprises publishing, by each computing device of the plurality of computing devices other than the first computing device, the gradient value associated with the loss associated with the respective subset of the plurality of computational nodes to the second publisher-subscriber model queue, wherein the previous computing device of the plurality of computing devices comprises a subscriber to the second publisher-subscriber model queue.
Step 2A Prong Two Analysis: With respect to the abstract idea, the judicial exception is not integrated into a practical application.
The claim recites additional element(s) – subscriber. The additional element(s) is/are recited at a high-level of generality (i.e., as generic computer components performing generic computer functions of executing instructions on the computers) such that it amounts to no more than mere instructions to apply the exception using generic computer components (MPEP 2106.05(b)).
The claim recites wherein communicating the intermediate output via the first message queue comprises publishing, by each computing device of the plurality of computing devices other than the last computing device, the intermediate output of the respective subset of the plurality of computational nodes to the first publisher-subscriber model queue ...; wherein communicating the gradient value via the second message queue comprises publishing, by each computing device of the plurality of computing devices other than the first computing device, the gradient value associated with the loss associated with the respective subset of the plurality of computational nodes to the second publisher-subscriber model queue ..., which is simply storing and/or transmitting data recited at a high level of generality. This is nothing more than insignificant extra-solution activity (MPEP 2106.05(g)).
The claim recites ... wherein the next computing device of the plurality of computing devices comprises a subscriber to the first publisher-subscriber model queue; ... wherein the previous computing device of the plurality of computing devices comprises a subscriber to the second publisher-subscriber model queue which is simply additional information regarding the computer components, and the element(s) do(es) not apply the exception in a meaningful way (MPEP 2106.05(e)).
Accordingly, the additional element(s) do(es) not integrate the abstract idea into a practical application because the additional element(s) do(es) not impose any meaningful limits on practicing the abstract idea, and, therefore, the claim is directed to an abstract idea.
Step 2B Analysis: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to the integration of the abstract idea into a practical application, the additional element(s) of:
subscriber amount(s) to no more than mere instructions to apply the exception using generic computer components (MPEP 2106.05(b))
storing and/or transmitting data amount(s) to no more than insignificant extra-solution activity (MPEP 2106.05(g)), wherein the insignificant extra-solution activity is the well-understood routine and conventional activit(y/ies) of receiving or transmitting data over a network and/or storing and retrieving information in memory (MPEP 2016.05(d))
additional information regarding the computer components do(es) not apply the exception in a meaningful way (MPEP 2106.05(e))
The additional element(s) do(es) not provide an inventive concept, and, therefore, the claim is not patent eligible.
Regarding claim 9, the claim is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1 Analysis: Claim 9 is directed to a method, which is directed to a process, one of the statutory categories.
Step 2A Prong One Analysis: The claim recites a(n) computer-implemented method.
The limitation of … wherein determining the loss comprises determining the loss for each sample of the plurality of samples based on the output and the label, as drafted, is a process that, under its broadest reasonable interpretation, covers a mathematical concept. The limitation encompasses calculating a loss.
If a claim limitation, under its broadest reasonable interpretation, covers performance of mathematical concepts, then it falls within the "Mathematical Concepts" grouping. Accordingly, the claim recites an abstract idea.
Step 2A Prong Two Analysis: With respect to the abstract idea, the judicial exception is not integrated into a practical application.
The claim recites additional element(s) – database. The additional element(s) is/are recited at a high-level of generality (i.e., as generic computer components performing generic computer functions of executing instructions on the computers) such that it amounts to no more than mere instructions to apply the exception using generic computer components (MPEP 2106.05(b)).
The claim recites storing, in a database, a label for each sample of the plurality of samples ..., which is simply storing data recited at a high level of generality. This is nothing more than insignificant extra-solution activity (MPEP 2106.05(g)).
Accordingly, the additional element(s) do(es) not integrate the abstract idea into a practical application because the additional element(s) do(es) not impose any meaningful limits on practicing the abstract idea, and, therefore, the claim is directed to an abstract idea.
Step 2B Analysis: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to the integration of the abstract idea into a practical application, the additional element(s) of:
database amount(s) to no more than mere instructions to apply the exception using generic computer components (MPEP 2106.05(b))
storing data amount(s) to no more than insignificant extra-solution activity (MPEP 2106.05(g)), wherein the insignificant extra-solution activity is the well-understood routine and conventional activit(y/ies) of storing and retrieving information in memory (MPEP 2016.05(d))
The additional element(s) do(es) not provide an inventive concept, and, therefore, the claim is not patent eligible.
Regarding claim 10, the claim is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1 Analysis: Claim 10 is directed to a method, which is directed to a process, one of the statutory categories.
Step 2A Prong One Analysis: The claim recites a(n) computer-implemented method.
The limitation of determining a variance in the at least one parameter for each computational node, as drafted, is a process that, under its broadest reasonable interpretation, covers a mathematical concept. The limitation encompasses calculating a variance.
The limitation of generating at least one new computational node on at least one computing device of the plurality of computing devices based on the variance of at least one computational node of the respective subset of the plurality of computational nodes, as drafted, is a process that, under its broadest reasonable interpretation, covers a mental process. The limitation is directed to observation, evaluation, judgment and opinion and is a process capable of being performed by a human mentally or using pen and paper.
If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind, then it falls within the "Mental Processes" grouping. If a claim limitation, under its broadest reasonable interpretation, covers performance of mathematical concepts, then it falls within the "Mathematical Concepts" grouping. Accordingly, the claim recites an abstract idea.
Step 2A Prong Two Analysis: With respect to the abstract idea, the judicial exception is not integrated
into a practical application. The claim does not recite any additional elements which integrate the
abstract idea into a practical application and, therefore, does not impose any meaningful limits on
practicing the abstract idea. Therefore, the claim is directed to an abstract idea.
Step 2B Analysis: The claim does not include additional elements that are sufficient to amount to
significantly more than the judicial exception. As discussed above with respect to the integration of the
abstract idea into a practical application, the claim does not recite any additional elements which
provide an inventive concept, and, therefore, the claim is not patent eligible.
Regarding claim 11, the claim is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1 Analysis: Claim 11 is directed to a method, which is directed to a process, one of the statutory categories.
Step 2A Prong One Analysis: The claim recites a(n) computer-implemented method.
The limitation of determining the loss for a first sample of the plurality of samples satisfies a threshold associated with at least one computing device of the plurality of computing devices becoming unavailable before determining the variance in the at least one parameter for each computational node, as drafted, is a process that, under its broadest reasonable interpretation, covers a mental process. The limitation is directed to observation, evaluation, judgment and opinion and is a process capable of being performed by a human mentally or using pen and paper.
If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind, then it falls within the "Mental Processes" grouping. Accordingly, the claim recites an abstract idea.
Step 2A Prong Two Analysis: With respect to the abstract idea, the judicial exception is not integrated
into a practical application. The claim does not recite any additional elements which integrate the
abstract idea into a practical application and, therefore, does not impose any meaningful limits on
practicing the abstract idea. Therefore, the claim is directed to an abstract idea.
Step 2B Analysis: The claim does not include additional elements that are sufficient to amount to
significantly more than the judicial exception. As discussed above with respect to the integration of the
abstract idea into a practical application, the claim does not recite any additional elements which
provide an inventive concept, and, therefore, the claim is not patent eligible.
Regarding claim 12, the claim is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1 Analysis: Claim 12 is directed to a system with computing devices, which is directed to a machine, one of the statutory categories.
Step 2A Prong One Analysis: The claim recites a(n) system.
The limitation of determine a loss for each sample of the plurality of samples based on the output, as drafted, is a process that, under its broadest reasonable interpretation, covers a mathematical concept. The limitation encompasses calculating a loss.
The limitation of asynchronously update the at least one parameter of each computational node based on the loss for each sample as the loss for each sample is backward propagated while at least one of the plurality of samples is forward propagating through the distributed machine learning model, as drafted, is a process that, under its broadest reasonable interpretation, covers a mathematical concept. The limitation encompasses calculating parameter changes.
If a claim limitation, under its broadest reasonable interpretation, covers performance of mathematical concepts, then it falls within the "Mathematical Concepts" grouping. Accordingly, the claim recites an abstract idea.
Step 2A Prong Two Analysis: With respect to the abstract idea, the judicial exception is not integrated into a practical application.
The claim recites additional element(s) – system, first message queue, second message queue, plurality of computing devices, at least one processor, at least one non-transitory computer-readable medium, one or more instructions. The additional element(s) is/are recited at a high-level of generality (i.e., as generic computer components performing generic computer functions of executing instructions on the computers) such that it amounts to no more than mere instructions to apply the exception using generic computer components (MPEP 2106.05(b)).
The claim recites additional element(s) – distributed machine learning model. The additional element(s) is/are recited at a high-level of generality such that it amounts to no more than indicating a field of use or technological environment in which to apply the judicial exception (MPEP 2106.05(h)).
The claim recites forward propagate each sample of a plurality of samples of training data through the distributed machine learning model to generate an output for each sample of the plurality of samples ... which is simply applying the model recited at a high level of generality and amounts to the recitation of the words “apply it” (or an equivalent) or amounts to no more than mere instructions to implement an abstract idea or other exception on a computer (MPEP 2106.05(f)).
The claim recites ... wherein forward propagating comprises communicating, from each computing device of the plurality of computing devices other than a last computing device, an intermediate output of the respective subset of the plurality of computational nodes to a next computing device of the plurality of computing devices via the first message queue; backward propagate the loss for each sample of the plurality of samples to each computing device of the plurality of computing devices, wherein backward propagating comprises communicating, from each computing device of the plurality of computing devices other than a first computing device, a gradient value associated with the loss associated with the respective subset of the plurality of computational nodes to a previous computing device of the plurality of computing devices via the second message queue; store, at each computing device of the plurality of computing devices, the at least one parameter of each computational node of the respective subset of the plurality of computational nodes as updated; communicate, from each computing device of the plurality of computing devices to all other computing devices of the plurality of computing devices, data associated with the at least one parameter of each computational node of the respective subset of the plurality of computational nodes as updated, each of the other computing devices of the plurality of computing devices storing the at least one parameter of each computational node as updated, which is simply transmitting data and storing data recited at a high level of generality. This is nothing more than insignificant extra-solution activity (MPEP 2106.05(g)).
The claim recites each computing device of the plurality of computing devices comprising a respective subset of a plurality of computational nodes of a distributed machine learning model, each computational node comprising at least one parameter which is simply additional information regarding the model, and the element(s) do(es) not apply the exception in a meaningful way (MPEP 2106.05(e)).
Accordingly, the additional element(s) do(es) not integrate the abstract idea into a practical application because the additional element(s) do(es) not impose any meaningful limits on practicing the abstract idea, and, therefore, the claim is directed to an abstract idea.
Step 2B Analysis: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to the integration of the abstract idea into a practical application, the additional element(s) of:
system, first message queue, second message queue, plurality of computing devices, at least one processor, at least one non-transitory computer-readable medium, one or more instructions amount(s) to no more than mere instructions to apply the exception using generic computer components (MPEP 2106.05(b))
applying the model amount(s) to no more than mere instructions to apply the exception (MPEP 2106.05(f))
transmitting data and storing data amount(s) to no more than insignificant extra-solution activity (MPEP 2106.05(g)), wherein the insignificant extra-solution activity is the well-understood routine and conventional activit(y/ies) of receiving or transmitting data over a network and/or storing and retrieving information in memory (MPEP 2016.05(d))
distributed machine learning model amount(s) to no more than indicating a field of use or technological environment in which to apply the judicial exception (MPEP 2106.05(h))
additional information regarding the model do(es) not apply the exception in a meaningful way (MPEP 2106.05(e))
The additional element(s) do(es) not provide an inventive concept, and, therefore, the claim is not patent eligible.
Regarding claim 13, the claim is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1 Analysis: Claim 13 is directed to a system with computing devices, which is directed to a machine, one of the statutory categories.
Step 2A Prong One Analysis: The claim recites a(n) system. The Step 2A Prong One Analysis for claim 12 is applicable here since claim 13 carries out the system of claim 12 but for the recitation of additional element(s) of wherein the distributed machine learning model comprises a deep neural network and each computational node of the plurality of computational nodes comprises a neuron, and wherein the plurality of computational nodes are divided into a plurality of layers, each subset of computational nodes being associated with at least one layer of the plurality of layers, each layer being implemented on at least one of the plurality of computing devices.
Step 2A Prong Two Analysis: With respect to the abstract idea, the judicial exception is not integrated into a practical application.
The claim recites wherein the distributed machine learning model comprises a deep neural network and each computational node of the plurality of computational nodes comprises a neuron, and wherein the plurality of computational nodes are divided into a plurality of layers, each subset of computational nodes being associated with at least one layer of the plurality of layers, each layer being implemented on at least one of the plurality of computing devices which is simply additional information regarding the model, and the element(s) do(es) not apply the exception in a meaningful way (MPEP 2106.05(e)).
The claim recites additional element(s) – deep neural network. The additional element(s) is/are recited at a high-level of generality such that it amounts to no more than indicating a field of use or technological environment in which to apply the judicial exception (MPEP 2106.05(h)).
Accordingly, the additional element(s) do(es) not integrate the abstract idea into a practical application because the additional element(s) do(es) not impose any meaningful limits on practicing the abstract idea, and, therefore, the claim is directed to an abstract idea.
Step 2B Analysis: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to the integration of the abstract idea into a practical application, the additional element(s) of:
deep neural network amount(s) to no more than indicating a field of use or technological environment in which to apply the judicial exception (MPEP 2106.05(h))
additional information regarding the model do(es) not apply the exception in a meaningful way (MPEP 2106.05(e))
The additional element(s) do(es) not provide an inventive concept, and, therefore, the claim is not patent eligible.
Regarding claim 14, the claim is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1 Analysis: Claim 14 is directed to a system with computing devices, which is directed to a machine, one of the statutory categories.
Step 2A Prong One Analysis: The claim recites a(n) system. The Step 2A Prong One Analysis for claim 12 is applicable here since claim 14 carries out the system of claim 12 but for the recitation of additional element(s) of wherein each computing device of the plurality of computing devices comprises at least one cache memory, and wherein storing the at least one parameter of each computational node comprises storing the at least one parameter of each computational node in the at least one cache memory of each computing device of the plurality of computing devices.
Step 2A Prong Two Analysis: With respect to the abstract idea, the judicial exception is not integrated into a practical application.
The claim recites additional element(s) – at least one cache memory. The additional element(s) is/are recited at a high-level of generality (i.e., as generic computer components performing generic computer functions of executing instructions on the computers) such that it amounts to no more than mere instructions to apply the exception using generic computer components (MPEP 2106.05(b)).
The claim recites wherein storing the at least one parameter of each computational node comprises storing the at least one parameter of each computational node in the at least one cache memory of each computing device of the plurality of computing devices, which is simply storing data recited at a high level of generality. This is nothing more than insignificant extra-solution activity (MPEP 2106.05(g)).
The claim recites wherein each computing device of the plurality of computing devices comprises at least one cache memory which is simply additional information regarding the computing devices, and the element(s) do(es) not apply the exception in a meaningful way (MPEP 2106.05(e)).
Accordingly, the additional element(s) do(es) not integrate the abstract idea into a practical application because the additional element(s) do(es) not impose any meaningful limits on practicing the abstract idea, and, therefore, the claim is directed to an abstract idea.
Step 2B Analysis: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to the integration of the abstract idea into a practical application, the additional element(s) of:
at least one cache memory amount(s) to no more than mere instructions to apply the exception using generic computer components (MPEP 2106.05(b))
storing data amount(s) to no more than insignificant extra-solution activity (MPEP 2106.05(g)), wherein the insignificant extra-solution activity is the well-understood routine and conventional activit(y/ies) of storing and retrieving information in memory (MPEP 2016.05(d))
additional information regarding the computing devices do(es) not apply the exception in a meaningful way (MPEP 2106.05(e))
The additional element(s) do(es) not provide an inventive concept, and, therefore, the claim is not patent eligible.
Regarding claim 15, the claim is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1 Analysis: Claim 15 is directed to a system with computing devices, which is directed to a machine, one of the statutory categories.
Step 2A Prong One Analysis: The claim recites a(n) system. The Step 2A Prong One Analysis for claim 14 is applicable here since claim 15 carries out the system of claim 14 but for the recitation of additional element(s) of a backup storage, wherein the backup storage is configured to store the at least one parameter of each computational node of the plurality of computational nodes.
Step 2A Prong Two Analysis: With respect to the abstract idea, the judicial exception is not integrated into a practical application.
The claim recites additional element(s) – backup storage. The additional element(s) is/are recited at a high-level of generality (i.e., as generic computer components performing generic computer functions of executing instructions on the computers) such that it amounts to no more than mere instructions to apply the exception using generic computer components (MPEP 2106.05(b)).
The claim recites a backup storage, wherein the backup storage is configured to store the at least one parameter of each computational node of the plurality of computational nodes, which is simply storing data recited at a high level of generality. This is nothing more than insignificant extra-solution activity (MPEP 2106.05(g)).
Accordingly, the additional element(s) do(es) not integrate the abstract idea into a practical application because the additional element(s) do(es) not impose any meaningful limits on practicing the abstract idea, and, therefore, the claim is directed to an abstract idea.
Step 2B Analysis: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to the integration of the abstract idea into a practical application, the additional element(s) of:
backup storage amount(s) to no more than mere instructions to apply the exception using generic computer components (MPEP 2106.05(b))
storing data amount(s) to no more than insignificant extra-solution activity (MPEP 2106.05(g)), wherein the insignificant extra-solution activity is the well-understood routine and conventional activit(y/ies) of storing and retrieving information in memory (MPEP 2016.05(d))
The additional element(s) do(es) not provide an inventive concept, and, therefore, the claim is not patent eligible.
Regarding claim 16, the claim is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1 Analysis: Claim 16 is directed to a system with computing devices, which is directed to a machine, one of the statutory categories.
Step 2A Prong One Analysis: The claim recites a(n) system. The Step 2A Prong One Analysis for claim 12 is applicable here since claim 16 carries out the system of claim 12 but for the recitation of additional element(s) of wherein the first message queue comprises a first publisher-subscriber model queue, and wherein the second message queue comprises a second publisher-subscriber model queue.
Step 2A Prong Two Analysis: With respect to the abstract idea, the judicial exception is not integrated into a practical application.
The claim recites wherein the first message queue comprises a first publisher-subscriber model queue, and wherein the second message queue comprises a second publisher-subscriber model queue which is simply additional information regarding the computer components, and the element(s) do(es) not apply the exception in a meaningful way (MPEP 2106.05(e)).
The claim recites additional element(s) – first publisher-subscriber model queue, second publisher-subscriber model queue. The additional element(s) is/are recited at a high-level of generality (i.e., as generic computer components performing generic computer functions of executing instructions on the computers) such that it amounts to no more than mere instructions to apply the exception using generic computer components (MPEP 2106.05(b)).
Accordingly, the additional element(s) do(es) not integrate the abstract idea into a practical application because the additional element(s) do(es) not impose any meaningful limits on practicing the abstract idea, and, therefore, the claim is directed to an abstract idea.
Step 2B Analysis: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to the integration of the abstract idea into a practical application, the additional element(s) of:
first publisher-subscriber model queue, second publisher-subscriber model queue amount(s) to no more than mere instructions to apply the exception using generic computer components (MPEP 2106.05(b))
additional information regarding the computer components do(es) not apply the exception in a meaningful way (MPEP 2106.05(e))
The additional element(s) do(es) not provide an inventive concept, and, therefore, the claim is not patent eligible.
Regarding claim 17, the claim is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1 Analysis: Claim 17 is directed to a system with computing devices, which is directed to a machine, one of the statutory categories.
Step 2A Prong One Analysis: The claim recites a(n) system. The Step 2A Prong One Analysis for claim 16 is applicable here since claim 17 carries out the system of claim 16 but for the recitation of additional element(s) of wherein communicating the intermediate output via the first message queue comprises publishing, by each computing device of the plurality of computing devices other than the last computing device, the intermediate output of the respective subset of the plurality of computational nodes to the first publisher-subscriber model queue, wherein the next computing device of the plurality of computing devices comprises a subscriber to the first publisher-subscriber model queue, and wherein communicating the gradient value via the second message queue comprises publishing, by each computing device of the plurality of computing devices other than the first computing device, the gradient value associated with the loss associated with the respective subset of the plurality of computational nodes to the second publisher-subscriber model queue, wherein the previous computing device of the plurality of computing devices comprises a subscriber to the second publisher-subscriber model queue.
Step 2A Prong Two Analysis: With respect to the abstract idea, the judicial exception is not integrated into a practical application.
The claim recites additional element(s) – subscriber. The additional element(s) is/are recited at a high-level of generality (i.e., as generic computer components performing generic computer functions of executing instructions on the computers) such that it amounts to no more than mere instructions to apply the exception using generic computer components (MPEP 2106.05(b)).
The claim recites wherein communicating the intermediate output via the first message queue comprises publishing, by each computing device of the plurality of computing devices other than the last computing device, the intermediate output of the respective subset of the plurality of computational nodes to the first publisher-subscriber model queue ...; wherein communicating the gradient value via the second message queue comprises publishing, by each computing device of the plurality of computing devices other than the first computing device, the gradient value associated with the loss associated with the respective subset of the plurality of computational nodes to the second publisher-subscriber model queue ..., which is simply storing and/or transmitting data recited at a high level of generality. This is nothing more than insignificant extra-solution activity (MPEP 2106.05(g)).
The claim recites ... wherein the next computing device of the plurality of computing devices comprises a subscriber to the first publisher-subscriber model queue; ... wherein the previous computing device of the plurality of computing devices comprises a subscriber to the second publisher-subscriber model queue which is simply additional information regarding the computer components, and the element(s) do(es) not apply the exception in a meaningful way (MPEP 2106.05(e)).
Accordingly, the additional element(s) do(es) not integrate the abstract idea into a practical application because the additional element(s) do(es) not impose any meaningful limits on practicing the abstract idea, and, therefore, the claim is directed to an abstract idea.
Step 2B Analysis: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to the integration of the abstract idea into a practical application, the additional element(s) of:
subscriber amount(s) to no more than mere instructions to apply the exception using generic computer components (MPEP 2106.05(b))
storing and/or transmitting data amount(s) to no more than insignificant extra-solution activity (MPEP 2106.05(g)), wherein the insignificant extra-solution activity is the well-understood routine and conventional activit(y/ies) of receiving or transmitting data over a network and/or storing and retrieving information in memory (MPEP 2016.05(d))
additional information regarding the computer components do(es) not apply the exception in a meaningful way (MPEP 2106.05(e))
The additional element(s) do(es) not provide an inventive concept, and, therefore, the claim is not patent eligible.
Regarding claim 18, the claim is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1 Analysis: Claim 18 is directed to a system with computing devices, which is directed to a machine, one of the statutory categories.
Step 2A Prong One Analysis: The claim recites a(n) system.
The limitation of ... wherein determining the loss comprises determining the loss for each sample of the plurality of samples based on the output and the label, as drafted, is a process that, under its broadest reasonable interpretation, covers a mathematical concept. The limitation encompasses calculating a loss.
If a claim limitation, under its broadest reasonable interpretation, covers performance of mathematical concepts, then it falls within the "Mathematical Concepts" grouping. Accordingly, the claim recites an abstract idea.
Step 2A Prong Two Analysis: With respect to the abstract idea, the judicial exception is not integrated into a practical application.
The claim recites additional element(s) – database. The additional element(s) is/are recited at a high-level of generality (i.e., as generic computer components performing generic computer functions of executing instructions on the computers) such that it amounts to no more than mere instructions to apply the exception using generic computer components (MPEP 2106.05(b)).
The claim recites a database configured to store a label for each sample of the plurality of samples ..., which is simply storing data recited at a high level of generality. This is nothing more than insignificant extra-solution activity (MPEP 2106.05(g)).
Accordingly, the additional element(s) do(es) not integrate the abstract idea into a practical application because the additional element(s) do(es) not impose any meaningful limits on practicing the abstract idea, and, therefore, the claim is directed to an abstract idea.
Step 2B Analysis: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to the integration of the abstract idea into a practical application, the additional element(s) of:
database amount(s) to no more than mere instructions to apply the exception using generic computer components (MPEP 2106.05(b))
storing data amount(s) to no more than insignificant extra-solution activity (MPEP 2106.05(g)), wherein the insignificant extra-solution activity is the well-understood routine and conventional activit(y/ies) of storing and retrieving information in memory (MPEP 2016.05(d))
The additional element(s) do(es) not provide an inventive concept, and, therefore, the claim is not patent eligible.
Regarding claim 19, the claim is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1 Analysis: Claim 19 is directed to a system with computing devices, which is directed to a machine, one of the statutory categories.
Step 2A Prong One Analysis: The claim recites a(n) system.
The limitation of determine a variance in the at least one parameter for each computational node, as drafted, is a process that, under its broadest reasonable interpretation, covers a mathematical concept. The limitation encompasses calculating a variance.
The limitation of generate at least one new computational node on at least one computing device of the plurality of computing devices based on the variance of at least one computational node of the respective subset of the plurality of computational nodes, as drafted, is a process that, under its broadest reasonable interpretation, covers a mental process. The limitation is directed to observation, evaluation, judgment and opinion and is a process capable of being performed by a human mentally or using pen and paper.
If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind, then it falls within the "Mental Processes" grouping. If a claim limitation, under its broadest reasonable interpretation, covers performance of mathematical concepts, then it falls within the "Mathematical Concepts" grouping. Accordingly, the claim recites an abstract idea.
Step 2A Prong Two Analysis: With respect to the abstract idea, the judicial exception is not integrated
into a practical application. The claim does not recite any additional elements which integrate the
abstract idea into a practical application and, therefore, does not impose any meaningful limits on
practicing the abstract idea. Therefore, the claim is directed to an abstract idea.
Step 2B Analysis: The claim does not include additional elements that are sufficient to amount to
significantly more than the judicial exception. As discussed above with respect to the integration of the
abstract idea into a practical application, the claim does not recite any additional elements which
provide an inventive concept, and, therefore, the claim is not patent eligible.
Regarding claim 20, the claim is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1 Analysis: Claim 20 is directed to a(n) computer program product, which is directed to an article of manufacture, one of the statutory categories.
Step 2A Prong One Analysis: The claim recites a(n) computer program product.
The limitation of initialize a respective subset of a plurality of computational nodes of a distributed machine learning model, each computational node comprising at least one parameter, as drafted, is a process that, under its broadest reasonable interpretation, covers a mental process. The limitation is directed to observation, evaluation, judgment and opinion and is a process capable of being performed by a human mentally or using pen and paper.
The limitation of determine a loss for each sample of the plurality of samples based on the output, as drafted, is a process that, under its broadest reasonable interpretation, covers a mathematical concept. The limitation encompasses calculating a loss.
The limitation of asynchronously update the at least one parameter of each computational node based on the gradient value associated with the loss for each sample as the gradient value associated with the loss for each sample is backward propagated while at least one of the plurality of samples is forward propagating through the distributed machine learning model, as drafted, is a process that, under its broadest reasonable interpretation, covers a mathematical concept. The limitation encompasses calculating parameter changes.
If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind, then it falls within the "Mental Processes" grouping. If a claim limitation, under its broadest reasonable interpretation, covers performance of mathematical concepts, then it falls within the "Mathematical Concepts" grouping. Accordingly, the claim recites an abstract idea.
Step 2A Prong Two Analysis: With respect to the abstract idea, the judicial exception is not integrated into a practical application.
The claim recites additional element(s) – computer program product, at least one computer-readable medium, one or more instructions, at least one processor, plurality of computing devices, first message queue, second message queue. The additional element(s) is/are recited at a high-level of generality (i.e., as generic computer components performing generic computer functions of executing instructions on the computers) such that it amounts to no more than mere instructions to apply the exception using generic computer components (MPEP 2106.05(b)).
The claim recites additional element(s) – distributed machine learning model. The additional element(s) is/are recited at a high-level of generality such that it amounts to no more than indicating a field of use or technological environment in which to apply the judicial exception (MPEP 2106.05(h)).
The claim recites forward propagate each sample of the plurality of samples through the respective subset of the plurality of computational nodes of the distributed machine learning model to generate an intermediate output for each sample of the plurality of samples ... which is simply applying the model recited at a high level of generality and amounts to the recitation of the words “apply it” (or an equivalent) or amounts to no more than mere instructions to implement an abstract idea or other exception on a computer (MPEP 2106.05(f)).
The claim recites receive training data associated with a plurality of samples; ... wherein forward propagating comprises communicating, from each computing device of the plurality of computing devices other than a last computing device, an intermediate output of the respective subset of the plurality of computational nodes to a next computing device of the plurality of computing devices via a first message queue; backward propagate the loss for each sample of the plurality of samples to each computing device of the plurality of computing devices, wherein backward propagating comprises communicating, from each computing device of the plurality of computing devices other than a first computing device, a gradient value associated with the loss associated with the respective subset of the plurality of computational nodes to a previous computing device of the plurality of computing devices via a second message queue; store the at least one parameter of each computational node of the respective subset of the plurality of computational nodes as updated; communicate, to at least one other computing device, data associated with the at least one parameter of each computational node of the respective subset of the plurality of computational nodes as updated to cause the at least one other computing device to store the at least one parameter of each computational node as updated, which is simply acquiring data, transmitting data and storing data recited at a high level of generality. This is nothing more than insignificant extra-solution activity (MPEP 2106.05(g)).
Accordingly, the additional element(s) do(es) not integrate the abstract idea into a practical application because the additional element(s) do(es) not impose any meaningful limits on practicing the abstract idea, and, therefore, the claim is directed to an abstract idea.
Step 2B Analysis: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to the integration of the abstract idea into a practical application, the additional element(s) of:
computer program product, at least one computer-readable medium, one or more instructions, at least one processor, plurality of computing devices, first message queue, second message queue amount(s) to no more than mere instructions to apply the exception using generic computer components (MPEP 2106.05(b))
applying the model amount(s) to no more than mere instructions to apply the exception (MPEP 2106.05(f))
acquiring data, transmitting data and storing data amount(s) to no more than insignificant extra-solution activity (MPEP 2106.05(g)), wherein the insignificant extra-solution activity is the well-understood routine and conventional activit(y/ies) of receiving or transmitting data over a network and/or storing and retrieving information in memory (MPEP 2016.05(d))
distributed machine learning model amount(s) to no more than indicating a field of use or technological environment in which to apply the judicial exception (MPEP 2106.05(h))
The additional element(s) do(es) not provide an inventive concept, and, therefore, the claim is not patent eligible.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 1-9, 12-18, 20 are rejected under 35 U.S.C. 103 as being unpatentable over Tomioka et al. (US 2018/0336458 A1 – Asynchronous Neural Network, hereinafter referred to as “Tomioka”) in view of Yancey et al. (US 2008/0120260 A1 – Reconfigurable Neural Network Systems and Methods Using FPGAs Having Packet Routers, hereinafter referred to as “Yancey”).
Regarding claim 1, Tomioka teaches a computer-implemented method (Tomioka, [0058] – teaches computer implemented), comprising:
initializing a distributed machine learning model on a plurality of computing devices, the distributed machine learning model comprising a plurality of computational nodes, each computing device of the plurality of computing devices comprising a respective subset of the plurality of computational nodes, each computational node comprising at least one parameter (Tomioka, [0026] - teaches a distributed neural network comprising a plurality of worker nodes [computing devices], wherein each worker node comprises a layer of the neural network and the associated parameters for that layer; see also Tomioka, [0022]; Tomioka, [0025]);
receiving training data associated with a plurality of samples at a first computing device of the plurality of computing devices (Tomioka, [0035]-[0036] - teaches a first worker node receiving a data instance; see also Tomioka, [0022]-[0023]);
forward propagating each sample of the plurality of samples through the distributed machine learning model to generate an output for each sample of the plurality of samples (Tomioka, [0036] – teaches forward propagating the data instances through the worker nodes of the distributed machine learning model to generate an output at the final worker node; see also Tomioka, Fig. 4), wherein forward propagating comprises communicating, from each computing device of the plurality of computing devices other than a last computing device, an intermediate output of the respective subset of the plurality of computational nodes to a next computing device of the plurality of computing devices (Tomioka, [0036] - teaches each worker node passing its intermediate result to the next worker node using messages during forward propagation; see also Tomioka, Fig. 4) via a first message queue (Tomioka, [0047]-[0050] - teaches message passing using message queues for communicating data both forward and backward; see also Tomioka, [0033]);
determining a loss for each sample of the plurality of samples based on the output (Tomioka, [0036] – teaches determining a loss for a training data sample);
backward propagating the loss for each sample of the plurality of samples to each computing device of the plurality of computing devices (Tomioka, [0038]-[0041] - teaches determining gradients based on the loss and back propagating the gradients through the worker nodes; see also Tomioka, Fig. 5), wherein backward propagating comprises communicating, from each computing device of the plurality of computing devices other than the first computing device, a gradient value associated with the loss associated with the respective subset of the plurality of computational nodes to a previous computing device of the plurality of computing devices (Tomioka, [0041] - teaches using messages to back propagate gradients to each worker node; see also Tomioka, Fig. 5) via a second message queue (Tomioka, [0047]-[0050] - teaches message passing using message queues for communicating data both forward and backward; see also Tomioka, [0033]);
asynchronously updating the at least one parameter of each computational node based on the loss for each sample as the loss for each sample is backward propagated while at least one of the plurality of samples is forward propagating through the distributed machine learning model (Tomioka, [0035] – teaches performing the forward and backward propagation at the same time; see also Tomioka, Fig. 9, [0053]);
storing, at each computing device of the plurality of computing devices, the at least one parameter of each computational node of the respective subset of the plurality of computational nodes as updated (Tomioka, [0036] – teaches storing the parameters; Tomioka, [0038] – teaches storing updated parameter values of the computational node; Tomioka, [0041] – storing updated parameters).
While Tomioka teaches that all of the worker nodes are connected, propagating data through all of the worker nodes and sending data from one worker node to the previous or subsequent worker node based on propagation direction, Tomioka does not explicitly teach sending data from one worker node to all of the other worker nodes.
Yancey teaches communicating, from each computing device of the plurality of computing devices to all other computing devices of the plurality of computing devices (Yancey, [0023] – teaches data packets being sent to all other nodes; see also Yancey, Figs. 1, 4), data associated with the at least one parameter of each computational node of the respective subset of the plurality of computational nodes as updated, each of the other computing devices of the plurality of computing devices storing the at least one parameter of each computational node as updated (Yancey, [0024] – teaches that the data send to each of the other devices can be weights [parameters] of the given device for the neurons of that device).
It would have been obvious to one of ordinary skill in the art before the filing date of the claimed invention to modify Tomioka with the teachings of Yancey in order to generate a more intelligent and dynamic operation of distributed models in the field of training distributed machine learning models (Yancey, [0005] – “Systems and methods are disclosed for reconfigurable neural networks that utilize interconnected FPGAs. Each FPGA has a packet router that provides communication connections to neural network nodes internal to the FPGA and neural network nodes in other external FPGAs. In this configuration, a control module in one FPGA can be configured to act as a neural network controller to provide control of neuron interconnections, weight values, and other control data to the other network nodes in order to control the operation of the neural network as a whole. In addition, the FPGAs can be connected to each other using high-speed interconnects, such as high-speed serial digital connections. By utilizing packet routing and high speed interconnects, a highly reconfigurable and advantageous neural network can be formed that allows for rapid reconfiguration and dynamic operation in the tasks the neural network is programmed to perform. In this way, more intelligent and dynamic operation of a neural network can be achieved than was previously possible. As described below, various embodiments and features can be implemented, as desired, and related systems and methods can be utilized, as well.”).
Regarding claim 2, Tomioka in view of Yancey teaches all of the limitations of the method of claim 1 as noted above. Tomioka further teaches wherein the distributed machine learning model comprises a deep neural network and each computational node of the plurality of computational nodes comprises a neuron (Tomioka, [0021], [0026], [0034], [0036] – teaches a deep neural network across a plurality of worker nodes where each computational node is a neural; see also Tomioka, Fig. 3).
It would have been obvious to one of ordinary skill in the art before the filing data of the claimed invention to combine the teaching of Tomioka and Yancey for the same reasons as disclosed in claim 1 above.
Regarding claim 3, Tomioka in view of Yancey teaches all of the limitations of the method of claim 1 as noted above. Tomioka further teaches wherein the plurality of computational nodes are divided into a plurality of layers, each subset of computational nodes being associated with at least one layer of the plurality of layers, each layer being implemented on at least one of the plurality of computing devices (Tomioka, [0026] – teaches the neural network being divided into layers, where each worker node comprises a network layer; see also Tomioka, Fig. 3).
It would have been obvious to one of ordinary skill in the art before the filing data of the claimed invention to combine the teaching of Tomioka and Yancey for the same reasons as disclosed in claim 1 above.
Regarding claim 4, Tomioka in view of Yancey teaches all of the limitations of the method of claim 1 as noted above. Tomioka further teaches wherein the at least one parameter of each computational node comprises at least one of a weight parameter, a bias parameter, or any combination thereof (Tomioka, [0022]-[0023] – teaches weight parameters).
It would have been obvious to one of ordinary skill in the art before the filing data of the claimed invention to combine the teaching of Tomioka and Yancey for the same reasons as disclosed in claim 1 above.
Regarding claim 5, Tomioka in view of Yancey teaches all of the limitations of the method of claim 1 as noted above. Tomioka further teaches wherein storing the at least one parameter of each computational node comprises storing the at least one parameter of each computational node in at least one cache memory of each computing device of the plurality of computing devices (Tomioka, [0029], [0034], [0041] – teaches storing parameters in cache memory; Tomioka, [0059] – teaches cache memory).
It would have been obvious to one of ordinary skill in the art before the filing data of the claimed invention to combine the teaching of Tomioka and Yancey for the same reasons as disclosed in claim 1 above.
Regarding claim 6, Tomioka in view of Yancey teaches all of the limitations of the method of claim 5 as noted above. Tomioka further teaches storing, in a backup storage, the at least one parameter of each computational node of the plurality of computational nodes (Tomioka, [0025] – teaches storing the parameters of the trained neural network until the network is transferred to a device for deployment).
It would have been obvious to one of ordinary skill in the art before the filing data of the claimed invention to combine the teaching of Tomioka and Yancey for the same reasons as disclosed in claim 5 above.
Regarding claim 7, Tomioka in view of Yancey teaches all of the limitations of the method of claim 1 as noted above. Tomioka further teaches wherein the first message queue comprises a first publisher-subscriber model queue, and wherein the second message queue comprises a second publisher-subscriber model queue (Tomioka, [0047]-[0050] - teaches message passing using message queues for communicating data both forward and backward; see also Tomioka, [0033]).
It would have been obvious to one of ordinary skill in the art before the filing data of the claimed invention to combine the teaching of Tomioka and Yancey for the same reasons as disclosed in claim 1 above.
Regarding claim 8, Tomioka in view of Yancey teaches all of the limitations of the method of claim 7 as noted above. Tomioka further teaches
wherein communicating the intermediate output via the first message queue comprises publishing, by each computing device of the plurality of computing devices other than the last computing device, the intermediate output of the respective subset of the plurality of computational nodes to the first publisher-subscriber model queue, wherein the next computing device of the plurality of computing devices comprises a subscriber to the first publisher-subscriber model queue (Tomioka, [0047]-[0050] - teaches message passing using message queues for communicating data both forward and backward; see also Tomioka, [0033]), and
wherein communicating the gradient value via the second message queue comprises publishing, by each computing device of the plurality of computing devices other than the first computing device, the gradient value associated with the loss associated with the respective subset of the plurality of computational nodes to the second publisher-subscriber model queue, wherein the previous computing device of the plurality of computing devices comprises a subscriber to the second publisher-subscriber model queue (Tomioka, [0047]-[0050] - teaches message passing using message queues for communicating data both forward and backward; see also Tomioka, [0033]).
It would have been obvious to one of ordinary skill in the art before the filing data of the claimed invention to combine the teaching of Tomioka and Yancey for the same reasons as disclosed in claim 7 above.
Regarding claim 9, Tomioka in view of Yancey teaches all of the limitations of the method of claim 1 as noted above. Tomioka further teaches storing, in a database, a label for each sample of the plurality of samples, wherein determining the loss comprises determining the loss for each sample of the plurality of samples based on the output and the label (Tomioka, [0022] – teaches training data comprising labels used for calculating loss during backpropagation training; Tomioka, [0033] – teaches storing the training data to use during training; see also Tomioka, [0036]).
It would have been obvious to one of ordinary skill in the art before the filing data of the claimed invention to combine the teaching of Tomioka and Yancey for the same reasons as disclosed in claim 1 above.
Regarding claim 12, it is the system embodiment of claim 1 with similar limitations to claim 1 and is rejected using the same reasoning found in claim 1. Tomioka further teaches a system, comprising:
a first message queue (Tomioka, [0047]-[0050] - teaches message passing using message queues for communicating data both forward and backward; see also Tomioka, [0033]);
a second message queue (Tomioka, [0047]-[0050] - teaches message passing using message queues for communicating data both forward and backward; see also Tomioka, [0033]);
a plurality of computing devices, each computing device of the plurality of computing devices comprising a respective subset of a plurality of computational nodes of a distributed machine learning model, each computational node comprising at least one parameter (Tomioka, [0026] - teaches a distributed neural network comprising a plurality of worker nodes [computing devices], wherein each worker node comprises a layer of the neural network and the associated parameters for that layer; see also Tomioka, [0022]; Tomioka, [0025]), each computing device comprising at least one processor and at least one non-transitory computer-readable medium including one or more instructions that, when executed by the at least one processor, cause the at least one processor to (Tomioka, [0058] – teaches computer implemented) …
It would have been obvious to one of ordinary skill in the art before the filing data of the claimed invention to combine the teaching of Tomioka and Yancey for the same reasons as disclosed in claim 1 above.
Regarding claim 13, Tomioka in view of Yancey teaches all of the limitations of the system of claim 12 as noted above. Tomioka further teaches
wherein the distributed machine learning model comprises a deep neural network and each computational node of the plurality of computational nodes comprises a neuron (Tomioka, [0021], [0026], [0034], [0036] – teaches a deep neural network across a plurality of worker nodes where each computational node is a neural; see also Tomioka, Fig. 3), and
wherein the plurality of computational nodes are divided into a plurality of layers, each subset of computational nodes being associated with at least one layer of the plurality of layers, each layer being implemented on at least one of the plurality of computing devices (Tomioka, [0026] – teaches the neural network being divided into layers, where each worker node comprises a network layer; see also Tomioka, Fig. 3).
It would have been obvious to one of ordinary skill in the art before the filing data of the claimed invention to combine the teaching of Tomioka and Yancey for the same reasons as disclosed in claim 12 above.
Regarding claim 14, the rejection of claim 12 is incorporated herein. Further, the limitations in this claim are taught by Tomioka in view of Yancey for the reasons set forth in the rejection of claim 5.
Regarding claim 15, the rejection of claim 14 is incorporated herein. Further, the limitations in this claim are taught by Tomioka in view of Yancey for the reasons set forth in the rejection of claim 6.
Regarding claim 16, the rejection of claim 12 is incorporated herein. Further, the limitations in this claim are taught by Tomioka in view of Yancey for the reasons set forth in the rejection of claim 7.
Regarding claim 17, the rejection of claim 16 is incorporated herein. Further, the limitations in this claim are taught by Tomioka in view of Yancey for the reasons set forth in the rejection of claim 8.
Regarding claim 18, the rejection of claim 12 is incorporated herein. Further, the limitations in this claim are taught by Tomioka in view of Yancey for the reasons set forth in the rejection of claim 9.
Regarding claim 20, Tomioka teaches a computer program product comprising at least one non-transitory computer-readable medium including one or more instructions that, when executed by at least one processor, cause the at least one processor (Tomioka, [0058] – teaches computer implemented) to:
initialize a respective subset of a plurality of computational nodes of a distributed machine learning model, each computational node comprising at least one parameter (Tomioka, [0026] – teaches the neural network being divided into layers, where each worker node comprises a network layer; see also Tomioka, Fig. 3);
receive training data associated with a plurality of samples (Tomioka, [0035]-[0036] - teaches a first worker node receiving a data instance; see also Tomioka, [0022]-[0023]);
forward propagate each sample of the plurality of samples through the respective subset of the plurality of computational nodes of the distributed machine learning model to generate an intermediate output for each sample of the plurality of samples (Tomioka, [0036] – teaches forward propagating the data instances through the worker nodes of the distributed machine learning model to generate an output at the final worker node; see also Tomioka, Fig. 4), wherein forward propagating comprises communicating, from each computing device of the plurality of computing devices other than a last computing device, an intermediate output of the respective subset of the plurality of computational nodes to a next computing device of the plurality of computing devices (Tomioka, [0036] - teaches each worker node passing its intermediate result to the next worker node using messages during forward propagation; see also Tomioka, Fig. 4) via a first message queue (Tomioka, [0047]-[0050] - teaches message passing using message queues for communicating data both forward and backward; see also Tomioka, [0033]);
determine a loss for each sample of the plurality of samples based on the output (Tomioka, [0036] – teaches determining a loss for a training data sample);
backward propagate the loss for each sample of the plurality of samples to each computing device of the plurality of computing devices (Tomioka, [0038]-[0041] - teaches determining gradients based on the loss and back propagating the gradients through the worker nodes; see also Tomioka, Fig. 5), wherein backward propagating comprises communicating, from each computing device of the plurality of computing devices other than a first computing device, a gradient value associated with the loss associated with the respective subset of the plurality of computational nodes to a previous computing device of the plurality of computing devices (Tomioka, [0041] - teaches using messages to back propagate gradients to each worker node; see also Tomioka, Fig. 5) via a second message queue (Tomioka, [0047]-[0050] - teaches message passing using message queues for communicating data both forward and backward; see also Tomioka, [0033]);
asynchronously update the at least one parameter of each computational node based on the gradient value associated with the loss for each sample as the gradient value associated with the loss for each sample is backward propagated while at least one of the plurality of samples is forward propagating through the distributed machine learning model (Tomioka, [0035] – teaches performing the forward and backward propagation at the same time; see also Tomioka, Fig. 9, [0053]);
store the at least one parameter of each computational node of the respective subset of the plurality of computational nodes as updated (Tomioka, [0036] – teaches storing the parameters; Tomioka, [0038] – teaches storing updated parameter values of the computational node; Tomioka, [0041] – storing updated parameters).
While Tomioka teaches that all of the worker nodes are connected, propagating data through all of the worker nodes and sending data from one worker node to the previous or subsequent worker node based on propagation direction, Tomioka does not explicitly teach sending data from one worker node to all of the other worker nodes.
Yancey teaches communicate, to at least one other computing device, data associated with the at least one parameter of each computational node of the respective subset of the plurality of computational nodes as updated to cause the at least one other computing device to store the at least one parameter of each computational node as updated (Yancey, [0023] – teaches data packets being sent to all other nodes; Yancey, [0024] – teaches that the data send to each of the other devices can be weights [parameters] of the given device for the neurons of that device; see also Yancey, Figs. 1, 4).
It would have been obvious to one of ordinary skill in the art before the filing date of the claimed invention to modify Tomioka with the teachings of Yancey in order to generate a more intelligent and dynamic operation of distributed models in the field of training distributed machine learning models (Yancey, [0005] – “Systems and methods are disclosed for reconfigurable neural networks that utilize interconnected FPGAs. Each FPGA has a packet router that provides communication connections to neural network nodes internal to the FPGA and neural network nodes in other external FPGAs. In this configuration, a control module in one FPGA can be configured to act as a neural network controller to provide control of neuron interconnections, weight values, and other control data to the other network nodes in order to control the operation of the neural network as a whole. In addition, the FPGAs can be connected to each other using high-speed interconnects, such as high-speed serial digital connections. By utilizing packet routing and high speed interconnects, a highly reconfigurable and advantageous neural network can be formed that allows for rapid reconfiguration and dynamic operation in the tasks the neural network is programmed to perform. In this way, more intelligent and dynamic operation of a neural network can be achieved than was previously possible. As described below, various embodiments and features can be implemented, as desired, and related systems and methods can be utilized, as well.”).
Claims 10, 19 are rejected under 35 U.S.C. 103 as being unpatentable over Tomioka in view of Yancey and further in view of Srinivasan et al. (US 2020/0117992 A1 – Distributed Training of Reinforcement Learning Systems, hereinafter referred to as “Srinivasan”).
Regarding claim 10, Tomioka in view of Yancey teaches all of the limitations of the method of claim 1 as noted above. However, Tomioka in view of Yancey does not explicitly teach determining a variance in the at least one parameter for each computational node; and generating at least one new computational node on at least one computing device of the plurality of computing devices based on the variance of at least one computational node of the respective subset of the plurality of computational nodes.
Srinivasan teaches
determining a variance in the at least one parameter for each computational node (Srinivasan, [0065] – teaches determining the variance for a parameter of a node of a learner); and
generating at least one new computational node on at least one computing device of the plurality of computing devices based on the variance of at least one computational node of the respective subset of the plurality of computational nodes (Srinivasan, [0070]-[0072] – teaches modifying the number of learners or actors, including adding learners, based on the variance in the parameters).
It would have been obvious to one of ordinary skill in the art before the filing date of the claimed invention to modify Tomioka in view of Yancey with the teachings of Srinivasan in order to perform faster training while improving performance in the field of training distributed machine learning models (Srinivasan, [0007] – “The subject matter described in this specification can be implemented in particular embodiments so as to realize one or more of the following advantages. By parallelizing training, a reinforcement learning system can be trained faster. Additionally, a reinforcement learning system trained using the distributed learning techniques described in this specification can, after training, have an improved performance on reinforcement learning tasks than the same reinforcement learning system trained using a non-distributed reinforcement learning training technique. By providing an architecture that allows a distributed reinforcement learning training system to include arbitrary numbers of learners, actors, and replay memories, the system can easily be adapted for training a system to perform various reinforcement learning tasks. Additionally, the numbers of learners, actors, and, optionally, replay memories can easily be adjusted during training, resulting in improved performance.”).
Regarding claim 19, the rejection of claim 12 is incorporated herein. Further, the limitations in this claim are taught by Tomioka in view of Yancey and further in view of Srinivasan for the reasons set forth in the rejection of claim 10.
Conclusion
Any inquiry concerning this communication or earlier communication from the examiner should be directed to MARSHALL WERNER whose telephone number is (469) 295-9143. The examiner can normally be reached on Monday – Thursday 7:30 AM – 4:30 PM ET.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kamran Afshar, can be reached at (571) 272-7796. The fax number for the organization where this application or proceeding is assigned is (571) 273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/MARSHALL L WERNER/ Primary Examiner, Art Unit 2125