Prosecution Insights
Last updated: August 14, 2026
Application No. 17/752,044

DATA PARALLELISM IN DISTRIBUTED TRAINING OF ARTIFICIAL INTELLIGENCE MODELS

Final Rejection §101§103§DOUBLEPATENT
Filed
May 24, 2022
Priority
Jul 15, 2019 — provisional 62/874,462 +3 more
Examiner
BRACERO, ANDREW ANGEL
Art Unit
2126
Tech Center
2100 — Computer Architecture & Software
Assignee
Microsoft Technology Licensing, LLC
OA Round
2 (Final)
90%
Grant Probability
Favorable
3-4
OA Rounds
4m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 90% — above average
90%
Career Allowance Rate
9 granted / 10 resolved
+35.0% vs TC avg
Strong +33% interview lift
Without
With
+33.3%
Interview Lift
resolved cases with interview
Typical timeline
4y 7m
Avg Prosecution
13 currently pending
Career history
33
Total Applications
across all art units

Statute-Specific Performance

§101
34.5%
-5.5% vs TC avg
§103
45.8%
+5.8% vs TC avg
§102
9.2%
-30.8% vs TC avg
§112
9.2%
-30.8% vs TC avg
Black line = Tech Center average estimate • Based on career data from 10 resolved cases

Office Action

§101 §103 §DOUBLEPATENT
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . DETAILED ACTION Claims 1-20 are presented for examination in this application, 17/752,044, filed 2022-05-24 having an effective filing date of provisional application, 62/874,462, of 2019-07-15. The Examiner cites particular sections in the references as applied to the claims below for the convenience of the applicant(s). Although the specified citations are representative of the teachings in the art and are applied to the specific limitations within the individual claim, other passages and figures may apply as well. It is respectfully requested that, in preparing responses, the applicant(s) fully consider the references in their entirety as potentially teaching all or part of the claimed invention, as well as the context of the passage as taught by the prior art or disclosed by the Examiner. Information Disclosure Statement Acknowledgement is made of the information disclosure statements filed on 2022-10-19, 2023-09-07, 2025-02-18, 2025-04-22, 2025-05-13, and 2025-08-25. Response to Arguments Applicant’s arguments and remarks filed 02/18/2026 have been fully considered. The arguments and remarks regarding the 35 U.S.C 101 and 35 U.S.C 103 rejections were not all found to be persuasive. The 35 U.S.C 101 and 35 U.S.C 103 rejections have been maintained. 35 U.S.C 101 Applicant’s response: Applicant asserts “the claimed embodiments are directed towards an improvement to computer-related technology to computer functionality versus being direct to an abstract idea. In particular, the present claims are directed towards improving the accuracy and/or relevancy of responses provided by a machine learning model without the need to retrain the machine learning model. Performing these operations contemporaneously improves the efficiency of the distributed execution of the model by hiding the communication latency. Id. Therefore, like ENFISH, the claimed embodiments are directed to an improvement to computer-related technology and are therefore not directed to an abstract idea.” Applicant further asserts “While Applicant respectfully asserts that the pending claims are not directed to abstract ideas, the pending claims also recite a "practical application" of any alleged abstract ideas. … As discussed above, some operations in the current invention, such as the updating of weights based on results from a second subportion of the model and the transmission of weights for a third subportion of the model, are performed contemporaneously with the execution of a first subportion of the model at the target device. See, e.g., [0040], [0056], [0058], [0063], [0066], [0067], and/or [0090]. Performing these operations contemporaneously improves the efficiency of the distributed execution of the model by hiding the communication latency. Id. As such, the pending claims are patent eligible because they are not directed towards abstract ideas, and, even assuming, arguendo, the claimed invention were directed towards abstract ideas, the claimed invention includes an inventive concept that provides a patent-eligible application of any alleged abstract idea. Like ENFISH, the claimed invention is directed towards a specific improvement to computer technology and is therefore not directed towards abstract ideas. Furthermore, similar DDR Holdings, the claimed invention is directed to a technical solution to a technical problem. Accordingly, Applicant respectfully requests that the rejection of claims 1-20 under 35 U.S.C. § 101 be reconsidered and withdrawn”. Examiner’s response: Examiner respectfully disagrees. The Examiner notes that even though the claims are examined with the broadest reasonable interpretation in light of the specification, the current claims do not recite additional elements that make the claims eligible. A consideration when determining whether a claim integrates the judicial exception into a practical application is whether the additional elements amount to adding insignificant extra-solution activities to the judicial exception, of which the courts have identified as not integrating a judicial exception into a practical application. As amended, the claims recite “a transmitter configured to transmit a portion of an artificial intelligence (AI) model to the target device, the device comprising an integrated circuit chip having an on-chip memory of a size less than an entirety of the AI model, and a size of the portion being based at least on the on-chip memory size and a size of one or more layers of the AI model” … “wherein the transmitter is further configured to send weights for a third subportion of the transmitted portion of the AI model to the target device”. This limitation amounts to the idea of outputting data based on types of information and availability of information, for collection, analysis, and/or display (Electric Power Group, LLC v Alstom S.A, 830 F.3d 1350, 1354-55, 119 USPQ2d 1739, 1742 (Fed. Cir. 2016)) as recited in MPEP 2106.05(g). Furthermore, under step 2B, a consideration for determining whether a claim recites significantly more than a judicial exception is whether the additional element(s) are well0understood, routine, conventional activities previously known to the industry. The aforementioned independent claim limitations of claims 1, 8, and 15 recite well-understood, routine, conventional functions claims as insignificant extra-solution activities that amount to merely receiving or transmitting data over a network, e.g., using the Internet to gather data (buySafe, Inc. v. Google, Inc., 765 F.3d 1350, 1335, 112 USPQ2d 1093, 1096 (Fed. Cir. 2014) computer receives and sends information over a network)). Additionally, the Examiner notes “if applicant amends a claim to add a generic computer or generic computer components and asserts that the claim recites significantly more because the generic computer is ‘specially programmed’ (as in Alappat, now considered superseded) or is a ‘particular machine’ (as in Bilski), the examiner should look at whether the added elements integrate the exception into a practical application or provide significantly more than the judicial exception. Merely adding a generic computer, generic computer components, or a programmed computer to perform generic computer functions does not automatically overcome an eligibility rejection. Alice Corp. Pty. Ltd. v. CLS Bank Int’l, 573 U.S. 208, 223-24, 110 USPQ2d 1976, 1983-84 (2014). See In re Alappat, 33 F.3d 1526, 1545, 31 USPQ2d 1545, 1558 (Fed. Cir. 1994); In re Bilski, 545 F.3d 943, 88 USPQ2d 1385 (Fed. Cir. 2008)”. MPEP 2106.05(b).“It is important to note that a general purpose computer that applies a judicial exception, such as an abstract idea, by use of conventional computer functions does not qualify as a particular machine”. MPEP 2016.05(b)(I). A person having ordinary skill in the art would find {insert limitation here} to be considered a conventional computer function within the field of machine learning and artificial intelligence. The claim as a whole is still directed to abstract idea mental process.” In regard to the arguments “As discussed above, some operations in the current invention, such as the updating of weights based on results from a second subportion of the model and the transmission of weights for a third subportion of the model, are performed contemporaneously with the execution of a first subportion of the model at the target device. See, e.g., [0040], [0056], [0058], [0063], [0066], [0067], and/or [0090]. Performing these operations contemporaneously improves the efficiency of the distributed execution of the model by hiding the communication latency. Id”, the Examiner notes that the spec recites at [0040]: “To increase balance and efficiency, this technique iterates on the same layer across a large number of input samples until either (a) the next layer is loaded onto the target device thereby completely hiding its latency, or (b) the next layer is loaded after the current layer finishes, exposing its latency, but minimizing the overhead with a long computation cycle for the current layer.”. The Examiner notes that as claimed, the limitations of the independent claims do not reflect that the “technique iterates on the same layer across a large number of input samples until … the (a) the next layer is loaded onto the target device thereby completely hiding its latency or (b) the next layer is loaded after the current layer finishes, exposing its latency, but minimizing the overhead with a long computation cycle for the current layer” The claimed invention of independent claims 1, 8, and 15 does not recite sufficient details to reflect that these techniques and therefore is not patent eligible. 35 U.S.C 103: Applicant’s response: Applicant asserts “that Sridharan and Ambrose, taken individually or in combination, fail to teach or suggest a "weight updater configured to perform a reduction of parameters for a second subportion of the transmitted portion of the AI model contemporaneously with an execution of a set of microbatches of a training dataset by the integrated circuit chip at the target device on a first subportion of the transmitted portion of the AI model stored in the on-chip memory," and a "transmitter...configured to send weights for a third subportion of the transmitted portion of the AI model to the target device contemporaneously with the execution of the set of microbatches of the training dataset by the integrated circuit chip at the target device on the first subportion of the transmitted portion of the AI model stored in the on-chip memory," as recited in independent claim 1. Sridharan relates to a system for distributed training of a model. See, e.g., Abstract. Sridharan arguably discloses a central parameter server that performs parameter averaging to update parameters (e.g., weights) for the model, and distributing of the updated weights to nodes. See, e.g., [0190] and/or [0220]. However, Applicant respectfully asserts that Sridharan fails to teach or suggest performing such operations contemporaneously with the execution of a set of microbatches on a different subportion of a model at a node. In particular, Applicant respectfully asserts that Sridharan fails to teach or suggest performing a reduction of parameters associated with a second subportion or sending weights for a third subportion contemporaneously with the execution of a set of microbatches on a first subportion. Accordingly, Applicant respectfully asserts that Sridharan fails to teach or suggest "a weight updater configured to perform a reduction of parameters for a second subportion of the transmitted portion of the AI model contemporaneously with an execution of a set of microbatches of a training dataset by the integrated circuit chip at the target device on a first subportion of the transmitted portion of the AI model stored in the on-chip memory," and a "transmitter...configured to send weights for a third subportion of the transmitted portion of the AI model to the target device contemporaneously with the execution of the set of microbatches of the training dataset by the integrated circuit chip at the target device on the first subportion of the transmitted portion of the AI model stored in the on-chip memory.". Examiner’s response: Examiner respectfully disagrees. Examiner notes that in at least para [0189], Sridharan details model parallelism where different layers, or subportions, of a neural network can be trained on by different processing nodes. In at least para [0190], Sridharan details of the weight updating being done by data parallelism. In at least para [0202] Sridharan details data parallelism being implemented with mini-batches. Claim Interpretation The following is a quotation of 35 U.S.C. 112(f): (f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph: An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. This application includes one or more claim limitations that that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f), because the claim limitations use a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitations are: “a weight updater configured to perform a reduction of parameters for a second subportion of the transmitted portion of the AI model contemporaneously with an execution of a set of microbatches of a training dataset by the integrated circuit chip at the target device on a first subportion of the transmitted portion of the AI model stored in the on-chip memory” as recited in claim 1. The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked. As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph: (A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function; (B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and (C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function. Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function. Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function. Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. The following appears to be the closest portions of the specification corresponding to the 35 U.S.C 112(f) invocations: weight updater Fig. 1 shows a weight updater within the output data manager that is housed within an AI model Manager that is part of the memory within the parameter server. Fig. 12 shows a flowchart of the weight updating process. Specification at para [0133] recites “and the weight updater is further configured to receive gradients from the another target device to perform reduction of parameters for the another portion of the AI model. ”. Regarding claim 1 and the above-noted three-prong test, the recited weight updater is a generic placeholder, wherein configured to perform a reduction of parameters for a second subportion of the transmitted portion of the AI model contemporaneously with an execution of a set of microbatches of a training dataset by the integrated circuit chip at the target device on a first subportion of the transmitted portion of the AI model stored in the on-chip memory” as recited in claim 1. Because this claim limitation is being interpreted under 35 U.S.C 112(f), it is being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof. If applicant does not intend to have these limitations interpreted under 35 U.S.C 112(f), applicant may (1) amend the claim limitations to avoid them being interpreted under 35 U.S.C. 112(f) (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitations recite sufficient structure to perform the claimed function so as to avoid them being interpreted under 35 U.S.C. 112(f). Double Patenting The nonstatutory double patenting rejection is based on judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg,140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman,11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969). A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on nonstatutory double patenting provided the reference application or patent either is shown to be commonly owned with the examined application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. See MPEP § 717.02 for applications subject to examination under the first inventor to file provisions of the AIA as explained in MPEP § 2159. See MPEP § 2146 et seq. for applications not subject to examination under the first inventor to file provisions of the AIA . A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b). The filing of a terminal disclaimer by itself is not a complete reply to a nonstatutory double patenting (NSDP) rejection. A complete reply requires that the terminal disclaimer be accompanied by a reply requesting reconsideration of the prior Office action. Even where the NSDP rejection is provisional the reply must be complete. See MPEP § 804, subsection I.B.1. For a reply to a non-final Office action, see 37 CFR 1.111(a). For a reply to a final Office action, see 37 CFR 1.113(c). A request for reconsideration while not provided for in 37 CFR 1.113(c) may be filed after final for consideration. See MPEP § § 706.07(e) and 714.13. The USPTO Internet website contains terminal disclaimer forms which may be used. Please visit www.uspto.gov/patent/patents-forms. The actual filing date of the application in which the form is filed determines what form (e.g., PTO/SB/25, PTO/SB/26, PTO/AIA /25, or PTO/AIA /26) should be used. A web-based eTerminal Disclaimer may be filed out completely online using web-screens. An eTerminal Disclaimer may be filed out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto-processed and approved upon submission. For more information about eTerminal Disclaimers, refer to www.uspto.gov/patents/apply/applying-online/eterminal-disclaimer. Examiner notes that claims 1, 5-8, 12-15, 19 and 20 are rejected on the ground of nonstatutory double patenting, as indicated below. Claims 1, 5-8, 12-15, 19 and 20 are rejected on the ground of nonstatutory double patenting as being unpatentable over claims 1, 5-8, 12-15, 19 and 20 of U.S. Patent No. US11436019B2. Although the claims at issue are not identical, they are not patentably distinct from each other because claims 1, 5- 8, 12-15, 19 and 20 of the instant application are anticipated by the claims of the issued patent by being broader than the claims of the issued patent. See the comparison below: Instant Application U.S. Patent No. US11436019B2 1. A system, comprising: a parameter server communicatively connected to a target device, the parameter server comprises: a transmitter configured to transmit a portion of an artificial intelligence (AI) model to the target device, the target device comprising an integrated circuit chip having an on-chip memory of a size less than an entirety of the Al model, and a size of the portion being based at least on the on-chip memory size and a size of one or more layers of the Al model, and contemporaneously, with a set of microbatches of a training dataset being executed by the integrated circuit chip at the target device on a first subportion of the transmitted portion of the Al model stored in the on-chip memory to generate gradients, a weight updater is configured to perform a reduction of parameters for a second subportion of the transmitted portion of the Al model, and the transmitter is further configured to send weights for a third subportion of the transmitted portion of the Al model to the target device. 5. The system of claim 1 wherein a microbatch size is configurable based on a rate of executing the set of microbatches at the target device and a rate of communication between the target device and the parameter server. 6. The system of claim 1, wherein the parameter server further comprises a precision formatter configured to: convert weights for a fourth subportion of the transmitted portion of the Al model to a first precision format prior to sending the weights to the target device; convert gradients received from the target device to a second precision format; and update the weights using the converted gradients. 7. The system of claim 1, wherein the transmitter is further configured to transmit another portion of the Al model to another target device; and the weight updater is further configured to receive gradients from the another target device to perform reduction of parameters for the another portion of the Al model. 8. A method implemented in a parameter server, comprising: transmitting a portion of a stored artificial intelligence (AI) model to a target device the target device comprising an integrated circuit chip having an on-chip memory of a size less than an entirety of the Al model, and a size of the portion being based at least on the on-chip memory size and a size of one or more layers of the Al model and contemporaneously, with a set of microbatches of a determined size, of a training dataset being executed by the integrated circuit chip at the target device on a first subportion of the transmitted portion of the Al model stored in the on-chip memory to generate gradients, performing a reduction of parameters for a second subportion of the transmitted portion of the Al model and sending weights for a third subportion of the transmitted portion of the Al model to the target device. 12. The method of claim 8, wherein a microbatch size is configurable based on a rate of executing the set of microbatches at the target device and a rate of communication between the target device and the parameter server. 13. The method of claim 8, further comprising: converting weights for a fourth subportion of the transmitted portion of the Al model to a first precision format prior to sending the weights to the target device; converting gradients received from the target device to a second precision format; and updating the weights using the converted gradients. 14. The method of claim 8, further comprising: transmitting another portion of the Al model to another target device; and receiving gradients from the another target device to perform reduction of parameters for the another portion of the Al model. 15. A computer program product comprising a computer-readable storage device having computer program logic recorded thereon that when executed by a processor- based computer system causes the processor-based system to perform a method, the method comprising: transmitting a portion of a stored artificial intelligence (AI) model to a target device, the target device comprising an integrated circuit chip having an on-chip memory of a size less than an entirety of the Al model, and a size of the portion being based at least on the on-chip memory size and a size of one or more layers of the Al model; contemporaneously, with a set of microbatches, of a determined size, of a training dataset being executed by the integrated circuit chip at the target device on a first subportion of the transmitted portion of the Al model stored in the on-chip memory to generate gradients, performing at least one of a reduction of parameters for a second subportion of the transmitted portion of the Al model or sending weights for a third subportion of the transmitted portion of the Al model to the target device. 19. The computer program product of claim 15, wherein a microbatch size is configurable based on a rate of executing the set of microbatches at the target device and a rate of communication between the target device and the parameter server. 20. The computer program product of claim 15, wherein the method further comprises: transmitting another portion of the Al model to another target device; and receiving gradients from the another target device to perform reduction of parameters for the another portion of the Al model. 1. A system, comprising: a parameter server communicatively connected to a target device, the parameter server comprises: a data manager configured to store a master copy of an artificial intelligence (AI) model a transmitter configured to transmit a portion of an artificial intelligence (AI) model to the target device, the target device comprising an integrated circuit chip having an on-chip memory of a size less than an entirety of the Al model, and a size of the portion being based at least on the on-chip memory size and a size of one or more layers of the Al model, a batch manager configured to determine a microbatch size suitable for the target device; and contemporaneously, with a set of microbatches of a training dataset being executed by the integrated circuit chip at the target device on a first subportion of the transmitted portion of the Al model stored in the on-chip memory to generate gradients, a weight updater is configured to perform a reduction of parameters for a second subportion of the transmitted portion of the Al model, and the transmitter is further configured to send weights for a third subportion of the transmitted portion of the Al model to the target device. 5. The system of claim 2, wherein the microbatch size is configurable based on a rate of executing the set of microbatches at the target device and a rate of communication between the target device and the parameter server. 6. A system, comprising: a parameter server communicatively connected to a target device, the parameter server comprises: a data manager configured to store a master copy of an artificial intelligence (Al) model; a transmitter configured to transmit a portion of the Al model to the target device; a batch manager configured to determine a microbatch size suitable for the target device; and contemporaneously, with a set of microbatches of a training dataset being executed at the target device on a first subportion of the transmitted portion of the Al model to generate gradients, a weight updater is configured to perform reduction of parameters for a second subportion of the transmitted portion of the Al model, and the transmitter is further configured to send weights for a third subportion of the transmitted portion of the Al model to the target device, wherein the parameter server further comprises a precision formatter configured to: convert weights for a fourth subportion of the transmitted portion of the Al model to a first precision format prior to sending the weights to the target device; convert gradients received from the target device to a second precision format; and update the weights using the converted gradients. 7. The system of claim 1, wherein the transmitter is further configured to transmit another portion of the Al model to another target device; and the weight updater is further configured to receive gradients from the another target device to perform reduction of parameters for the another portion of the Al model. 8. A method implemented in a parameter server, comprising: storing a master copy in an artificial intelligence (AI) model; transmitting a portion of the Al model to a target device the target device comprising an integrated circuit chip having an on-chip memory of a size less than an entirety of the Al model, and a size of the portion being based at least on the on-chip memory size and a size of one or more layers of the Al model determining a microbatch size suitable for the target device and contemporaneously, with a set of microbatches of a training dataset being executed by the integrated circuit chip at the target device on a first subportion of the transmitted portion of the Al model stored in the on-chip memory to generate gradients, performing a reduction of parameters for a second subportion of the transmitted portion of the Al model and sending weights for a third subportion of the transmitted portion of the Al model to the target device. 12. The method of claim 9, wherein the microbatch size is configurable based on a rate of executing the set of microbatches at the target device and a rate of communication between the target device and the parameter server. 13. A method implemented in a parameter server, comprising: storing a master copy of an artificial intelligence (Al) model; transmitting a portion of the Al model to a target device; determining a microbatch size suitable for the target device; contemporaneously, with a set of microbatches of a training dataset being executed at the target device on a first subportion of the transmitted portion of the Al model to generate gradients, performing reduction of parameters for a second subportion of the transmitted portion of the Al model and sending weights for a third subportion of the transmitted portion of the Al model to the target device; converting weights for a fourth subportion of the transmitted portion of the Al model to a first precision format prior to sending the weights to the target device; converting gradients received from the target device to a second precision format; and updating the weights using the converted gradients. 14. The method of claim 8, further comprising: transmitting another portion of the Al model to another target device; and receiving gradients from the another target device to perform reduction of parameters for the another portion of the Al model. 15. A computer program product comprising a computer-readable storage device having computer program logic recorded thereon that when executed by a processor- based computer system causes the processor-based system to perform a method, the method comprising: storing a master copy of an artificial intelligence (Al) model at a parameter server; transmitting a portion of the Al model to a target device, the target device comprising an integrated circuit chip having an on-chip memory of a size less than an entirety of the Al model, and a size of the portion being based at least on the on-chip memory size and a size of one or more layers of the Al model; determining a microbatch size suitable for the target device; and contemporaneously, with a set of microbatches of a training dataset being executed by the integrated circuit chip at the target device on a first subportion of the transmitted portion of the Al model stored in the on-chip memory to generate gradients, performing reduction of parameters for a second subportion of the transmitted portion of the Al model and sending weights for a third subportion of the transmitted portion of the Al model to the target device. 19. The computer program product of claim 15, wherein the microbatch size is configurable based on a rate of executing the set of microbatches at the target device and a rate of communication between the target device and the parameter server. 20. The computer program product of claim 15, wherein the method further comprises: transmitting another portion of the Al model to another target device; and receiving gradients from the another target device to perform reduction of parameters for the another portion of the Al model. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-20 are rejected under 35 U.S.C 101 as being unpatentable because the claimed invention in these claims is directed to an abstract idea without significantly more. The analysis of the claims will follow the 2019 Revised Patent Subject Matter Eligibility Guidance, 84 Fed. Reg. 50-57 (January 7, 2019) (“2019 PEG”). Regarding claim 1 (Currently Amended): Step 1 – Is the claim directed to a process, machine, manufacture, or a composition of matter? Yes, the claim is directed to a system. Step 2A – Prong 1 – Does the claim recite an abstract idea, law of nature, or a natural phenomenon? Yes, the claim recites abstract ideas: a weight updater configured to perform a reduction of parameters for a second subportion of the transmitted portion of the Al model perform a reduction of parameters for a second subportion of the transmitted portion of the Al model — this limitation is directed to a mathematical calculation (see MPEP 2106.04(a)(2) I. C.). Step 2A – Prong 2 – Does the claim recite additional elements that integrate the judicial exception into a practical application? No, the claim recites additional elements that do not integrate the judicial exception into a practical application: a system, comprising: a parameter server communicatively connected to a target device — this limitation amounts to mere instructions to apply an exception, as the use of a computer or other machinery in its ordinary capacity amounts to invoking computer components merely as a tool to perform an existing process (see MPEP 2106.05(f)(2)). a transmitter configured to transmit a portion of an artificial intelligence (AI) model to the target device — this limitation is directed to mere data gathering and outputting which has been recognized by the courts (as per Ultramercial, 772 F.3d at 715, 112 USPQ2d at 1754) as insignificant extra-solution activity (see MPEP 2106.05(g)). the target device comprising an integrated circuit chip having an on-chip memory of a size less than an entirety of the AI model — this limitation amounts to mere instructions to apply an exception, as the use of a computer or other machinery in its ordinary capacity amounts to invoking computer components merely as a tool to perform an existing process (see MPEP 2106.05(f)(2)). a size of the portion being based at least on the on-chip memory size and a size of one or more layers of the Al model — this limitation amounts to mere instructions to apply an exception, as the use of a computer or other machinery in its ordinary capacity amounts to invoking computer components merely as a tool to perform an existing process (see MPEP 2106.05(f)(2)). wherein the transmitter is further configured to send weights for a third subportion of the transmitted portion of the Al model to the target device contemporaneously with the execution of the set of microbatches of the training dataset by the integrated circuit chip at the target device on the first subportion of the transmitted portion of the Al model stored in the on-chip memory — this limitation is directed to mere data gathering and outputting which has been recognized by the courts (as per Ultramercial, 772 F.3d at 715, 112 USPQ2d at 1754) as insignificant extra-solution activity (see MPEP 2106.05(g)). Step 2B – Does the claim recite additional elements that amount to significantly more than the abstract idea itself? No, there are no additional elements that amount to significantly more than the judicial exception. Any additional elements that were determined to be insignificant extra-solution activities in step 2A prong 2 are further evaluated in step 2B on whether they are well-understood, routine, and conventional activities. The “a transmitter configured to transmit a portion of an artificial intelligence (AI) model to the target device” and “wherein the transmitter is further configured to send weights for a third subportion of the transmitted portion of the Al model to the target device contemporaneously with the execution of the set of microbatches of the training dataset by the integrated circuit chip at the target device on the first subportion of the transmitted portion of the Al model stored in the on-chip memory” limitations were found to be insignificant extra-solution activities in claim 1. These limitations are recited at a high level of generality and amounts to transmitting data over a network, which is a well-understood, routine, and conventional activity (see MPEP 2106.05(d) II.). Thus, the claim is not patent eligible. Regarding claim 2 (Currently Amended): Claim 2 recites a data parallelism process using machine learning at a high level of generality, which amounts to apply the judicial exception on a computer (see MPEP 2106.05(f)). Claims 9 and 16 are analogous. Regarding claim 3: Claim 3 recites a data parallelism process using machine learning at a high level of generality, which amounts to apply the judicial exception on a computer (see MPEP 2106.05(f)). Claims 10 and 17 are analogous. Regarding claim 4: Claim 4 recites receiving gradients from a target device, which amounts to mere data gathering, considered a pre-solution activity which is an insignificant extra-solution activity (see MPEP 2106.05(g) (3)) which has been recognized by the courts (as per Ultramercial, 772 F.3d at 715, 112 USPQ2d at 1754). Claim 4 also recites gradients being generated by executing microbatches, which is an abstract idea (a calculation) see MPEP 2106.04(a)(2) I. C.), generating an average of gradients, which is an abstract idea (a calculation) see MPEP 2106.04(a)(2) I. C.), and updating an AI model which amounts to a machine learning process at a high level of generality, which amounts to apply the judicial exception on a computer (see MPEP 2106.05(f)). Claims 11 and 18 are analogous. Regarding claim 5: Claim 5 recites a machine learning process (allowing for a batch, or subset of data, size to be configurable) recited at a high level of generality, which amounts to mere instructions to apply the judicial exception on a computer (see MPEP 2106.05(f)).Claims 12 and 19 are analogous. Regarding claim 6: Claim 6 recites machine learning processes (converting weights and updating weights), at a high level of generality, which amounts to mere instructions to apply the judicial exception on a computer (see MPEP 2106.05(f)). Claim 13 is analogous. Regarding claim 7: Claim 7 recites transmitting a portion of an AI model, as well as receiving gradients from a device, which amount to mere data gathering, considered a pre-solution activity which is an insignificant extra-solution activity (see MPEP 2106.05(g) (3)) which has been recognized by the courts (as per Ultramercial, 772 F.3d at 715, 112 USPQ2d at 1754). Furthermore,gathering data to be manipulated and used within a system is a well-understood, routine, and conventional activity (WURC) that cannot provide significantly more than the judicial exception (see MPEP 2106.05(d)(II)). Claims 14 and 20 are analogous. Regarding claim 8: Step 1 – Is the claim directed to a process, machine, manufacture, or a composition of matter? Yes, the claim is directed to a method. Step 2A – Prong 1 – Does the claim recite an abstract idea, law of nature, or a natural phenomenon? Yes, the claim recites abstract ideas: contemporaneously, with a set of microbatches, of a determined size, of a training dataset being executed by the integrated circuit chip at the target device on a first subportion of the transmitted portion of the Al model stored in the on-chip memory to generate gradients, performing a reduction of parameters for a second subportion of the transmitted portion of the Al model — this limitation is directed to a mathematical calculation (see MPEP 2106.04(a)(2) I. C.). Step 2A – Prong 2 – Does the claim recite additional elements that integrate the judicial exception into a practical application? No, the claim recites additional elements that do not integrate the judicial exception into a practical application: transmitting a portion of a stored artificial intelligence (AI) model to a target device — this limitation is directed to mere data gathering and outputting which has been recognized by the courts (as per Ultramercial, 772 F.3d at 715, 112 USPQ2d at 1754) as insignificant extra-solution activity (see MPEP 2106.05(g)). the target device comprising an integrated circuit chip having an on-chip memory of a size less than an entirety of the AI model — this limitation amounts to mere instructions to apply an exception, as the use of a computer or other machinery in its ordinary capacity amounts to invoking computer components merely as a tool to perform an existing process (see MPEP 2106.05(f)(2)). a size of the portion being based at least on the on-chip memory size and a size of one or more layers of the Al model — this limitation amounts to mere instructions to apply an exception, as the use of a computer or other machinery in its ordinary capacity amounts to invoking computer components merely as a tool to perform an existing process (see MPEP 2106.05(f)(2)). contemporaneously, with a set of microbatches, of a determined size, of a training dataset being executed by the integrated circuit chip at the target device on a first subportion of the transmitted portion of the Al model stored in the on-chip memory to generate gradients, sending weights for a third subportion of the transmitted portion of the Al model to the target device — this limitation is directed to mere data gathering and outputting which has been recognized by the courts (as per Ultramercial, 772 F.3d at 715, 112 USPQ2d at 1754) as insignificant extra-solution activity (see MPEP 2106.05(g)). Step 2B – Does the claim recite additional elements that amount to significantly more than the abstract idea itself? No, there are no additional elements that amount to significantly more than the judicial exception. Any additional elements that were determined to be insignificant extra-solution activities in step 2A prong 2 are further evaluated in step 2B on whether they are well-understood, routine, and conventional activities. The “transmitting a portion of a stored artificial intelligence (AI) model to a target device” and “sending weights for a third subportion of the transmitted portion of the Al model to the target device” limitations were found to be insignificant extra-solution activities in claim 8. This limitation is recited at a high level of generality and amounts to transmitting data over a network, which is a well-understood, routine, and conventional activity (see MPEP 2106.05(d) II.). Thus, the claim is not patent eligible. Regarding claim 15: Step 1 – Is the claim directed to a process, machine, manufacture, or a composition of matter? Yes, the claim is directed to a manufacture (a computer program product). Step 2A – Prong 1 – Does the claim recite an abstract idea, law of nature, or a natural phenomenon? Yes, the claim recites abstract ideas: contemporaneously, with a set of microbatches, of a determined size, of a training dataset being executed by the integrated circuit chip at the target device on a first subportion of the transmitted portion of the Al model stored in the on-chip memory to generate gradients, performing at least one of a reduction of parameters for a second subportion of the transmitted portion of the Al model or sending weights for a third subportion of the transmitted portion of the Al model to the target device — this limitation is directed to a mathematical calculation (see MPEP 2106.04(a)(2) I. C.). Step 2A – Prong 2 – Does the claim recite additional elements that integrate the judicial exception into a practical application? No, the claim recites additional elements that do not integrate the judicial exception into a practical application: a computer program product comprising a computer-readable storage device having computer program logic recorded thereon that when executed by a processor-based computer system causes the processor-based system to perform a method — this limitation amounts to mere instructions to apply an exception, as the use of a computer or other machinery in its ordinary capacity amounts to invoking computer components merely as a tool to perform an existing process (see MPEP 2106.05(f)(2)). transmitting a portion of a stored artificial intelligence (AI) model to a target device — this limitation is directed to mere data gathering and outputting which has been recognized by the courts (as per Ultramercial, 772 F.3d at 715, 112 USPQ2d at 1754) as insignificant extra-solution activity (see MPEP 2106.05(g)). the target device comprising an integrated circuit chip having an on-chip memory of a size less than an entirety of the AI model — this limitation amounts to mere instructions to apply an exception, as the use of a computer or other machinery in its ordinary capacity amounts to invoking computer components merely as a tool to perform an existing process (see MPEP 2106.05(f)(2)). a size of the portion being based at least on the on-chip memory size and a size of one or more layers of the Al model — this limitation amounts to mere instructions to apply an exception, as the use of a computer or other machinery in its ordinary capacity amounts to invoking computer components merely as a tool to perform an existing process (see MPEP 2106.05(f)(2)). contemporaneously, with a set of microbatches, of a determined size, of a training dataset being executed by the integrated circuit chip at the target device on a first subportion of the transmitted portion of the Al model stored in the on-chip memory to generate gradients, performing at least one of a reduction of parameters for a second subportion of the transmitted portion of the AI model or sending weights for a third subportion of the transmitted portion of the Al model to the target device — this limitation is directed to mere data gathering and outputting which has been recognized by the courts (as per Ultramercial, 772 F.3d at 715, 112 USPQ2d at 1754) as insignificant extra-solution activity (see MPEP 2106.05(g)). Step 2B – Does the claim recite additional elements that amount to significantly more than the abstract idea itself? No, there are no additional elements that amount to significantly more than the judicial exception. Any additional elements that were determined to be insignificant extra-solution activities in step 2A prong 2 are further evaluated in step 2B on whether they are well-understood, routine, and conventional activities. The “transmitting a portion of a stored artificial intelligence (AI) model to a target device” and “sending weights for a third subportion of the transmitted portion of the Al model to the target device” limitations were found to be insignificant extra-solution activities in claim 15. This limitation is recited at a high level of generality and amounts to transmitting data over a network, which is a well-understood, routine, and conventional activity (see MPEP 2106.05(d) II.). Thus, the claim is not patent eligible. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claims 1-4, 6-11, 13-18, and 20 are rejected under 35 U.S.C 103 as being unpatentable over Sridharan et al. (US20190205745A1 hereinafter referred to as Sridharan) in view of Ambrose et al. (US20170344882A1 hereinafter referred to as Ambrose). Regarding claim 1 (Currently Amended): Sridharan teaches a system, comprising: a parameter server communicatively connected to a target device (see [0047]: “FIG. 1 is a block diagram of a processing system 100, according to an embodiment. In various embodiments the system 100 includes one or more processors 102 and one or more graphics processors 108, and may be a single processor desktop system, a multiprocessor workstation system, or a server system having a large number of processors 102 or processor cores 107. In one embodiment, the system 100 is a processing platform incorporated within a system-on-a-chip (SoC) integrated circuit for use in mobile, handheld, or embedded devices.”. Also see [0223]: “The communications framework can then create a second group 2230 of worker nodes can includes additional sets of worker nodes 2236A-2236B having low mutual latency. The specific number of groups that are created can vary based on the topology of the network. One or more parameter servers 2220 can then be instantiated. The parameter servers can then be configured to enable efficient inter-group communication between the various groups 2210, 2230.”.) the parameter server comprises: a transmitter configured to transmit a portion of an artificial intelligence (AI) model to the target device (see [0222]: “Described herein is a communication system that makes use of a topology-aware algorithm for flexible node grouping. In one embodiment, a distributed training system can be constructed in a manner that is sensitive to the existing network topology of the worker nodes, such that local nodes can be assembled into compute groups based on network topology. Nodes within a compute group communicate with each other using operations such as all-reduce, while distant nodes are bridged via synchronization operations performed with a parameter server.”.) and a weight updater configured to perform a reduction of parameters for a second subportion of the transmitted portion of the Al model (see [0202]: “ An allreduce operation 2005 is used to update the weights of each layer for the next forward pass”. Also see [0206]: “During back propagation 2028, distributed stochastic gradient descent is performed to generate updated weight data. An initial Allreduce operation 2012 is performed for Layer N and a set of Allreduce operations 2011A, 2011B, 2011N are performed to update the weights of each layer for the next forward pass.”.) contemporaneously with an execution of a set of microbatches of a training dataset by the integrated circuit chip at the target device on a first subportion of the transmitted portion of the AI model stored in the on-chip memory (see para [0202]: “As shown in FIG. 20A, data parallelism can be implemented in which input data 2002 is split along a mini-batch dimension and the same model is replicated across the nodes. The mini-batch is split across several compute nodes, with each node responsible for computing gradients with respect to all model parameters using a subset of the samples in the mini-batch. Forward propagation is performed independently on each node. In one embodiment only one communication is performed during the backward pass to calculate an average for the gradients with respect to learnable parameters. An allreduce operation 2005 is used to update the weights of each layer for the next forward pass.”. Also see para [0017]: “FIG. 12 is a block diagram illustrating an exemplary system on a chip integrated circuit, according to an embodiment;”. Also see fig. 12 which shows the on-chip memory. Also see fig. 19 which shows model parallelism of different sub portions, or layers, of the model) wherein the transmitter is further configured to send weights for a third subportion of the transmitted portion of the Al model to the target device contemporaneously with the execution of the set of microbatches of the training dataset by the integrated circuit chip at the target device on the first subportion of the transmitted portion of the AI model stored in the on-chip memory (see para [0202]: “As shown in FIG. 20A, data parallelism can be implemented in which input data 2002 is split along a mini-batch dimension and the same model is replicated across the nodes. The mini-batch is split across several compute nodes, with each node responsible for computing gradients with respect to all model parameters using a subset of the samples in the mini-batch. Forward propagation is performed independently on each node. In one embodiment only one communication is performed during the backward pass to calculate an average for the gradients with respect to learnable parameters. An allreduce operation 2005 is used to update the weights of each layer for the next forward pass.”. Also see para [0017]: “FIG. 12 is a block diagram illustrating an exemplary system on a chip integrated circuit, according to an embodiment;”. Also see fig. 12 which shows the on-chip memory. Also see fig. 19 which shows model parallelism of different sub portions, or layers, of the model) Sridharan does not explicitly teach the target device comprising an integrated circuit chip having an on-chip memory of a size less than an entirety of the Al model or a size of the portion being based at least on the on-chip memory size and a size of one or more layers of the Al model. Ambrose, however, teaches in analogous teach the target device comprising an integrated circuit chip having an on-chip memory of a size less than an entirety of the Al model (see [0005]: “Graphical Processing Units (GPUs) are a strong candidate for implementing CNN algorithms because GPUs, which are suitable for parallel computation, are well adapted to exploit the high level of data parallelism in the CNN algorithms.”. Also see [0046]: “The best scheduling scheme 1722 for the particular CNN algorithm layer 1704 is defined as the scheduling scheme which satisfies given criteria with minimal design costs to execute on the SoC 1714 for the layer 17104 in question. The criteria and design costs are for the entire SoC. For example, the best scheduling schemes can be ones which require a local shared memory (such as 1710) whose size is within a given constraint while consuming minimal accesses to the external memory 1709 thus resulting in a smaller execution time for the entire SoC. In order to reduce the number of simulations but still be able to find the best selection of scheduling schemes to layers of the CNN algorithm, an estimation method is required in order to accurately predict the design costs, such as the required local memory size, the external memory accesses and the execution time of the CNN algorithm 1703.”.) and a size of the portion being based at least on the on-chip memory size and a size of one or more layers of the Al model (see [0005]: “Graphical Processing Units (GPUs) are a strong candidate for implementing CNN algorithms because GPUs, which are suitable for parallel computation, are well adapted to exploit the high level of data parallelism in the CNN algorithms.”. Also see [0046]: “The best scheduling scheme 1722 for the particular CNN algorithm layer 1704 is defined as the scheduling scheme which satisfies given criteria with minimal design costs to execute on the SoC 1714 for the layer 17104 in question. The criteria and design costs are for the entire SoC. For example, the best scheduling schemes can be ones which require a local shared memory (such as 1710) whose size is within a given constraint while consuming minimal accesses to the external memory 1709 thus resulting in a smaller execution time for the entire SoC. In order to reduce the number of simulations but still be able to find the best selection of scheduling schemes to layers of the CNN algorithm, an estimation method is required in order to accurately predict the design costs, such as the required local memory size, the external memory accesses and the execution time of the CNN algorithm 1703.”.) Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art, having the teachings of Sridharan and Ambrose before him or her, to modify the system of claim 1 to include attributes of having the target device comprising an integrated circuit chip having an on-chip memory of a size less than an entirety of the Al model or a size of the portion being based at least on the on-chip memory size and a size of one or more layers of the Al model in order to allow for optimal scheduling schemes (see [0046]: “Also see [0046]: “The best scheduling scheme 1722 for the particular CNN algorithm layer 1704 is defined as the scheduling scheme which satisfies given criteria with minimal design costs to execute on the SoC 1714 for the layer 17104 in question. The criteria and design costs are for the entire SoC. For example, the best scheduling schemes can be ones which require a local shared memory (such as 1710) whose size is within a given constraint while consuming minimal accesses to the external memory 1709 thus resulting in a smaller execution time for the entire SoC.”.). Regarding claim 2 (Currently Amended): Sridharan in view of Ambrose teaches the system of claim 1. Sridharan further teaches wherein the first subportion of the transmitted portion of the Al model comprises a current layer of the Al model (see [0204]: “As shown in FIG. 20C, hybrid parallelism can be performed in which a partitioning is performed across activations and weights to minimize skewed matrices. For a layer of a neural network, the input data 2002, weight data 2004, and/or activation data 2006 is partitioned and distributed across multiple compute nodes (e.g., Node 0-Node 3).”) and and the second subportion of the transmitted portion of the Al model comprises a prior layer of the Al model (see [0205]: “FIG. 20D illustrates the transfer of partial activation data 2006A-2006B for a given layer of a neural network (Layer N-1) to a successive layer of the neural network (Layer N). Via multiple nodes (Node 0, Node 1), a set of partial activations 2006A-2006B is generated by based on the application of a mathematical operation (e.g., convolution) to the input data 2002A-2002B and weight data 2004A-2004B. For example, in one embodiment a reduce_scatter operation 2010 is used which performs a reduce operation on the partial activations 2006A-2006B of layer N-1 from the multiple nodes and scatters the result to the multiple nodes as activations for use in Layer N of the neural network.”). Regarding claims 9 (Currently Amended) and 16: Claims 9 and 16 recite analogous limitations to claim 2 and are therefore rejected on the same grounds. Regarding claim 3: Sridharan in view of Ambrose teaches the system of claim 1. Sridharan further teaches wherein the first subportion of the transmitted portion of the Al model comprises a current layer of the Al model (see [0205]: “FIG. 20D illustrates the transfer of partial activation data 2006A-2006B for a given layer of a neural network (Layer N-1) to a successive layer of the neural network (Layer N). Via multiple nodes (Node 0, Node 1), a set of partial activations 2006A-2006B is generated by based on the application of a mathematical operation (e.g., convolution) to the input data 2002A-2002B and weight data 2004A-2004B. For example, in one embodiment a reduce_scatter operation 2010 is used which performs a reduce operation on the partial activations 2006A-2006B of layer N-1 from the multiple nodes and scatters the result to the multiple nodes as activations for use in Layer N of the neural network.”.) the third subportion of the transmitted portion of the Al model comprises a next layer of the Al model (see [0205]: “FIG. 20D illustrates the transfer of partial activation data 2006A-2006B for a given layer of a neural network (Layer N-1) to a successive layer of the neural network (Layer N). Via multiple nodes (Node 0, Node 1), a set of partial activations 2006A-2006B is generated by based on the application of a mathematical operation (e.g., convolution) to the input data 2002A-2002B and weight data 2004A-2004B. For example, in one embodiment a reduce_scatter operation 2010 is used which performs a reduce operation on the partial activations 2006A-2006B of layer N-1 from the multiple nodes and scatters the result to the multiple nodes as activations for use in Layer N of the neural network.”.). Regarding claims 10 and 17: Claims 10 and 17 recite analogous limitations to claim 3 and are therefore rejected on the same grounds. Regarding claim 4: Sridharan in view of Ambrose teaches the system of claim 1. Sridharan further teaches wherein the weight updater is configured to: perform reduction of parameters by: receiving gradients from the target device, the gradients being generated by the target device executing the set of microbatches of the training dataset on the second subportion at the target device (see [0202]: “As shown in FIG. 20A, data parallelism can be implemented in which input data 2002 is split along a mini-batch dimension and the same model is replicated across the nodes. The mini-batch is split across several compute nodes, with each node responsible for computing gradients with respect to all model parameters using a subset of the samples in the mini-batch.”.) generating an average of the received gradients (see [0202]: “Forward propagation is performed independently on each node. In one embodiment only one communication is performed during the backward pass to calculate an average for the gradients with respect to learnable parameters. ”.) update the Al model with the average of the received gradients (see [0202]: “An allreduce operation 2005 is used to update the weights of each layer for the next forward pass. In one embodiment, distributed weight update can be enabled in which a reduce_scatter is used calculate an average for gradients before stochastic gradient descent is performed and an allgather operation is used after stochastic gradient descent to synchronize weights across nodes.”.) Regarding claims 11 and 18: Claims 11 and 18 recite analogous limitations to claim 4 and are therefore rejected on the same grounds. Regarding claim 6: Sridharan in view of Ambrose teaches the system of claim 1. Sridharan further teaches convert weights for a fourth subportion of the transmitted portion of the Al model to a first precision format prior to sending the weights to the target device (see [0202]: “As shown in FIG. 20A, data parallelism can be implemented in which input data 2002 is split along a mini-batch dimension and the same model is replicated across the nodes. The mini-batch is split across several compute nodes, with each node responsible for computing gradients with respect to all model parameters using a subset of the samples in the mini-batch.”.) convert gradients received from the target device to a second precision format (see [0202]: “Forward propagation is performed independently on each node. In one embodiment only one communication is performed during the backward pass to calculate an average for the gradients with respect to learnable parameters. ”.) update the weights using the converted gradients (see [0202]: “An allreduce operation 2005 is used to update the weights of each layer for the next forward pass. In one embodiment, distributed weight update can be enabled in which a reduce_scatter is used calculate an average for gradients before stochastic gradient descent is performed and an allgather operation is used after stochastic gradient descent to synchronize weights across nodes.”.) Regarding claim 13: Claim 13 recites analogous limitations to claim 6 and is therefore rejected on the same grounds. Regarding claim 7: Sridharan in view of Ambrose teaches the system of claim 1. Sridharan further teaches wherein the transmitter is further configured to transmit another portion of the Al model to another target device (see [0204]: “As shown in FIG. 20C, hybrid parallelism can be performed in which a partitioning is performed across activations and weights to minimize skewed matrices. For a layer of a neural network, the input data 2002, weight data 2004, and/or activation data 2006 is partitioned and distributed across multiple compute nodes (e.g., Node 0-Node 3).”) and the weight updater is further configured to receive gradients from the another target device to perform reduction of parameters for the another portion of the Al model (see [0202]: “ An allreduce operation 2005 is used to update the weights of each layer for the next forward pass”. Also see [0206]: “During back propagation 2028, distributed stochastic gradient descent is performed to generate updated weight data. An initial Allreduce operation 2012 is performed for Layer N and a set of Allreduce operations 2011A, 2011B, 2011N are performed to update the weights of each layer for the next forward pass.”.) Regarding claims 14 and 20: Claims 14 and 20 recite analogous limitations to claim 7 and are therefore rejected on the same grounds. Regarding claim 8: Sridharan teaches a method implemented in a parameter server, comprising, (see [0223]: “The communications framework can then create a second group 2230 of worker nodes can includes additional sets of worker nodes 2236A-2236B having low mutual latency. The specific number of groups that are created can vary based on the topology of the network. One or more parameter servers 2220 can then be instantiated. The parameter servers can then be configured to enable efficient inter-group communication between the various groups 2210, 2230.”.) transmitting a portion of an artificial intelligence (AI) model to the target device (see [0222]: “Described herein is a communication system that makes use of a topology-aware algorithm for flexible node grouping. In one embodiment, a distributed training system can be constructed in a manner that is sensitive to the existing network topology of the worker nodes, such that local nodes can be assembled into compute groups based on network topology. Nodes within a compute group communicate with each other using operations such as all-reduce, while distant nodes are bridged via synchronization operations performed with a parameter server.”.) and contemporaneously, with a set of microbatches, of a determined size, of a training dataset being executed by the integrated circuit chip at the target device on a first subportion of the transmitted portion of the Al model stored in the on-chip memory to generate gradients (see [0202]: “As shown in FIG. 20A, data parallelism can be implemented in which input data 2002 is split along a mini-batch dimension and the same model is replicated across the nodes. The mini-batch is split across several compute nodes, with each node responsible for computing gradients with respect to all model parameters using a subset of the samples in the mini-batch. Forward propagation is performed independently on each node. In one embodiment only one communication is performed during the backward pass to calculate an average for the gradients with respect to learnable parameters.”.), performing a reduction of parameters for a second subportion of the transmitted portion of the Al model (see [0202]: “ An allreduce operation 2005 is used to update the weights of each layer for the next forward pass”. Also see [0206]: “During back propagation 2028, distributed stochastic gradient descent is performed to generate updated weight data. An initial Allreduce operation 2012 is performed for Layer N and a set of Allreduce operations 2011A, 2011B, 2011N are performed to update the weights of each layer for the next forward pass.”.) and sending weights for a third subportion of the transmitted portion of the Al model to the target device (see [0202]: “ An allreduce operation 2005 is used to update the weights of each layer for the next forward pass”. Also see [0206]: “During back propagation 2028, distributed stochastic gradient descent is performed to generate updated weight data. An initial Allreduce operation 2012 is performed for Layer N and a set of Allreduce operations 2011A, 2011B, 2011N are performed to update the weights of each layer for the next forward pass.”.) [(Examiner’s note: A person having ordinary skill in the art using broadest reasonable interpretation in light of the specification could take “ set of microbatches” to be taken as at least one minibatch. In the instant application, the specification at [0040] reads: “A group of microbatches forms a minibatch, which is the term for the number of samples per update (for training) or the number served in every inference cycle (for inference).”, which further gives reason to believe that here a group and set could be taken to be synonymous.)] Sridharan does not explicitly teach the target device comprising an integrated circuit chip having an on-chip memory of a size less than an entirety of the Al model or a size of the portion being based at least on the on-chip memory size and a size of one or more layers of the Al model. Ambrose, however, teaches in analogous teach the target device comprising an integrated circuit chip having an on-chip memory of a size less than an entirety of the Al model (see [0005]: “Graphical Processing Units (GPUs) are a strong candidate for implementing CNN algorithms because GPUs, which are suitable for parallel computation, are well adapted to exploit the high level of data parallelism in the CNN algorithms.”. Also see [0046]: “The best scheduling scheme 1722 for the particular CNN algorithm layer 1704 is defined as the scheduling scheme which satisfies given criteria with minimal design costs to execute on the SoC 1714 for the layer 17104 in question. The criteria and design costs are for the entire SoC. For example, the best scheduling schemes can be ones which require a local shared memory (such as 1710) whose size is within a given constraint while consuming minimal accesses to the external memory 1709 thus resulting in a smaller execution time for the entire SoC. In order to reduce the number of simulations but still be able to find the best selection of scheduling schemes to layers of the CNN algorithm, an estimation method is required in order to accurately predict the design costs, such as the required local memory size, the external memory accesses and the execution time of the CNN algorithm 1703.”.) and a size of the portion being based at least on the on-chip memory size and a size of one or more layers of the Al model (see [0005]: “Graphical Processing Units (GPUs) are a strong candidate for implementing CNN algorithms because GPUs, which are suitable for parallel computation, are well adapted to exploit the high level of data parallelism in the CNN algorithms.”. Also see [0046]: “The best scheduling scheme 1722 for the particular CNN algorithm layer 1704 is defined as the scheduling scheme which satisfies given criteria with minimal design costs to execute on the SoC 1714 for the layer 17104 in question. The criteria and design costs are for the entire SoC. For example, the best scheduling schemes can be ones which require a local shared memory (such as 1710) whose size is within a given constraint while consuming minimal accesses to the external memory 1709 thus resulting in a smaller execution time for the entire SoC. In order to reduce the number of simulations but still be able to find the best selection of scheduling schemes to layers of the CNN algorithm, an estimation method is required in order to accurately predict the design costs, such as the required local memory size, the external memory accesses and the execution time of the CNN algorithm 1703.”.) Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art, having the teachings of Sridharan and Ambrose before him or her, to modify the method of claim 8 to include attributes of having the target device comprising an integrated circuit chip having an on-chip memory of a size less than an entirety of the Al model or a size of the portion being based at least on the on-chip memory size and a size of one or more layers of the Al model in order to allow for optimal scheduling schemes (see [0046]: “Also see [0046]: “The best scheduling scheme 1722 for the particular CNN algorithm layer 1704 is defined as the scheduling scheme which satisfies given criteria with minimal design costs to execute on the SoC 1714 for the layer 17104 in question. The criteria and design costs are for the entire SoC. For example, the best scheduling schemes can be ones which require a local shared memory (such as 1710) whose size is within a given constraint while consuming minimal accesses to the external memory 1709 thus resulting in a smaller execution time for the entire SoC.”.). Regarding claim 15: Sridharan teaches a computer program product comprising a computer-readable storage device having computer program logic recorded thereon that when executed by a processor-based computer system causes the processor-based system to perform a method, the method comprising: (see [0047]: “FIG. 1 is a block diagram of a processing system 100, according to an embodiment. In various embodiments the system 100 includes one or more processors 102 and one or more graphics processors 108, and may be a single processor desktop system, a multiprocessor workstation system, or a server system having a large number of processors 102 or processor cores 107. In one embodiment, the system 100 is a processing platform incorporated within a system-on-a-chip (SoC) integrated circuit for use in mobile, handheld, or embedded devices.”. Also see [0223]: “The communications framework can then create a second group 2230 of worker nodes can includes additional sets of worker nodes 2236A-2236B having low mutual latency. The specific number of groups that are created can vary based on the topology of the network. One or more parameter servers 2220 can then be instantiated. The parameter servers can then be configured to enable efficient inter-group communication between the various groups 2210, 2230.”.) transmitting a portion of an artificial intelligence (AI) model to the target device (see [0222]: “Described herein is a communication system that makes use of a topology-aware algorithm for flexible node grouping. In one embodiment, a distributed training system can be constructed in a manner that is sensitive to the existing network topology of the worker nodes, such that local nodes can be assembled into compute groups based on network topology. Nodes within a compute group communicate with each other using operations such as all-reduce, while distant nodes are bridged via synchronization operations performed with a parameter server.”.) and contemporaneously, with a set of microbatches of a training dataset being executed by the integrated circuit chip at the target device on a first subportion of the transmitted portion of the Al model stored in the on-chip memory to generate gradients (see [0202]: “As shown in FIG. 20A, data parallelism can be implemented in which input data 2002 is split along a mini-batch dimension and the same model is replicated across the nodes. The mini-batch is split across several compute nodes, with each node responsible for computing gradients with respect to all model parameters using a subset of the samples in the mini-batch. Forward propagation is performed independently on each node. In one embodiment only one communication is performed during the backward pass to calculate an average for the gradients with respect to learnable parameters.”.), performing a reduction of parameters for a second subportion of the transmitted portion of the Al model (see [0202]: “ An allreduce operation 2005 is used to update the weights of each layer for the next forward pass”. Also see [0206]: “During back propagation 2028, distributed stochastic gradient descent is performed to generate updated weight data. An initial Allreduce operation 2012 is performed for Layer N and a set of Allreduce operations 2011A, 2011B, 2011N are performed to update the weights of each layer for the next forward pass.”.) or sending weights for a third subportion of the transmitted portion of the Al model to the target device (see [0202]: “ An allreduce operation 2005 is used to update the weights of each layer for the next forward pass”. Also see [0206]: “During back propagation 2028, distributed stochastic gradient descent is performed to generate updated weight data. An initial Allreduce operation 2012 is performed for Layer N and a set of Allreduce operations 2011A, 2011B, 2011N are performed to update the weights of each layer for the next forward pass.”.) [(Examiner’s note: A person having ordinary skill in the art using broadest reasonable interpretation in light of the specification could take “ set of microbatches” to be taken as at least one minibatch. In the instant application, the specification at [0040] reads: “A group of microbatches forms a minibatch, which is the term for the number of samples per update (for training) or the number served in every inference cycle (for inference).”, which further gives reason to believe that here a group and set could be taken to be synonymous.)] Sridharan does not explicitly teach the target device comprising an integrated circuit chip having an on-chip memory of a size less than an entirety of the Al model or a size of the portion being based at least on the on-chip memory size and a size of one or more layers of the Al model. Ambrose, however, teaches in analogous teach the target device comprising an integrated circuit chip having an on-chip memory of a size less than an entirety of the Al model (see [0005]: “Graphical Processing Units (GPUs) are a strong candidate for implementing CNN algorithms because GPUs, which are suitable for parallel computation, are well adapted to exploit the high level of data parallelism in the CNN algorithms.”. Also see [0046]: “The best scheduling scheme 1722 for the particular CNN algorithm layer 1704 is defined as the scheduling scheme which satisfies given criteria with minimal design costs to execute on the SoC 1714 for the layer 17104 in question. The criteria and design costs are for the entire SoC. For example, the best scheduling schemes can be ones which require a local shared memory (such as 1710) whose size is within a given constraint while consuming minimal accesses to the external memory 1709 thus resulting in a smaller execution time for the entire SoC. In order to reduce the number of simulations but still be able to find the best selection of scheduling schemes to layers of the CNN algorithm, an estimation method is required in order to accurately predict the design costs, such as the required local memory size, the external memory accesses and the execution time of the CNN algorithm 1703.”.) and a size of the portion being based at least on the on-chip memory size and a size of one or more layers of the Al model (see [0005]: “Graphical Processing Units (GPUs) are a strong candidate for implementing CNN algorithms because GPUs, which are suitable for parallel computation, are well adapted to exploit the high level of data parallelism in the CNN algorithms.”. Also see [0046]: “The best scheduling scheme 1722 for the particular CNN algorithm layer 1704 is defined as the scheduling scheme which satisfies given criteria with minimal design costs to execute on the SoC 1714 for the layer 17104 in question. The criteria and design costs are for the entire SoC. For example, the best scheduling schemes can be ones which require a local shared memory (such as 1710) whose size is within a given constraint while consuming minimal accesses to the external memory 1709 thus resulting in a smaller execution time for the entire SoC. In order to reduce the number of simulations but still be able to find the best selection of scheduling schemes to layers of the CNN algorithm, an estimation method is required in order to accurately predict the design costs, such as the required local memory size, the external memory accesses and the execution time of the CNN algorithm 1703.”.) Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art, having the teachings of Sridharan and Ambrose before him or her, to modify the computer program product of claim 15 to include attributes of having the target device comprising an integrated circuit chip having an on-chip memory of a size less than an entirety of the Al model or a size of the portion being based at least on the on-chip memory size and a size of one or more layers of the Al model in order to allow for optimal scheduling schemes (see [0046]: “Also see [0046]: “The best scheduling scheme 1722 for the particular CNN algorithm layer 1704 is defined as the scheduling scheme which satisfies given criteria with minimal design costs to execute on the SoC 1714 for the layer 17104 in question. The criteria and design costs are for the entire SoC. For example, the best scheduling schemes can be ones which require a local shared memory (such as 1710) whose size is within a given constraint while consuming minimal accesses to the external memory 1709 thus resulting in a smaller execution time for the entire SoC.”.). Claims 5, 12, and 19 are rejected under 35 U.S.C 103 as being unpatentable over Sridharan et al. (US20190205745A1 hereinafter referred to as Sridharan) in view of Ambrose et al. (US20170344882A1 hereinafter referred to as Ambrose) in further view of Oyama et al. (“Accelerating Deep Learning Frameworks with Micro-Batches” hereinafter referred to as Oyama). Regarding claim 5: Sridharan in view of Ambrose teaches the system of claim 1. Sridharan in view of Ambrose does not explicitly teach wherein a microbatch size is configurable based on a rate of executing the set of microbatches at the target device and a rate of communication between the target device and the parameter server. Oyama, however, teaches in analogous wherein a microbatch size is configurable based on a rate of executing the set of microbatches at the target device and a rate of communication between the target device and the parameter server (see section III B page 5 : “The goal of the WR policy is to minimize T(B), the total execution time with mini-batch size of B using Dynamic Programming (DP), where T(b) is defined as follows: PNG media_image1.png 86 608 media_image1.png Greyscale where Tμ(b) is the fastest execution time of one convolution kernel with a micro-batch size of b, within the workspace constraint.” ) Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art, having the teachings of Sridharan, Ambrose, and Oyama before him or her, to modify the system of claim 5 to include attributes of a microbatch size is configurable based on a rate of executing the set of microbatches at the target device and a rate of communication between the target device and the parameter server in order to minimize total workspace time (see section III B page 5: “The goal of the WR policy is to minimize T(B), the total execution time with mini-batch size of B”.). Regarding claims 12 and 19: Claims 12 and 19 recite analogous limitations to claim 5 and are therefore rejected on the same grounds. Pertinent Prior Art The prior art made of record and not relied upon is considered pertinent to applicant’s disclosure: US20190073590A1 — Wu et al. — discloses system-on-chip architecture using data parallelism and mini-batches “Performance Characterization and Optimization of In-Memory Data Analytics on a Scale-up Server” — Awan — discloses system-on-chip as well as server-on-chip architecture in the context of data parallelism “PipeDream: Fast and Efficient Pipeline Parallel DNN” — Harlap et al. — discloses data parallelism and the use of minibatches “Demystifying Parallel and Distributed Deep Learning: An In-Depth Concurrency Analysis” — Ben-Nun et al.(1) — discloses single-machine parallelism architecture in the context of data parallelism “A Modular Benchmarking Infrastructure for High-Performance and Reproducible Deep Learning” — Ben-Nun et al.(2) — discloses the use of micro-batches in the context of data parallelism Conclusion THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to Andrew A Bracero whose telephone number is (571)270-0592. The examiner can normally be reached Monday - Friday 7:30a.m. - 5:00 p.m. ET. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, David Yi can be reached at Monday – Friday 9:00 a.m. – 5:00 p.m. E.T. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /ANDREW BRACERO/Examiner, Art Unit 2126 /DAVID YI/Supervisory Patent Examiner, Art Unit 2126
Read full office action

Prosecution Timeline

May 24, 2022
Application Filed
Sep 18, 2025
Non-Final Rejection mailed — §101, §103, §DOUBLEPATENT
Dec 18, 2025
Applicant Interview (Telephonic)
Dec 19, 2025
Examiner Interview Summary
Feb 18, 2026
Response Filed
May 27, 2026
Final Rejection mailed — §101, §103, §DOUBLEPATENT (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12688517
METHOD AND APPARATUS FOR TRAINING ONLINE PREDICTION MODEL, DEVICE AND STORAGE MEDIUM
5y 3m to grant Granted Jul 21, 2026
Patent 12651149
NEURAL NETWORK COMPUTATION APPARATUS HAVING SYSTOLIC ARRAY
5y 4m to grant Granted Jun 09, 2026
Patent 12651156
METHOD AND APPARATUS FOR DETERMINING CAUSALITY, ELECTRONIC DEVICE AND STORAGE MEDIUM
5y 2m to grant Granted Jun 09, 2026
Patent 12619870
PROGRAMMABLE NON-LINEAR ACTIVATION ENGINE FOR NEURAL NETWORK ACCELERATION
4y 1m to grant Granted May 05, 2026
Study what changed to get past this examiner. Based on 4 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
90%
Grant Probability
99%
With Interview (+33.3%)
4y 7m (~4m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 10 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month