Prosecution Insights
Last updated: October 02, 2026
Application No. 19/257,848

ABSTRACTION LAYERS FOR SCALABLE DISTRIBUTED MACHINE LEARNING

Non-Final OA §103§DOUBLEPATENT
Filed
Jul 02, 2025
Priority
Apr 10, 2017 — continuation of 11/094,029 +2 more
Examiner
PUENTES, DANIEL CALRISSIAN
Art Unit
2836
Tech Center
2800 — Semiconductors & Electrical Systems
Assignee
Intel Corporation
OA Round
1 (Non-Final)
89%
Grant Probability
Favorable
1-2
OA Rounds
10m
Est. Remaining
92%
With Interview

Examiner Intelligence

Grants 89% — above average
89%
Career Allowance Rate
830 granted / 936 resolved
+20.7% vs TC avg
Minimal +3% lift
Without
With
+3.1%
Interview Lift
resolved cases with interview
Fast prosecutor
2y 1m
Avg Prosecution
19 currently pending
Career history
958
Total Applications
across all art units

Statute-Specific Performance

§101
1.1%
-38.9% vs TC avg
§103
42.1%
+2.1% vs TC avg
§102
32.3%
-7.7% vs TC avg
§112
18.5%
-21.5% vs TC avg
Black line = Tech Center average estimate • Based on career data from 936 resolved cases

Office Action

§103 §DOUBLEPATENT
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Double Patenting The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969). A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on nonstatutory double patenting provided the reference application or patent either is shown to be commonly owned with the examined application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. See MPEP § 717.02 for applications subject to examination under the first inventor to file provisions of the AIA as explained in MPEP § 2159. See MPEP § 2146 et seq. for applications not subject to examination under the first inventor to file provisions of the AIA . A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b). The filing of a terminal disclaimer by itself is not a complete reply to a nonstatutory double patenting (NSDP) rejection. A complete reply requires that the terminal disclaimer be accompanied by a reply requesting reconsideration of the prior Office action. Even where the NSDP rejection is provisional the reply must be complete. See MPEP § 804, subsection I.B.1. For a reply to a non-final Office action, see 37 CFR 1.111(a). For a reply to final Office action, see 37 CFR 1.113(c). A request for reconsideration while not provided for in 37 CFR 1.113(c) may be filed after final for consideration. See MPEP §§ 706.07(e) and 714.13. The USPTO Internet website contains terminal disclaimer forms which may be used. Please visit www.uspto.gov/patent/patents-forms. The actual filing date of the application in which the form is filed determines what form (e.g., PTO/SB/25, PTO/SB/26, PTO/AIA /25, or PTO/AIA /26) should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto-processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to www.uspto.gov/patents/apply/applying-online/eterminal-disclaimer. Claims 21-40 are rejected on the ground of nonstatutory double patenting as being unpatentable over claims 1-20 of U.S. Patent No. 12,387,287. Although the claims at issue are not identical, they are not patentably distinct from each other because the patent teaches all of the limitations of the current application. Although only claims 21, 31 and 37 are shown below, it can be understood that similar reasons apply for the remaining claims. For claim 37, the patent does not explicitly teach a memory and fabric interface. However, Examiner takes official notice that it is notoriously old and well-known for processors to use these components as they are required for operation. Before the effective filing date of the invention it would have been obvious to one of ordinary skill in the art to use the patent’s processor with a memory and fabric interface as these conventional components that are well known by those of ordinary skill in the art to be used with a processor for implementing a computer. US 19/257,848 (current application) US 12,387,287 (patent) 21. A non-transitory machine readable medium storing instructions which, when executed by one or more processors, cause the one or more processors to perform operations comprising: creating a session for a neural network using an application programming interface (API) of a machine learning library, wherein the API operates on a session object, and wherein the session of the neural network is associated with communication operations and compute operations to be performed by the neural network in a distributed compute environment; specifying, using the API, a distribution object for the session that indicates a degree of parallelism of the neural network; enabling, via the API, the communication operations to be performed between multiple compute nodes of the distributed compute environment used to perform the compute operations for one or more layers of the neural network; and performing a machine learning framework workflow of the neural network in the distributed compute environment with the compute operations using API calls of the API. 1. A non-transitory machine-readable storage medium having stored thereon executable computer program instructions that, when executed by one or more processors, cause the one or more processors to perform operations comprising: defining a distributed compute environment for a neural network using a machine learning scaling library (MLSL) application programming interface (API), wherein the MLSL API operates on: a session object to represent a collection of operation objects, the operation objects created for a communication session associated with communication operations to be performed by the distributed compute environment; and a distribution object that indicates a number of partitions for data parallelism and a number of partitions for model parallelism of the neural network; enabling, via the MLSL API, the communication operations to be performed between multiple compute nodes of the distributed compute environment used to perform compute operations for one or more layers of the neural network; and optimizing the communication operations by enabling specification of resources for managing the communication operations; wherein the neural network is defined using operations of a neural network domain specific framework. 7. The non-transitory machine-readable storage medium as in claim 1, wherein the operations further comprising: performing a machine learning framework workflow; and while performing the machine learning framework workflow, automatically exchanging gradients with respect to machine learning parameters. 31. A method comprising: creating a session for a neural network using an application programming interface (API) of a machine learning library, wherein the API operates on a session object, and wherein the session of the neural network is associated with communication operations and compute operations to be performed by the neural network in a distributed compute environment; specifying, using the API, a distribution object for the session that indicates a degree of parallelism of the neural network; enabling, via the API, the communication operations to be performed between multiple compute nodes of the distributed compute environment used to perform the compute operations for one or more layers of the neural network; and performing a machine learning framework workflow of the neural network in the distributed compute environment with the compute operations using API calls of the API. 10. A method comprising: defining a distributed compute environment for a neural network using a machine learning scaling library (MLSL) application programming interface (API), wherein the MLSL API operates on: a session object to represent a collection of operation objects, the operation objects created for a communication session associated with communication operations to be performed by the distributed compute environment; and a distribution object that indicates a number of partitions for data parallelism and a number of partitions for model parallelism of the neural network; enabling, via the MLSL API, the communication operations to be performed between multiple compute nodes of the distributed compute environment used to perform compute operations for one or more layers of the neural network; and optimizing the communication operations by enabling specification of resources for managing the communication operations; wherein the neural network is defined using operations of a neural network domain specific framework. 15. The method as in claim 10, wherein the operations further comprising: performing a machine learning framework workflow; and while performing the machine learning framework workflow, automatically exchanging gradients with respect to machine learning parameters. 37. A system comprising: a system memory to store a set of trainable machine learning parameters and a machine learning library to facilitate data transmission during distributed training of the neural network; a fabric interface to enable transmission and receipt of data associated with the set of trainable machine learning parameters; and a general-purpose graphics processor to: create a session for the neural network using an application programming interface (API) of the machine learning library, wherein the API operates on a session object, and wherein the session of the neural network is associated with communication operations and compute operations to be performed by the neural network in a distributed compute environment; specify, using the API, a distribution object for the session that indicates a degree of parallelism of the neural network; enable, via the API, the communication operations to be performed between multiple compute nodes of the distributed compute environment used to perform the compute operations for one or more layers of the neural network; and perform the machine learning framework workflow of the neural network in the distributed compute environment with the compute operations using API calls of the API. 16. An apparatus comprising: processor hardware circuitry to: define a distributed compute environment for a neural network using a machine learning scaling library (MLSL) application programming interface (API), wherein the MLSL API operates on: a session object to represent a collection of operation objects, the operation objects created for a communication session associated with communication operations to be performed by the distributed compute environment; and a distribution object that indicates a number of partitions for data parallelism and a number of partitions for model parallelism of the neural network; enable, via the MLSL API, the communication operations to be performed between multiple compute nodes of the distributed compute environment used to perform compute operations for one or more layers of the neural network; and optimize the communication operations by enabling specification of resources for managing the communication operations; wherein the neural network is defined using operations of a neural network domain specific framework. 17. The apparatus as in claim 16, wherein the processor hardware circuitry is further to: create a global view of the communication operations to be performed between the multiple compute nodes of the distributed compute environment; and utilize the global view to track an overlap of the compute operations and the communication operations. 18. The apparatus as in claim 17, wherein each operation object of the collection of operation objects is to store machine learning parameters for the communication session associated with the communication operations to be performed. Claims 21-40 are rejected on the ground of nonstatutory double patenting as being unpatentable over claims 1-28 of U.S. Patent No. 11,094,029. Although the claims at issue are not identical, they are not patentably distinct from each other because the patent teaches all of the limitations of the current application. Claims 1 and 8-10 of the patent teaches all of the limitations of claim 25 of the application except for API calls. However, the patent requires the use of session objects communicated between nodes of a distributed compute system. Examiner takes official notice that it is notoriously old and well-known to use API calls to retrieve data or trigger an action from an external source. Before the effective filing date of the invention it would have been obvious to one of ordinary skill in the art to implement the patent such that session objects are communicated between nodes of a distributed compute system via API calls since the particular known technique was recognized as part of the ordinary capabilities of one skilled in the art. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claim(s) 21-40 is/are rejected under 35 U.S.C. 103 as being unpatentable over Das et al (NPL: “Distributed Deep Learning Using Synchronous Stochastic Gradient Descent”) in view of PyTorch (NPL: PyTorch Github pages). For claim 21, Das teaches operations comprising: a session of a neural network is associated with communication operations (determining communication volume, §3.2) and compute operations (forward and backpropagate operations are partitioned across multiple minibatches into jobs, then equally distributed across different threads, §2.5) to be performed by the neural network in a distributed compute environment (PCL-DNN Software Framework, §4, and the corresponding hardware, §5); specifying a degree of parallelism of the neural network (data parallelism, model parallelism and hybrid parallelism, §3.1-§3.3); enabling the communication operations to be performed between multiple compute nodes of the distributed compute environment used to perform the compute operations for one or more layers of the neural network (§2.5, §5.2, §5.3); and performing a machine learning framework workflow of the neural network PCL-DNN Software Framework, §4) in the distributed compute environment (e.g., multi-node experiment, §5). Das fails to teach: a non-transitory machine readable medium storing instructions which, when executed by one or more processors, cause the one or more processors to perform operations comprising: creating a session for a neural network using an application programming interface (API) of a machine learning library, wherein the API operates on a session object; specifying, using the API, a distribution object for the session that indicates a degree of parallelism of the neural network; enabling, via the API, the communication operations to be performed between multiple compute nodes of the distributed compute environment used to perform the compute operations for one or more layers of the neural network;; and performing a machine learning framework workflow of the neural network in the distributed compute environment with the compute operations using API calls of the API. It is noted that Das teaches that his invention can be used with machine learning frameworks such as TensorFlow, FireCaffe and DeepImage (¶1 of §1). PyTorch is a machine learning framework that teaches: creating a session object (custom nn modules comprising a plurality of nn modules, “nn module” section from tutorials/beginner_source/pytorch_with_examples.rst page, hereinafter PWE) to represent a collection of operation objects (e.g., different neural network layers and loss functions, “PyTorch:nn” from PWE), the operation objects created for a communication session associated with the communication operations to be performed by the distributed compute environment (e.g., scatter and gather, “DataParallel” from tutorials/_downloads/parallelism_tutorial.ipynb page, hereinafter PT; and creating a distribution object (DataParallel Model class and DistributedModel class, “Multi-GPU examples” from PT) that indicates a number of partitions for data parallelism and a number of partitions for model parallelism (as understood by the code segments in PT). Before the effective filing date of the invention it would have been obvious to one of ordinary skill in the art to implement Das’ method of determining which strategy (model parallelism, data parallelism or hybrid parallelism) is best suited to reduce time in training a neural network by using PyTorch since PyTorch enables a user create custom parallelism (via torch.nn.DataParallel, PT) and build neural networks using a modular approach (“PyTorch: nn” and “PyTorch: Custom nn Modules”, PWE). The combination of Das and PyTorch teaches: a non-transitory machine-readable storage medium (memory inherently required to run PyTorch and the computations and communication of Das) having stored thereon executable computer program instructions that, when executed by one or more processors (§5, Das), cause the one or more processors to perform operations comprising said method (as understood by the combination of references); creating a session for a neural network using an API of a machine learning library (PyTorch Variables and Tensors both use APIs, see “PyTorch: Variables and autograd”) specifying, using the API, a distribution object for the session (DataParallel Model class and DistributedModel class, “Multi-GPU examples” from PT) that indicates a degree of parallelism of the neural network (replicate, scatter, see page 2); enabling, via the API (via gather, parallel_apply), the communication operations to be performed between multiple compute nodes of the distributed compute environment used to perform the compute operations for one or more layers of the neural network (as understood by pages 2-3 of PyTorch); performing a machine learning framework workflow of the neural network in the distributed compute environment with the compute operations using API calls of the API (see “Multi-GPU examples” of pages 2-3, PyTorch). For claim 22, the combination of Das and PyTorch as defined above teaches the limitations of claim 21 and Das further teaches: creating a global view of the communication operations to be performed (task graph, §2) between multiple compute nodes (e.g., determining the amount of data received for forward pass and the amount of data sent out by the previous layer, ¶1-3 of §3.2) of the distributed compute environment; and utilizing the global view to track an overlap of the compute operations and the communication operations (§2 discusses using a task graph to understand the compute and memory bandwidth needs and to identify optimal threading). For claim 23, the combination of Das and PyTorch as defined above teaches the limitations of claim 22 and further teaches: the session represents a collection of operation objects is to store machine learning parameters for the communication session associated with the communications operations to be performed (as discussed in the rejection of claim 21). For claim 24, the combination of Das and PyTorch as defined above teaches the limitations of claim 23 and Das further teaches: the collection of operation objects in the session object are set with a same batch size for the communication session (§5.2 teaches minibatch sizes of 256 and 512). For claim 25, the combination of Das and PyTorch as defined above teaches the limitations of claim 21 and Das further teaches: the operations further comprise determining a type of parallelism to use for the compute operations for the one or more layers of the neural network (“model parallelism is better if…” ¶7 of §3.2, see also comparison of data-parallelism, model-parallelism and hybrid parallelism in §3.3). For claim 26, the combination of Das and PyTorch as defined above teaches the limitations of claim 25 and Das further teaches: the type of parallelism comprises one or more of the data parallelism (§3.1), the model parallelism (§3.2), or a hybrid of the data parallelism and the model parallelism (§3.3). For claim 27, the combination of Das and PyTorch as defined above teaches the limitations of claim 21 and Das further teaches: the operations further comprise optimizing the communication operations by enabling specification of resources for managing the communication operations (blocked loop structure of Algorithm 2 optimizes bandwidth and performance, §2.3). For claim 28, the combination of Das and PyTorch as defined above teaches the limitations of claim 21 and Das further teaches: the operations further comprise: while performing the machine learning framework workflow, automatically exchanging gradients with respect to machine learning parameters (“the data handling library…executes the various computations – forward propagation, backpropagation and weight updates – on the underlying hardware”, ¶1 of §4); and updating the machine learning parameters based on the machine learning framework workflow (weight updates, ¶1 of §4). For claim 29, the combination of Das and PyTorch as defined above teaches the limitations of claim 21 and Das further teaches: performing the machine learning framework workflow comprises performing forward propagation computation to generate a set of activation data and performing a backward propagation computation to determine a gradient with respect to a set of trainable machine learning parameters (“the data handling library…executes the various computations – forward propagation, backpropagation and weight updates – on the underlying hardware”, ¶1 of §4). For claim 30, the combination of Das and PyTorch as defined above teaches the limitations of claim 21 and Das further teaches: the degree of parallelism comprises a number of partitions for data parallelism (§3.1) and a number of partitions for model parallelism (§3.2). For claim 31, Das teaches a method comprising: creating a session for a neural network using communication operations (determining communication volume, §3.2) and compute operations (forward and backpropagate operations are partitioned across multiple minibatches into jobs, then equally distributed across different threads, §2.5) to be performed by the neural network in a distributed compute environment (PCL-DNN Software Framework, §4, and the corresponding hardware, §5); specifying a degree of parallelism of the neural network (data parallelism, model parallelism and hybrid parallelism, §3.1-§3.3); enabling the communication operations to be performed between multiple compute nodes of the distributed compute environment used to perform the compute operations for one or more layers of the neural network (§2.5, §5.2, §5.3); and performing a machine learning framework workflow of the neural network PCL-DNN Software Framework, §4) in the distributed compute environment (e.g., multi-node experiment, §5). Das fails to teach: using an API of a machine learning library as claimed. It is noted that Das teaches that his invention can be used with machine learning frameworks such as TensorFlow, FireCaffe and DeepImage (¶1 of §1). PyTorch is a machine learning framework that teaches: creating a session object (custom nn modules comprising a plurality of nn modules, “nn module” section from tutorials/beginner_source/pytorch_with_examples.rst page, hereinafter PWE) to represent a collection of operation objects (e.g., different neural network layers and loss functions, “PyTorch:nn” from PWE), the operation objects created for a communication session associated with the communication operations to be performed by the distributed compute environment (e.g., scatter and gather, “DataParallel” from tutorials/_downloads/parallelism_tutorial.ipynb page, hereinafter PT; and creating a distribution object (DataParallel Model class and DistributedModel class, “Multi-GPU examples” from PT) that indicates a number of partitions for data parallelism and a number of partitions for model parallelism (as understood by the code segments in PT). Before the effective filing date of the invention it would have been obvious to one of ordinary skill in the art to implement Das’ method of determining which strategy (model parallelism, data parallelism or hybrid parallelism) is best suited to reduce time in training a neural network by using PyTorch since PyTorch enables a user create custom parallelism (via torch.nn.DataParallel, PT) and build neural networks using a modular approach (“PyTorch: nn” and “PyTorch: Custom nn Modules”, PWE). The combination of Das and PyTorch teaches: a non-transitory machine-readable storage medium (memory inherently required to run PyTorch and the computations and communication of Das) having stored thereon executable computer program instructions that, when executed by one or more processors (§5, Das), cause the one or more processors to perform operations comprising said method (as understood by the combination of references); creating a session for a neural network using an API of a machine learning library (PyTorch Variables and Tensors both use APIs, see “PyTorch: Variables and autograd”) specifying, using the API, a distribution object for the session (DataParallel Model class and DistributedModel class, “Multi-GPU examples” from PT) that indicates a degree of parallelism of the neural network (replicate, scatter, see page 2); enabling, via the API (via gather, parallel_apply), the communication operations to be performed between multiple compute nodes of the distributed compute environment used to perform the compute operations for one or more layers of the neural network (as understood by pages 2-3 of PyTorch); performing a machine learning framework workflow of the neural network in the distributed compute environment with the compute operations using API calls of the API (see “Multi-GPU examples” of pages 2-3, PyTorch). For claim 32, the combination of Das and PyTorch as defined above teaches the limitations of claim 31 and Das further teaches: creating a global view of the communication operations to be performed (task graph, §2) between multiple compute nodes (e.g., determining the amount of data received for forward pass and the amount of data sent out by the previous layer, ¶1-3 of §3.2) of the distributed compute environment; and utilizing the global view to track an overlap of the compute operations and the communication operations (§2 discusses using a task graph to understand the compute and memory bandwidth needs and to identify optimal threading). For claim 33, the combination of Das and PyTorch as defined above teaches the limitations of claim 32 and further teaches: the session represents a collection of operation objects is to store machine learning parameters for the communication session associated with the communications operations to be performed (as discussed in the rejection of claim 31). For claim 34, the combination of Das and PyTorch as defined above teaches the limitations of claim 21 and Das further teaches: the operations further comprise determining a type of parallelism to use for the compute operations for the one or more layers of the neural network, and wherein the type of parallelism comprises one or more of the data parallelism, the model parallelism or a hybrid of the data parallelism and the model parallelism (“model parallelism is better if…” ¶7 of §3.2, see also comparison of data-parallelism, model-parallelism and hybrid parallelism in §3.3). For claim 35, the combination of Das and PyTorch as defined above teaches the limitations of claim 31 and Das further teaches: while performing the machine learning framework workflow, automatically exchanging gradients with respect to machine learning parameters (“the data handling library…executes the various computations – forward propagation, backpropagation and weight updates – on the underlying hardware”, ¶1 of §4); and updating the machine learning parameters based on the machine learning framework workflow (weight updates, ¶1 of §4). For claim 36, the combination of Das and PyTorch as defined above teaches the limitations of claim 31 and Das further teaches: the degree of parallelism comprises a number of partitions for data parallelism (§3.1) and a number of partitions for model parallelism (§3.2). For claim 37, Das teaches A system comprising: a system memory to store a set of trainable machine learning parameters and a machine learning library to facilitate data transmission during distributed training of the neural network (§5); a fabric interface to enable transmission and receipt of data associated with the set of trainable machine learning parameters (Cray Aries high speed ”dragonfly” topology interconnect, §5); and a general-purpose graphics processor (CPU, §5) to: create a session for a neural network using communication operations (determining communication volume, §3.2) and compute operations (forward and backpropagate operations are partitioned across multiple minibatches into jobs, then equally distributed across different threads, §2.5) to be performed by the neural network in a distributed compute environment (PCL-DNN Software Framework, §4, and the corresponding hardware, §5); specify a degree of parallelism of the neural network (data parallelism, model parallelism and hybrid parallelism, §3.1-§3.3); enable the communication operations to be performed between multiple compute nodes of the distributed compute environment used to perform the compute operations for one or more layers of the neural network (§2.5, §5.2, §5.3); and perform a machine learning framework workflow of the neural network PCL-DNN Software Framework, §4) in the distributed compute environment (e.g., multi-node experiment, §5). Das fails to teach: using an API of a machine learning library as claimed. It is noted that Das teaches that his invention can be used with machine learning frameworks such as TensorFlow, FireCaffe and DeepImage (¶1 of §1). PyTorch is a machine learning framework that teaches: creating a session object (custom nn modules comprising a plurality of nn modules, “nn module” section from tutorials/beginner_source/pytorch_with_examples.rst page, hereinafter PWE) to represent a collection of operation objects (e.g., different neural network layers and loss functions, “PyTorch:nn” from PWE), the operation objects created for a communication session associated with the communication operations to be performed by the distributed compute environment (e.g., scatter and gather, “DataParallel” from tutorials/_downloads/parallelism_tutorial.ipynb page, hereinafter PT; and creating a distribution object (DataParallel Model class and DistributedModel class, “Multi-GPU examples” from PT) that indicates a number of partitions for data parallelism and a number of partitions for model parallelism (as understood by the code segments in PT). Before the effective filing date of the invention it would have been obvious to one of ordinary skill in the art to implement Das’ method of determining which strategy (model parallelism, data parallelism or hybrid parallelism) is best suited to reduce time in training a neural network by using PyTorch since PyTorch enables a user create custom parallelism (via torch.nn.DataParallel, PT) and build neural networks using a modular approach (“PyTorch: nn” and “PyTorch: Custom nn Modules”, PWE). The combination of Das and PyTorch teaches: a non-transitory machine-readable storage medium (memory inherently required to run PyTorch and the computations and communication of Das) having stored thereon executable computer program instructions that, when executed by one or more processors (§5, Das), cause the one or more processors to perform operations comprising said method (as understood by the combination of references); creating a session for a neural network using an API of a machine learning library (PyTorch Variables and Tensors both use APIs, see “PyTorch: Variables and autograd”) specifying, using the API, a distribution object for the session (DataParallel Model class and DistributedModel class, “Multi-GPU examples” from PT) that indicates a degree of parallelism of the neural network (replicate, scatter, see page 2); enabling, via the API (via gather, parallel_apply), the communication operations to be performed between multiple compute nodes of the distributed compute environment used to perform the compute operations for one or more layers of the neural network (as understood by pages 2-3 of PyTorch); performing a machine learning framework workflow of the neural network in the distributed compute environment with the compute operations using API calls of the API (see “Multi-GPU examples” of pages 2-3, PyTorch). For claim 38, the combination of Das and PyTorch as defined above teaches the limitations of claim 37 and further teaches: the session object represents a collection of operation objects that are to store machine learning parameters for the session (e.g., different neural network layers and loss functions, “PyTorch:nn” from PWE), and wherein the collection of operation objects of the session object are set with a same batch size for the session (§5.2 teaches minibatch sizes of 256 and 512). For claim 39, the combination of Das and PyTorch as defined above teaches the limitations of claim 37 and further teaches: the general-purpose graphics processor is further to determine a type of parallelism to use for the compute operations for the one or more layers of the neural network (“model parallelism is better if…” ¶7 of §3.2, see also comparison of data-parallelism, model-parallelism and hybrid parallelism in §3.3), and wherein the type of parallelism comprises one or more of the data parallelism (§3.1), the model parallelism (§3.2), or a hybrid of the data parallelism and the model parallelism (§3.3). For claim 40, the combination of Das and PyTorch as defined above teaches the limitations of claim 37 and further teaches: the general-purpose graphics processor is further to: while performing the machine learning framework workflow, automatically exchanging gradients with respect to machine learning parameters (“the data handling library…executes the various computations – forward propagation, backpropagation and weight updates – on the underlying hardware”, ¶1 of §4); and updating the machine learning parameters based on the machine learning framework workflow (weight updates, ¶1 of §4). Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure: Chapelle et al (US 2013/0290223) teaches a method for distributed machine learning. Any inquiry concerning this communication or earlier communications from the examiner should be directed to DANIEL CALRISSIAN PUENTES whose telephone number is (571)270-5070. The examiner can normally be reached M-F 9-6:30 (flex). Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Taelor Kim can be reached at (571) 270-7166. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /DANIEL C PUENTES/Primary Examiner, Art Unit 2836
Read full office action

Prosecution Timeline

Jul 02, 2025
Application Filed
Aug 11, 2026
Non-Final Rejection mailed — §103, §DOUBLEPATENT (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12750039
FINE TRIMMING OF A RADIO FREQUENCY GAIN BY MODULATING THE PERIPHERY OF A RADIO FREQUENCY SWITCH
3y 1m to grant Granted Sep 29, 2026
Patent 12749901
WIRELESS POWER TRANSMITTING AND CHARGING SYSTEM
2y 8m to grant Granted Sep 29, 2026
Patent 12744539
PHASE LOCKED LOOP CIRCUIT AND METHOD OF OPERATION THEREOF
2y 1m to grant Granted Sep 22, 2026
Patent 12738888
Safety Switch for Photovoltaic Systems
3y 6m to grant Granted Sep 15, 2026
Patent 12732103
POWER CONVERTER CONTROL LOOP GAIN ADAPTATION
2y 6m to grant Granted Sep 08, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
89%
Grant Probability
92%
With Interview (+3.1%)
2y 1m (~10m remaining)
Median Time to Grant
Low
PTA Risk
Based on 936 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month