Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
DETAILED ACTION
1. This action is in response to the amendment filed 8/7/2026.
2. Claims 1 and 4-20 have been examined and are pending in the application.
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
3. Claims 1 and 4-20 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Zhao U.S Patent No. 10,728,091.
As to claim 1, Zhao teaches a network (Fig. 1 and associated specifications) for communicatively coupling a plurality of graphics processing units (GPUs) (GPU devices 164, Fig. 1 and associated specifications), the network comprising:
a plurality of intra-cluster switches each of which connects two or more GPUs within one cluster (…two separate switches 305-1 and 305-2. The switch 305-1 enables direct communication between the CPU 303 and the GPUO and GPU1 devices, and the switch 305-2 enables direct communication between the CPU 304 and the GPU2 and GPU3 devices…, lines 13-17 column 14); and
a plurality of inter-cluster switches each to connect at least two different sets of GPUs (…reporting agents may run on switch devices that are configured within the backbone networking infrastructure of the service platform network 150…, lines 22-25 column 8) organized in two different clusters (GPU server nodes 160-1, 160-2, 160-n…, Fig. 1 and associated specifications), each of a set of GPUs having one or more endpoint interface (EPI) to connect the GPU to an intra-cluster switch and inter-cluster switch, each EPI connected to each intra- or inter-cluster switch through at least one physical link, wherein each EPI in at least a subset of EPIs also used to connect at least one inter-cluster switch to an inter-cluster switch (…A reporting agent 162 executing on a given GPU server node may report computing resource information to service control 140 such as: (i) GPU model and usage (e.g., NVidia P100, P40, etc.); (ii) intra-node bus topology information (PCIe, NVLink, NUMA/QPI, etc.); and (iii) inter-node connection information (e.g., NIC, SmartNIC, RDMA- enabled NIC, switch, 10/25/40/100GbE, port number, NUMA node connection, RDMA-enabled or not, etc…, lines 25-32 column 8).
As to claim 4, Zhao further teaches each EPI in the subset of EPIs includes an embedded switch to connect to at least one inter-cluster switch and one intra-cluster switch (…A reporting agent 162 executing on a given GPU server node may report computing resource information to service control 140 such as: (i) GPU model and usage (e.g., NVidia P100, P40, etc.); (ii) intra-node bus topology information (PCIe, NVLink, NUMA/QPI, etc.); and (iii) inter-node connection information (e.g., NIC, SmartNIC, RDMA- enabled NIC, switch, 10/25/40/100GbE, port number, NUMA node connection, RDMA-enabled or not, etc…, lines 25-32 column 8).
As to claim 5, Zhao further teaches each GPU includes an EPI that connects at least one inter- cluster switch to an inter-cluster switch (…inter-node topology information for a given server node can include port numbers of the servers, the type of network interface circuitry (and number of interface cards) that a given server utilizes to connect to other servers (and network components) including, but not limited to, network interface controllers (NICs) (e.g. SmartNICs, RDMA-enabled NICs), Host Bus Adapter (HBA) cards, Host Channel Adapter (HCA) cards…, line 65 column 6 to line 4 column 7).
As to claim 6, Zhao further teaches not every GPU includes an EPI that connects at least one inter-cluster switch to an inter-cluster switch (…inter-node topology information for a given server node can include port numbers of the servers, the type of network interface circuitry (and number of interface cards) that a given server utilizes to connect to other servers (and network components) including, but not limited to, network interface controllers (NICs) (e.g. SmartNICs, RDMA-enabled NICs), Host Bus Adapter (HBA) cards, Host Channel Adapter (HCA) cards…, line 65 column 6 to line 4 column 7)..
As to claim 7, Zhao further teaches the plurality of GPUs perform computations for a plurality of operations of a distributed application in order to collectively execute the distributed application (…dynamically schedules and provisions hardware accelerator resources (e.g., GPU resources) for pending jobs over one or more of the GPU server nodes 160-1, 160-2, ..., 160-n in the GPU server cluster 160 to execute HPC workloads associated with service requests received from the client systems 110…, lines 36-41 column 5).
As to claim 8, Zhao further teaches each GPU sharing a result of at least one computation with another GPU by sending the result in a plurality of segments (…A parameter server framework can implement parallel processing across the worker server nodes for deep learning application using data parallelism programming models. With data parallelism, each worker server node has access to a complete copy of a given deep learning model, but each worker server node operates on a different portion of the overall dataset, wherein the computation results from each worker server node are combined by the parameter server nodes. For neural networks, data parallelism involves each executing thread using the same weights (model parameters), but with each executing thread processing different mini-batches of data, wherein processing results (e.g., gradients) are synchronized (e.g., averaged) after each processing iteration of a mini-batch dataset. For example, in a parameter server framework, each worker GPU will compute a gradient on its subset of the minibatch, and then each worker GPU sends its computed gradient to a single parameter server, which takes the average of all the gradients, and sends the computed average back to the worker GPU devices…, lines 41-60 column 10) that are stored in a plurality of payloads of a plurality of data messages in a data message flow from the particular GPU to the other GPU, each data message flow of each GPU transmitted through one or more EPIs of the GPU (…The GPU devices can communicate, point-to-point, using a communication protocol such the known Message Passing Interface (MPI) communication protocol…, lines 6-9 column 13).
As to claim 9, Zhao further teaches the connection between the intra- and inter-cluster switches through the EPIs allows the network to forego using GPUs for transitioning between intra- and inter- cluster switches (…A reporting agent 162 executing on a given GPU server node may report computing resource information to service control 140 such as: (i) GPU model and usage (e.g., NVidia P100, P40, etc.); (ii) intra-node bus topology information (PCIe, NVLink, NUMA/QPI, etc.); and (iii) inter-node connection information (e.g., NIC, SmartNIC, RDMA- enabled NIC, switch, 10/25/40/100GbE, port number, NUMA node connection, RDMA-enabled or not, etc…, lines 25-32 column 8). .
As to claim 10, Zhao further teaches each intra-cluster switch is associated with a rack of GPUs (Fig. 3A and associated specifications).
As to claim 11, note the discussions of claims 1 and 7 above.
As to claims 12-17, note the discussions of claims 4-6, 9, 8 and 10 above, respectively.
As to claim 18, Zhao teaches an endpoint interface (EPI) for connecting a particular graphics processing unit (GPU) to a plurality of other GPUs through a network (Fig. 1 and associated specifications), the EPI comprising:
a network interface controller (NIC) for retrieving results of computations performed by the particular GPU (…A reporting agent 162 executing on a given GPU server node may report computing resource information to service control 140 such as: (i) GPU model and usage (e.g., NVidia P100, P40, etc.); (ii) intra-node bus topology information (PCIe, NVLink, NUMA/QPI, etc.); and (iii) inter-node connection information (e.g., NIC, SmartNIC, RDMA- enabled NIC, switch, 10/25/40/100GbE, port number, NUMA node connection, RDMA-enabled or not, etc…, lines 25-32 column 8); and
an embedded switch for interfacing with the network to forward data messages through the network that carry the results of the computations performed by the particular GPU (…A parameter server framework can implement parallel processing across the worker server nodes for deep learning application using data parallelism programming models. With data parallelism, each worker server node has access to a complete copy of a given deep learning model, but each worker server node operates on a different portion of the overall dataset, wherein the computation results from each worker server node are combined by the parameter server nodes. For neural networks, data parallelism involves each executing thread using the same weights (model parameters), but with each executing thread processing different mini-batches of data, wherein processing results (e.g., gradients) are synchronized (e.g., averaged) after each processing iteration of a mini-batch dataset. For example, in a parameter server framework, each worker GPU will compute a gradient on its subset of the minibatch, and then each worker GPU sends its computed gradient to a single parameter server, which takes the average of all the gradients, and sends the computed average back to the worker GPU devices…, lines 41-60 column 10).
As to claim 19, Zhao further teaches the GPUs are organized into a plurality of groups, the GPUs within each group are connected through at least one intra-group forwarding element (…two separate switches 305-1 and 305-2. The switch 305-1 enables direct communication between the CPU 303 and the GPUO and GPU1 devices, and the switch 305-2 enables direct communication between the CPU 304 and the GPU2 and GPU3 devices…, lines 13-17 column 14);
the GPUs in different groups are connected through at least one inter-group forwarding element (…reporting agents may run on switch devices that are configured within the backbone networking infrastructure of the service platform network 150…, lines 22-25 column 8), and
the embedded switch connecting at least one intra-group forwarding element to at least one inter-group forwarding element (…A reporting agent 162 executing on a given GPU server node may report computing resource information to service control 140 such as: (i) GPU model and usage (e.g., NVidia P100, P40, etc.); (ii) intra-node bus topology information (PCIe, NVLink, NUMA/QPI, etc.); and (iii) inter-node connection information (e.g., NIC, SmartNIC, RDMA- enabled NIC, switch, 10/25/40/100GbE, port number, NUMA node connection, RDMA-enabled or not, etc…, lines 25-32 column 8).
As to claim 20, Zhao further teaches the embedded switch further forwarding receive data messages through the network that are destined to the particular GPU and passing the data messages to the NIC to forward to the particular GPU (…The GPU devices can communicate, point-to-point, using a communication protocol such the known Message Passing Interface (MPI) communication protocol…, lines 6-9 column 13).
Response to Arguments
4. Applicant’s arguments have been fully considered but they are not persuasive.
Applicant argues Zhao reference does not teach the limitations of claim 1 (Remarks, page 6). In response, while the applicant argues Zhao reference does not teach these limitations, the applicant does not disclose any detail of how the cited portions from Zhao reference (as disclosed in the claim rejection above) do not meet the claim limitations. Therefore, it is unclear how the reference does not teach the claim invention as argued by the applicant.
Conclusion
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any extension fee pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Andy Ho whose telephone number is (571) 272-3762. A voice mail service is also available for this number. The examiner can normally be reached on Monday – Friday, 8:30 am – 5:00 pm.
If attempts to reach the examiner by telephone are unsuccessful, the examiner's supervisor, Kevin Young can be reached on (571) 270-3180.
Any inquiry of a general nature or relating to the status of this application or proceeding should be directed to the receptionist whose telephone number is 571-272-2100.
Any response to this action should be mailed to:
Commissioner for Patents
P.O Box 1450
Alexandria, VA 22313-1450
Or fax to:
AFTER-FINAL faxes must be signed and sent to (571) 273 - 8300.
OFFICAL faxes must be signed and sent to (571) 273 - 8300.
NON OFFICAL faxes should not be signed, please send to (571) 273 – 3762
/Andy Ho/
Primary Examiner
Art Unit 2194