DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Preliminary Amendment
The preliminary amendment filed on 2024/11/11 has been entered. Claims 1-20 are pending in the application.
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 2024/09/13; 2025/01/13; 2025/02/26 & 2025/04/30. The submission is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 102 that forms the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claims 10, 11, and 20 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Shen et al. (Shen), WO 2020/199560 A1, "AI Training Network and Method," published October 8, 2020 an English translation of which is available as US 2022/0012590 A1, published January 13, 2022.
As to independent Claim 10, Shen teaches a model training method, comprising:
"performing, by a first processor of a first node, model training in the first processor, to obtain first target data" — Shen discloses an AI training network formed by a server array that includes a server 11, a server 12, a server 13, and a server 14, and that the server 11 includes a CPU 111, a CPU 112, a GPU 113, and a GPU 114 (Shen, ¶¶ [0027], [0029], FIG. 1). Shen discloses that the master server 11 runs an AI training program, loads a training data set and a data flow diagram, and splits both into level-1 data subsets and level-1 data flow sub-diagrams that are sent to the servers, and that the CPUs of the servers split those again based on a quantity of GPUs into level-2 data subsets and level-2 data flow sub-diagrams which are sent to the corresponding GPUs, each GPU being instructed to perform calculation on a received level-2 data subset based on a received level-2 data flow sub-diagram (Shen, ¶¶ [0037]–[0039], FIG. 3, steps S11 and S12). Shen further discloses that the first graphics processing unit performs AI training calculation on a first dataset based on a first data flow diagram (Shen, ¶ [0006]), and that the training data set and the data flow diagram undertaken by the GPU 113 are referred to as a first training data set and a first data flow diagram, the calculation result of the GPU 113 being the training data set undertaken by the GPU 123 (Shen, ¶ [0043]). The recited first node reads on the server 11 of Shen, the recited first processor reads on the GPU 113 of Shen, and the recited first target data reads on the calculation result of the GPU 113 of Shen.
"sending, by the first processor, the first target data to a second processor of a second node through an optical transmission channel established by a micro-electro-mechanical system (MEMS)" — Shen discloses that the server 12 includes a GPU 123 and a GPU 124 (Shen, ¶ [0040], FIG. 1), and that the optical cross-connect is a micro-electro-mechanical system (MEMS) or a silicon photonics (SiP) (Shen, ¶ [0012]). Shen further discloses that the MEMS 15 includes a micro-electro-mechanical controller 150 and two reflection lens arrays, each lens array including a plurality of reflectors whose deflection angle is physically adjustable, that an electrical signal sent by the GPU 113 is converted into an optical signal and reaches a graphics processing unit after passing through a fiber channel 151, a reflector 152, a reflector 153, and a fiber channel 154, and that after the increased reflection angle of the reflector 153 is adjusted to 30° a channel between the GPU 113 and the GPU 123 is successfully established (Shen, ¶ [0042], FIG. 4). Shen further discloses that after the adjustment a reflection path 155-156-157 of the optical signal is established and channel switching between the GPU 113 and the GPU 123 is completed (Shen, ¶¶ [0049], [0050], FIG. 3, step S14). Shen further discloses that the calculation result may be immediately sent to the GPU 123 by using the optical path 155-156-157 (Shen, ¶ [0051], FIG. 3, step S15), and that once the GPU 113 completes the calculation the data may be immediately sent to the GPU 123 by using the MEMS 15 (Shen, ¶ [0052]). The recited second node reads on the server 12 of Shen, the recited second processor reads on the GPU 123 of Shen, and the recited optical transmission channel established by a MEMS reads on the reflection path 155-156-157 established through the MEMS 15 of Shen.
"wherein the MEMS is located between the first node and the second node" — Shen discloses that, where the input of the GPU 123 depends on the output of the GPU 113, the lens that needs to be adjusted is an optical cross-connect located between the GPU 123 and the GPU 113, namely, a MEMS 15 (Shen, ¶ [0042], FIG. 4). Shen further discloses that where the two GPUs that communicate with each other belong to different servers the communication needs to be performed by using a communication channel outside the servers, in other words by using the optical cross-connects of FIG. 1 (Shen, ¶ [0031]). The MEMS 15 of Shen is accordingly disposed between the server 11 containing the GPU 113 and the server 12 containing the GPU 123.
"and the first target data is for the second processor to adjust a parameter for model training in the second processor" — Shen discloses that the second graphics processing unit performs AI training calculation on the calculation result by using a second data flow diagram (Shen, ¶ [0006]), and that after receiving the calculation result of the GPU 113 the GPU 123 further performs calculation based on the data flow sub-diagram of the GPU 123 (Shen, ¶ [0051], FIG. 3, step S15). Shen further discloses that in a complete AI training process the steps of forward propagation, loss calculation, and backward propagation are iteratively performed until the calculation result converges to sufficient precision, the backward propagation propagating the loss backward level by level according to a chain derivation rule to obtain gradients of all parameters (Shen, ¶ [0028]), that the respective learning results of the accelerators are combined and distributed to each accelerator again before a next iteration is performed (Shen, ¶ [0004]), and that a task of the AI training is to find an assignment of each parameter in the neural network architecture (Shen, ¶ [0003]).
As to dependent Claim 11, the limitations of claim 10 from which claim 11 depends are rejected under the same rationale set forth above with respect to claim 10.
Regarding the additional limitations of claim 11, Shen teaches:
"sending, by the first processor, the first target data to the second processor through the optical transmission channel established by the MEMS and an intra-node channel" — Shen discloses that an electrical signal sent by the GPU 113 is converted into an optical signal and reaches a graphics processing unit after passing through a fiber channel 151, a reflector 152, a reflector 153, and a fiber channel 154, and that after the reflection angle of the reflector 153 is adjusted to 30° a channel between the GPU 113 and the GPU 123 is successfully established (Shen, ¶ [0042], FIG. 4). Shen further discloses that after the adjustment the reflection path 155-156-157 of the optical signal is established (Shen, ¶ [0050]), that the calculation result may be immediately sent to the GPU 123 by using the optical path 155-156-157 (Shen, ¶ [0051], FIG. 3, step S15), and that the data may be immediately sent to the GPU 123 by using the MEMS 15 (Shen, ¶ [0052]). The recited optical transmission channel established by the MEMS reads on the reflection path 155-156-157 of Shen, and the recited intra-node channel reads on the fiber channel 151 and the fiber channel 154 of Shen, which are traversed by the same transmission in addition to the reflector path within the MEMS 15.
"wherein the intra-node channel comprises: a channel that is in the first node and that is between the first processor and the MEMS" — Shen discloses that the MEMS 15 includes a micro-electro-mechanical controller 150 and two reflection lens arrays, and that the optical signal converted from the electrical signal sent by the GPU 113 passes through the fiber channel 151 before reaching the reflector 152 of the MEMS 15 (Shen, ¶ [0042], FIG. 4). Shen further discloses that where the two GPUs that communicate with each other belong to the same server the communication may be performed by using a bus inside the server (Shen, ¶ [0031]), and that communication between the CPU 111 and the GPU 113 within the server 11 is performed by using a peripheral component interconnect express (PCIe) bus (Shen, ¶ [0029]). The recited channel that is in the first node and that is between the first processor and the MEMS reads on the fiber channel 151 of Shen, which extends within the server 11 from the GPU 113 to the reflector 152 of the MEMS 15.
"and/or a channel that is in the second node and that is between the second processor and the MEMS" — Shen discloses that the optical signal reaches the destination graphics processing unit after passing through the reflector 153 and the fiber channel 154 (Shen, ¶ [0042], FIG. 4). The recited channel that is in the second node and that is between the second processor and the MEMS reads on the fiber channel 154 of Shen, which extends within the server 12 from the reflector 153 of the MEMS 15 to the GPU 123. The claim recites the two channels in the alternative, and Shen discloses both.
As to independent Claim 20, the recitations "perform model training in the at least one processor, to obtain first target data" and "send the first target data to a second processor of a second node through an optical transmission channel established by a micro-electro-mechanical system (MEMS), wherein the MEMS is located between the first node and the second node, and the first target data is for the second processor to adjust a parameter for model training in the second processor" correspond to the recitations of claim 10 and are rejected under the same rationale set forth above with respect to claim 10.
Regarding the remaining limitations of claim 20, Shen teaches:
"A chip of a first node, comprising:" — Shen discloses an AI training network formed by a server array that includes a server 11, a server 12, a server 13, and a server 14 (Shen, ¶ [0027], FIG. 1), and discloses that a server includes a central processing unit (CPU) and a graphics processing unit (GPU), the server 11 including a CPU 111, a CPU 112, a GPU 113, and a GPU 114 (Shen, ¶ [0029], FIG. 1). Shen further discloses that data is transmitted between GPU chips more frequently as the scale of an accelerator cluster with high computing power becomes larger, and that reducing the duration of establishing an optical channel and transmitting the data between the GPU chips is an urgent problem (Shen, ¶ [0005]). The recited first node reads on the server 11 of Shen, and the recited chip of a first node reads on the GPU 113 of Shen, which Shen identifies as a GPU chip.
"at least one processor" — Shen discloses that the server 11 includes a GPU 113 and a GPU 114 (Shen, ¶ [0029], FIG. 1), and that the GPU is more suitable for the iterative operation of AI training than a CPU and is therefore more widely applied to AI training (Shen, ¶ [0004]). The recited at least one processor reads on the GPU 113 of Shen.
"an interface, wherein the interface is configured to provide program instructions or data for the at least one processor" — Shen discloses that communication between the CPU 111 and the GPU 113 and between the CPU 112 and the GPU 114 may be performed by using a peripheral component interconnect express (PCIe) bus, that in addition to PCIe the standards ISA, PCI, AGP, AGI, and AGU are also available GPU interface standards, and that the CPU delivers a calculation command to the GPU and the GPU completes the calculation command delivered by the CPU (Shen, ¶ [0029], FIG. 1). Shen further discloses that the CPUs of the servers split the level-1 training data subsets and the level-1 data flow sub-diagrams into level-2 data subsets and level-2 data flow sub-diagrams, that the level-2 subsets of the data and the level-2 data flow sub-diagrams are sent to corresponding GPUs, and that each GPU is instructed to perform calculation on a received level-2 data subset based on a received level-2 data flow sub-diagram (Shen, ¶ [0039], FIG. 3, step S12). Shen further discloses that the master server has a processor and an interface (Shen, ¶ [0038]). The recited interface reads on the PCIe bus of Shen by which the CPU delivers to the GPU both the calculation command, which reads on the recited program instructions, and the level-2 data subset and level-2 data flow sub-diagram, which read on the recited data. The claim recites program instructions and data in the alternative, and Shen discloses both.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The text of those sections of Title 35, U.S. Code not included in this action can be found in a prior Office action.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 1, 2, 9, 12, and 13 are rejected under 35 U.S.C. 103 as being unpatentable over Shen et al. (Shen), WO 2020/199560 A1, "AI Training Network and Method," published October 8, 2020, an English translation of which is available as US 2022/0012590 A1, published January 13, 2022, in view of Chen et al. (Chen), Non-Patent Literature, "OSA: An Optical Switching Architecture for Data Center Networks With Unprecedented Flexibility," IEEE/ACM Transactions on Networking, vol. 22, no. 2, pp. 498–511, published April 2014, cited in the Information Disclosure Statement filed February 26, 2025.
As to independent Claim 1, Shen teaches a model training system, comprising:
"a first group comprising a micro-electro-mechanical system (MEMS) and S×C processors, wherein S is a quantity of nodes in the first group, C is a quantity of processors in a node, and both S and C are positive integers" — Shen discloses an AI training network formed by a server array that includes a server 11, a server 12, a server 13, and a server 14, and that further includes an optical cross-connect 15, an optical cross-connect 16, and an optical cross-connect 17 (Shen, ¶ [0027], FIG. 1). Shen discloses that the optical cross-connect is a micro-electro-mechanical system (MEMS) (Shen, ¶ [0012]), and identifies the optical cross-connect located between the GPU 123 and the GPU 113 as a MEMS 15 comprising a micro-electro-mechanical controller 150 and two reflection lens arrays (Shen, ¶ [0042], FIG. 4). Shen further discloses that the server 11 includes a CPU 111, a CPU 112, a GPU 113, and a GPU 114 (Shen, ¶ [0029], FIG. 1), and that the server 12 includes a GPU 123 and a GPU 124 (Shen, ¶ [0040], FIG. 1). Shen further discloses that there are four servers in total, that each server processes a one-quarter training data set and the corresponding one-quarter data flow diagram (Shen, ¶ [0038]), that the CPUs of the slave servers split the level-1 subsets again based on a quantity of GPUs (Shen, ¶ [0039], FIG. 3, step S12), and that the GPU 123 and the GPU 124 each undertake a one-eighth training data set and a one-eighth data flow diagram (Shen, ¶ [0040]). The recited first group reads on the server array of Shen together with the MEMS 15, the recited S reads on the four servers of Shen, the recited C reads on the two graphics processing units per server of Shen, both being positive integers, and the recited S×C processors read on the eight graphics processing units of Shen.
"wherein the S×C processors are configured to jointly train a model" — Shen discloses that the data flow diagram, which is the structure of the neural network, comprises a plurality of operators, and that the operators need to be distributed to a plurality of GPUs so that the plurality of GPUs may jointly complete data flow diagram calculation (Shen, ¶ [0030]). Shen further discloses that the master server 11 runs an AI training program, loads a training data set and a data flow diagram, and splits both into level-1 subsets sent to the slave servers, whose CPUs split them again into level-2 data subsets and level-2 data flow sub-diagrams sent to the corresponding GPUs, each GPU being instructed to perform calculation on its received level-2 data subset based on its received level-2 data flow sub-diagram (Shen, ¶¶ [0037]–[0039], FIG. 3, steps S11 and S12).
"wherein at least two processors of the S×C processors are configured to transmit target data through the optical transmission channel" — Shen discloses that after the GPU 113 completes the calculation, the calculation result may be immediately sent to the GPU 123 by using the optical path 155-156-157 (Shen, ¶ [0051], FIG. 3, step S15), and that once the GPU 113 completes the calculation the data may be immediately sent to the GPU 123 by using the MEMS 15 (Shen, ¶ [0052]). Shen designates that calculation result of the GPU 113 as the training data set undertaken by the GPU 123 (Shen, ¶ [0043]). The recited target data reads on the calculation result of the GPU 113 of Shen, and the recited at least two processors read on the GPU 113 and the GPU 123 of Shen.
"wherein a processor of the S×C processors that is configured to receive the target data and to adjust a parameter for model training in the processor based on the target data" — Shen discloses that the second graphics processing unit performs AI training calculation on the calculation result by using a second data flow diagram (Shen, ¶ [0006]), and that after receiving the calculation result of the GPU 113 the GPU 123 further performs calculation based on the data flow sub-diagram of the GPU 123 (Shen, ¶ [0051], FIG. 3, step S15). Shen further discloses that in a complete AI training process the steps of forward propagation, loss calculation, and backward propagation are iteratively performed until the calculation result converges to sufficient precision, the backward propagation propagating the loss backward level by level according to a chain derivation rule to obtain gradients of all parameters (Shen, ¶ [0028]), that the respective learning results of the accelerators are combined and distributed to each accelerator again before a next iteration is performed (Shen, ¶ [0004]), and that a task of the AI training is to find an assignment of each parameter in the neural network architecture (Shen, ¶ [0003]).
Shen teaches establishing an optical transmission channel between nodes, in that the MEMS 15 establishes an optical transmission channel between the GPU 113 and the GPU 123 by adjusting the reflector 153 from a deflection angle of 45° to a deflection angle of 30°, thereby changing the reflection path from 155-156-158 to 155-156-157 (Shen, ¶¶ [0042], [0049], [0050], FIG. 4).
However, Shen does not teach "wherein the MEMS is configured to establish an optical transmission channel between any two nodes in the S nodes."
In the same field of endeavor, Chen teaches "wherein the MEMS is configured to establish an optical transmission channel between any two nodes in the S nodes." Chen discloses that optical switching matrix modules are a bipartite switching matrix where any input port can connect to any of the output ports, and that a microelectromechanical system (MEMS) is the most popular optical switching matrix technology and achieves a reconfigurable one-to-one circuit by mechanically adjusting micro mirrors (Chen, page 500, section II-B, subsection 3) (“Optical Switching Matrix (OSM)”, fifth paragraph in column 1). Chen further discloses that, by connecting each top-of-rack switch to k ports at the MEMS, each top-of-rack switch can communicate with k other top-of-rack switches simultaneously, that the MEMS configuration determines which set of top-of-rack switches are connected, and that hop-by-hop stitching of the MEMS optical circuits achieves network-wide connectivity to top-of-rack switches that are not directly connected (Chen, page 500, section III-A (“Building Blocks”), second paragraph in column 2). Chen further discloses that each top-of-rack switch can communicate simultaneously with any four other top-of-rack switches and that the MEMS configuration allows construction of all possible 4-regular graphs among the top-of-rack switches (Chen, page 501, section III-B (“Putting It All Together: OSA-2560”), fifth paragraph in column 1). The recited nodes read on the top-of-rack switches of Chen.
Shen and Chen are analogous to the claimed invention as both are from the same field of endeavor of establishing reconfigurable optical transmission channels between computing nodes by means of a micro-electro-mechanical optical switch. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the AI training network of Shen with the any-port-to-any-port MEMS configuration of Chen. The motivation to combine Shen and Chen is that Chen teaches that the bipartite MEMS matrix permits construction of every possible regular graph among the nodes and adapts the topology and link capacities to dynamic traffic patterns, thereby delivering high bisection bandwidth (Chen, page 498, Abstract; page 500, column 2, section III-A, paragraph 1), which addresses the problem Shen identifies of data being transmitted between GPU chips increasingly frequently as the accelerator cluster scale grows, such that the speed of data transmission between the GPU chips has an increasingly obvious impact on the duration of the entire training process (Shen, ¶ [0005]).
As to dependent Claim 2, the limitations of claim 1 from which claim 2 depends are rejected under the same rationale set forth above with respect to claim 1.
Regarding the additional limitations of claim 2, Shen teaches:
"wherein the first group comprises a first node and a second node, the first node comprises a first processor, and the second node comprises a second processor" — Shen discloses a server array including a server 11 and a server 12 (Shen, ¶ [0027], FIG. 1), discloses that the server 11 includes a CPU 111, a CPU 112, a GPU 113, and a GPU 114 (Shen, ¶ [0029], FIG. 1), and discloses that the server 12 includes a GPU 123 and a GPU 124 (Shen, ¶ [0040], FIG. 1). The recited first node reads on the server 11 of Shen, the recited first processor reads on the GPU 113 of Shen, the recited second node reads on the server 12 of Shen, and the recited second processor reads on the GPU 123 of Shen.
"wherein the first processor is configured to perform model training in the first processor to obtain intermediate data of the first processor, and obtain first target data based on the intermediate data of the first processor, wherein the first target data is all or a part of the intermediate data of the first processor" — Shen discloses that the CPUs of the slave servers split the level-1 training data subsets and the level-1 data flow sub-diagrams again based on a quantity of GPUs into level-2 data subsets and level-2 data flow sub-diagrams, that the level-2 subsets and level-2 data flow sub-diagrams are sent to the corresponding GPUs, and that each GPU is instructed to perform calculation on a received level-2 data subset based on a received level-2 data flow sub-diagram (Shen, ¶ [0039], FIG. 3, step S12). Shen further discloses that the training data set and the data flow diagram undertaken by the GPU 113 are referred to as a first training data set and a first data flow diagram, and that the calculation result of the GPU 113 is the training data set undertaken by the GPU 123 (Shen, ¶ [0043]). Shen further discloses that an output result calculated by one operator is used as input data of another operator, and that an output result of a previous GPU is used as input data of a next GPU (Shen, ¶ [0030]). The recited intermediate data of the first processor reads on the result of the calculation performed by the GPU 113 on its level-2 data subset, and the recited first target data reads on the calculation result of the GPU 113 that is forwarded, which is the entirety of that result and therefore satisfies the "all" alternative of the recitation.
"wherein the first processor is further configured to send the first target data to the second processor through an optical transmission channel established by a first MEMS and an intra-node channel" — Shen discloses that after the GPU 113 completes the calculation, the calculation result may be immediately sent to the GPU 123 by using the optical path 155-156-157 (Shen, ¶ [0051], FIG. 3, step S15), and that once the GPU 113 completes the calculation the data may be immediately sent to the GPU 123 by using the MEMS 15 (Shen, ¶ [0052]). Shen further discloses that an electrical signal sent by the GPU 113 is converted into an optical signal and reaches the receiving graphics processing unit after passing through a fiber channel 151, a reflector 152, a reflector 153, and a fiber channel 154 (Shen, ¶ [0042], FIG. 4). The recited optical transmission channel established by a first MEMS reads on the reflection path 155-156-157 established through the MEMS 15 of Shen, and the recited intra-node channel reads on the fiber channel 151 and the fiber channel 154 of Shen.
"wherein the second processor is configured to adjust a parameter for model training in the second processor based on the first target data" — Shen discloses that after receiving the calculation result of the GPU 113, the GPU 123 further performs calculation based on the data flow sub-diagram of the GPU 123 (Shen, ¶ [0051]). Shen further discloses that the AI training process iteratively performs forward propagation, loss calculation, and backward propagation, the backward propagation propagating the loss backward level by level to obtain the gradients of all parameters, until convergence is formed (Shen, ¶ [0028]), and that the respective learning results of the accelerators are combined and distributed to each accelerator again before a next iteration is performed (Shen, ¶ [0004]).
"wherein the first MEMS is located between the first node and the second node" — Shen discloses that, where the input of the GPU 123 depends on the output of the GPU 113, the lens that needs to be adjusted is an optical cross-connect located between the GPU 123 and the GPU 113, namely, a MEMS 15 (Shen, ¶ [0042], FIG. 4). The recited first MEMS reads on the MEMS 15 of Shen, which is disposed between the server 11 containing the GPU 113 and the server 12 containing the GPU 123.
"wherein the intra-node channel comprises: a channel that is in the first node and that is between the first processor and the first MEMS, and/or a channel that is in the second node and that is between the second processor and the first MEMS" — Shen discloses that the MEMS 15 includes a micro-electro-mechanical controller 150 and two reflection lens arrays, and that the optical signal converted from the electrical signal sent by the GPU 113 passes through a fiber channel 151 before reaching the reflector 152 and passes through a fiber channel 154 after leaving the reflector 153 (Shen, ¶ [0042], FIG. 4). Shen further discloses that where the two GPUs that communicate with each other belong to the same server, the communication may be performed by using a bus inside the server (Shen, ¶ [0031]). The recited channel that is in the first node and that is between the first processor and the first MEMS reads on the fiber channel 151 of Shen, which extends within the server 11 from the GPU 113 to the reflector 152 of the MEMS 15, and the recited channel that is in the second node and that is between the second processor and the first MEMS reads on the fiber channel 154 of Shen, which extends within the server 12 from the reflector 153 of the MEMS 15 to the receiving graphics processing unit.
As to dependent Claim 9, the limitations of claim 1 from which claim 9 depends are rejected under the same rationale set forth above with respect to claim 1.
Regarding the additional limitations of claim 9, Shen teaches:
"wherein the target data comprises one or more of a gradient, a feature, or a model parameter for model iteration" — Shen discloses that in a complete AI training process the following steps are iteratively performed until a calculation result converges to sufficient precision: forward propagation, in which input data is input into a neural network and operators are run in sequence based on an operator dependency relationship; calculation of a loss, being the difference between the obtained result and the correct answer; and backward propagation, in which, according to a chain derivation rule, the loss is propagated backward level by level to obtain gradients of all parameters (Shen, ¶ [0028]). Shen further discloses that a plurality of accelerators separately perform calculation by using a training algorithm, that respective learning results are combined and distributed to each accelerator again, and that a next iteration is then performed, so that after a plurality of rounds of iterative calculation the machine learns more key details (Shen, ¶ [0004]). Shen further discloses that a task of the AI training is to find an appropriate neural network architecture and an assignment of each parameter in the neural network architecture (Shen, ¶ [0003]). The recited target data reads on the calculation result that the GPU 113 sends to the GPU 123 over the optical path 155-156-157 (Shen, ¶ [0051], FIG. 3, step S15), which is the output of the training operators executed by the GPU 113 on its level-2 data subset (Shen, ¶¶ [0039], [0043]) and which, under the iterative process of Shen, comprises the gradients obtained by backward propagation that are used to update the parameter assignment on the succeeding iteration. The claim recites the three enumerated items in the alternative and requires only one of them, and Shen expressly discloses the gradient.
As to dependent Claim 12, the limitations of claim 10 from which claim 12 depends are rejected under the same rationale set forth above with respect to claim 10.
Regarding the additional limitations of claim 12, Shen teaches sending the first target data through an optical transmission channel established by the MEMS, in that after the GPU 113 completes the calculation the calculation result may be immediately sent to the GPU 123 by using the optical path 155-156-157 established through the MEMS 15 (Shen, ¶¶ [0051], [0052], FIG. 4).
However, Shen does not teach:
"sending, by the first processor, the first target data to the second processor sequentially through an optical transmission channel established by a wavelength selective switch (WSS) and the optical transmission channel established by the MEMS"
"wherein the second node and the MEMS belong to a same group"
"wherein the WSS is located between the MEMS and the first node"
In the same field of endeavor, Chen teaches "sending, by the first processor, the first target data to the second processor sequentially through an optical transmission channel established by a wavelength selective switch (WSS) and the optical transmission channel established by the MEMS." Chen discloses that a WSS consists of one common port and a plurality of wavelength ports and partitions, runtime-configurably within a few milliseconds, the set of wavelengths coming in through the common port among the wavelength ports (Chen, page 500, section II-B (“Optical Networking Technologies”), subsection 2) (“Wavelength Selective Switch (WSS)”), fourth paragraph in column 1). Chen further discloses that the send fiber from the transceiver of each of the thirty-two optical-facing ports at a top-of-rack switch is connected to an optical multiplexer, that the multiplexer feeds a 1×4 WSS, that the WSS splits the set of thirty-two wavelengths it sees into four groups each transmitted on its own fiber, and that these fibers are connected to the MEMS via circulators (Chen, page 501, section III-B (“Putting It All Together: OSA-2560”), fourth paragraph in column 1). Chen further discloses that a circulator connects the send channel of the transceiver from a top-of-rack switch to the MEMS after the channel has passed through the multiplexer and WSS (Chen, page 501, section III-A (“Building Blocks”), “Efficient Port Usage”, second paragraph in column 1), and that a micro-electro-mechanical system achieves a reconfigurable one-to-one circuit by mechanically adjusting micro mirrors (Chen, page 500, section II-B (“Optical Networking Technologies”), subsection 3 (“Optical Switching Matrix (OSM)”), fifth paragraph in column 1). The transmitted data of Chen therefore traverses the channel established by the WSS and thereafter the channel established by the MEMS, in that sequence.
Chen further teaches "wherein the second node and the MEMS belong to a same group." Chen discloses that the WSS splits the wavelengths into four groups, each group being transmitted on its own fiber to a distinct MEMS port, and that the WSS splits these wavelengths to the appropriate MEMS port that has a circuit to the destination (Chen, page 500, section III-A (“Building Blocks”), fourth paragraph in column 2). Chen further discloses that each top-of-rack switch can communicate simultaneously with any four other top-of-rack switches, and that the MEMS configuration determines which set of top-of-rack switches are connected (Chen, page 501, section III-B (“Putting It All Together: OSA-2560”), fifth paragraph in column 1; page 500, section III-A (“Building Blocks”), first paragraph in column 2). The recited second node and MEMS belonging to a same group read on the destination top-of-rack switch of Chen and the MEMS port associated with the particular wavelength group over which that destination is reached.
Chen further teaches "wherein the WSS is located between the MEMS and the first node." Chen discloses the send path ordering in which the transceiver ports of a top-of-rack switch feed the optical multiplexer, the multiplexer feeds the 1×4 WSS, and the four output fibers of the WSS connect to the MEMS via circulators (Chen, page 501, section III-B (“Putting It All Together: OSA-2560”), fourth paragraph in column 1; page 501, section III-A (“Building Blocks”), “Efficient Port Usage”, second paragraph in column 1). The WSS of Chen is accordingly interposed between the source top-of-rack switch and the MEMS.
Shen and Chen are analogous to the claimed invention as both are from the same field of endeavor of optical interconnection of computing nodes through a micro-electro-mechanical optical switch. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to interpose the wavelength selective switch of Chen between the server of Shen and the micro-electro-mechanical system of Shen, such that the calculation result sent by the GPU 113 traverses the WSS and then the MEMS. The motivation to combine Shen and Chen is that Chen teaches that the WSS configuration permits the capacity of each of the four links to be varied at runtime (Chen, page 501, section III-B (“Putting It All Together: OSA-2560”), fifth paragraph in column 1), which addresses the problem Shen identifies of increasingly frequent data transmission between GPU chips as the accelerator cluster scale grows (Shen, ¶ [0005]).
As to dependent Claim 13, the limitations of claims 10 and 12 from which claim 13 depends are rejected under the same rationale set forth above with respect to claims 10 and 12.
Regarding the additional limitations of claim 13, Shen teaches converting the first target data into optical form for transmission, in that an electrical signal sent by the GPU 113 is converted into an optical signal and reaches the destination graphics processing unit after passing through a fiber channel 151, a reflector 152, a reflector 153, and a fiber channel 154 (Shen, ¶ [0042], FIG. 4).
However, Shen does not teach:
"wherein the WSS comprises a mapping relationship between a wavelength of a carrier and a group, and in one mapping relationship, a wavelength of a carrier is a preset wavelength of a corresponding group"
"modulating, by the first processor, the first target data to a carrier, wherein a wavelength of the carrier is a preset wavelength corresponding to a group to which the second node belongs"
"sending, by the first processor, the carrier carrying the first target data to the WSS, to enable the WSS to send the carrier carrying the first target data to the MEMS"
In the same field of endeavor, Chen teaches "wherein the WSS comprises a mapping relationship between a wavelength of a carrier and a group, and in one mapping relationship, a wavelength of a carrier is a preset wavelength of a corresponding group." Chen discloses that a WSS consists of one common port and a plurality of wavelength ports, that it partitions, runtime-configurably within a few milliseconds, the set of wavelengths coming in through the common port among the wavelength ports, and that if the common port receives eighty wavelengths it can route wavelengths 1–20 on port 1 and 30–40 on port 2 (Chen, page 500, section II-B (“Optical Networking Technologies”), subsection 2) (“Wavelength Selective Switch (WSS)”), fourth paragraph in column 1). Chen further discloses that the 1×4 WSS splits the set of thirty-two wavelengths it sees into four groups, each group being transmitted on its own fiber to the MEMS (Chen, page 501, section III-B (“Putting It All Together: OSA-2560”), fourth paragraph in column 1). Chen further discloses that each port facing the optical interconnect has a transceiver associated with a fixed and unique wavelength for sending and receiving data, that this wavelength-to-port association is a design-time decision, and that the same set of wavelengths is recycled across top-of-rack switches (Chen, page 501, section III-B (“Putting It All Together: OSA-2560”), third paragraph in column 1; page 501, section III-A (“Building Blocks”), first paragraph in column 1). The recited mapping relationship between a wavelength of a carrier and a group reads on the wavelength-to-output-port partition maintained by the WSS of Chen, and the recited preset wavelength of a corresponding group reads on the fixed and unique design-time wavelength of Chen that falls within the group directed to that output port.
Chen further teaches "modulating, by the first processor, the first target data to a carrier, wherein a wavelength of the carrier is a preset wavelength corresponding to a group to which the second node belongs." Chen discloses that dense wavelength division multiplexing transceivers are used, which support higher bit rates and more wavelength channels in a single piece of fiber (Chen, page 500, section II-B (“Optical Networking Technologies”), subsection 5), (“Optical Transceivers”), last paragraph in column 1), and that each port facing the optical interconnect has a transceiver associated with a fixed and unique wavelength for sending data, each port thus sending traffic at one fixed wavelength (Chen, page 501, section III-B (“Putting It All Together: OSA-2560”), third paragraph in column 1; page 501, section III-A (“Building Blocks”), first paragraph in column 1). Chen further discloses that, to serve a request, the source node uses the ports whose associated wavelengths are those the WSS splits to the appropriate MEMS port having a circuit to the destination (Chen, page 500, column 2, section III-A (“Building Blocks”), fourth paragraph in column 2). The recited preset wavelength corresponding to a group to which the second node belongs reads on the fixed wavelength of Chen falling within the wavelength group that the WSS directs toward the destination top-of-rack switch.
Chen further teaches "sending, by the first processor, the carrier carrying the first target data to the WSS, to enable the WSS to send the carrier carrying the first target data to the MEMS." Chen discloses that the send fiber from the transceiver of each optical-facing port at a top-of-rack switch is connected to an optical multiplexer, that wavelength division multiplexing enables these wavelengths to be multiplexed into one optical fiber that feeds the WSS, and that the WSS splits these wavelengths to the appropriate MEMS port that has a circuit to the destination, the four resulting fibers being connected to the MEMS via circulators (Chen, page 500, column 2, section III-A (“Building Blocks”), fourth paragraph in column 2; page 501, section III-A (“Building Blocks”), “Efficient Port Usage”, second paragraph in column 1). Chen further discloses that the MEMS and WSS configurations are decided by a central manager which estimates the traffic demand, calculates the appropriate configurations, and pushes them to the MEMS, WSS units, and top-of-rack switches (Chen, page 501, section III-B (“Putting It All Together: OSA-2560”), fifth paragraph in column 1).
Shen and Chen are analogous to the claimed invention as both are from the same field of endeavor of optical interconnection of computing nodes through a micro-electro-mechanical optical switch. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to convert the electrical signal of the GPU 113 of Shen onto a carrier at the destination-associated wavelength of Chen and to forward that carrier to the wavelength selective switch of Chen for delivery to the micro-electro-mechanical system of Shen under the wavelength-to-port mapping of Chen. The motivation to combine Shen and Chen is that Chen teaches that this wavelength assignment permits all wavelengths at a source node to be multiplexed onto a single fiber and, after demultiplexing, delivered to individual ports at the destination nodes (Chen, page 501, section III-A (“Building Blocks”), first paragraph in column 1), which addresses the problem Shen identifies of increasingly frequent data transmission between GPU chips as the accelerator cluster scale grows (Shen, ¶ [0005]).
Claims 3-8 are rejected under 35 U.S.C. 103 as being unpatentable over Shen et al. (Shen), WO 2020/199560 A1, "AI Training Network and Method," published October 8, 2020, an English translation of which is available as US 2022/0012590 A1, published January 13, 2022, in view of Chen et al. (Chen), Non-Patent Literature, "OSA: An Optical Switching Architecture for Data Center Networks With Unprecedented Flexibility," IEEE/ACM Transactions on Networking, vol. 22, no. 2, pp. 498–511, published April 2014, cited in the Information Disclosure Statement filed February 26, 2025, and further in view of Lu et al. (Lu), Non-Patent Literature, "X-NEST: A Scalable, Flexible, and High-Performance Network Architecture for Distributed Machine Learning," Journal of Lightwave Technology, vol. 39, no. 13, pp. 4247–4254, published July 1, 2021, cited in the Information Disclosure Statement filed February 26, 2025.
As to dependent Claim 3, the limitations of claim 1 from which claim 3 depends are rejected under the same rationale set forth above with respect to claim 1.
Regarding the additional limitations of claim 3, Shen teaches an array containing more than one optical cross-connect, in that the server array further includes an optical cross-connect 15, an optical cross-connect 16, and an optical cross-connect 17, and data sent by the GPU 113 reaches the GPU 144 after successively passing through the OXC 15, the OXC 16, and the OXC 17 (Shen, ¶¶ [0027], [0031], FIG. 1).
However, Shen does not teach:
"a wavelength selective switch (WSS)"
"wherein the WSS is connected to each of the W groups"
In the same field of endeavor, Chen teaches "a wavelength selective switch (WSS)." Chen discloses that a WSS consists of one common port and a plurality of wavelength ports, and that it partitions, runtime-configurably within a few milliseconds, the set of wavelengths arriving at the common port among the wavelength ports (Chen, page 500, section II-B (“Optical Networking Technologies”), subsection 2) (“Wavelength Selective Switch (WSS)”), fourth paragraph in column 1). Chen further discloses that the send fiber from the transceiver of each of the thirty-two optical-facing ports at a top-of-rack switch is connected to an optical multiplexer, and that the multiplexer feeds a 1×4 WSS (Chen, page 501, section III-B (“Putting It All Together: OSA-2560”), fourth paragraph in column 1).
Chen further teaches "wherein the WSS is connected to each of the W groups." Chen discloses that the WSS splits the set of thirty-two wavelengths it sees into four groups, each group being transmitted on its own fiber, and that these fibers are connected to the MEMS via circulators to enable bidirectional communications (Chen, page 500, column 2, section III-A (“Building Blocks”), fourth paragraph in column 2). Chen further discloses that a circulator connects the send channel of the transceiver from a top-of-rack switch to the MEMS after the channel has passed through the multiplexer and WSS (Chen, page 501, section III-A (“Building Blocks”), “Efficient Port Usage”, second paragraph in column 1), and that, as a result, each top-of-rack switch can communicate simultaneously with any four other top-of-rack switches (Chen, page 501, section III-B (“Putting It All Together: OSA-2560”), fifth paragraph in column 1). The recited W reads on the four wavelength ports of the 1×4 WSS of Chen, each of which is connected through a distinct MEMS port to a distinct destination, and W is therefore an integer greater than or equal to 2.
Shen and Chen are analogous to the claimed invention as both are from the same field of endeavor of optical interconnection of computing nodes through a micro-electro-mechanical optical switch. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to interpose the wavelength selective switch of Chen between each server of Shen and the micro-electro-mechanical system of Shen. The motivation to combine Shen and Chen is that Chen teaches that the WSS configuration permits the capacity of each of the four links to be varied at runtime (Chen, page 501, section III-B (“Putting It All Together: OSA-2560”), fifth paragraph in column 1), which addresses the problem Shen identifies of increasingly frequent data transmission between GPU chips as the accelerator cluster scale grows (Shen, ¶ [0005]).
The combination of Shen and Chen, however, does not teach "and (W-1) extended groups, wherein W is an integer greater than or equal to 2, and the first group and the (W-1) extended groups form W groups."
In the same field of endeavor, Lu teaches "and (W-1) extended groups, wherein W is an integer greater than or equal to 2, and the first group and the (W-1) extended groups form W groups." Lu discloses that a scalable topology is achieved by a hierarchical network with three levels, namely group, pod, and system, that there are k pods in the network, that each pod has k groups, and that each group has a top-of-rack switch connected to k computing nodes or servers, giving a total of k² top-of-rack switches and k³ computing nodes (Lu, page 4248, section III-A (“Topology”), first paragraph in column 2, continuing to page 4249, column 1). Lu further discloses that the pods are connected through MEMS switches with 100 Gbps optical links, that top-of-rack switches having the same x coordinate are connected to 2m different MEMS switches, and that those having the same y coordinate are connected to the same MEMS switch (Lu, page 4248, section III-A (“Topology”), first paragraph in column 2, continuing to page 4249, column 1; Lu, page 4249, FIG. 1 and caption). The recited W groups read on the k pods of Lu, each pod comprising a plurality of nodes each containing a plurality of processors and being served by MEMS switches, the recited first group reads on one such pod, and the recited (W-1) extended groups read on the remaining k−1 pods.
Shen, Chen, and Lu are analogous to the claimed invention as all three are from the same field of endeavor of optical interconnection of computing nodes through a micro-electro-mechanical optical switch. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to arrange the server array of Shen, as modified by the wavelength selective switch of Chen, into the plurality of MEMS-interconnected groups taught by Lu. The motivation to combine Shen, Chen, and Lu is that Lu teaches that the hierarchical group-and-pod arrangement achieves a scalable topology reaching k³ computing nodes (Lu, page 4248, column 2, section III-A, paragraph 1), which addresses the problem Shen identifies of an accelerator cluster of increasing scale (Shen, ¶ [0005]).
As to dependent Claim 4, the limitations of claims 1 and 3 from which claim 4 depends are rejected under the same rationale set forth above with respect to claims 1 and 3. The recitation "the W groups" in claim 4 finds its antecedent in claim 3, and Lu is relied upon for the W groups as set forth above with respect to claim 3.
Regarding the additional limitations of claim 4, Shen teaches ports of an optical switch, in that the optical cross-connects are connected to the data exchange network 18 by using an Ethernet or a fiber channel so as to receive a command from the server CPU and adjust a connection between an input and an output of an optical switch based on that command (Shen, ¶ [0031], FIG. 1).
However, Shen does not teach:
"wherein the WSS comprises W first WSS ports and W second WSS ports"
"wherein the W first WSS ports are respectively connected to W node ports, wherein the W node ports respectively belong to the W groups, and the W node ports are located at corresponding positions in respective groups"
"wherein the W node ports correspond to respective MEMS ports in the respective groups, and the MEMS ports corresponding to the W node ports are respectively connected to the W second WSS ports"
In the same field of endeavor, Chen teaches "wherein the WSS comprises W first WSS ports and W second WSS ports." Chen discloses that a WSS consists of one common port and a plurality of wavelength ports, and that it partitions the set of wavelengths arriving at the common port among the wavelength ports (Chen, page 500, section II-B (“Optical Networking Technologies”), subsection 2) (“Wavelength Selective Switch (WSS)”), fourth paragraph in column 1). Chen further discloses that the send fiber from the transceiver of each of the thirty-two optical-facing ports at a top-of-rack switch is connected to an optical multiplexer, that the multiplexer feeds a 1×4 WSS, and that the WSS splits the set of thirty-two wavelengths it sees into four groups, each group being transmitted on its own fiber (Chen, page 501, section III-B (“Putting It All Together: OSA-2560”), fourth paragraph in column 1). The recited W first WSS ports read on the four wavelength-bearing inputs presented to the WSS of Chen, one for each of the four groups into which the wavelength set is partitioned, and the recited W second WSS ports read on the four output fibers of the 1×4 WSS of Chen, W being four.
Chen further teaches "wherein the W first WSS ports are respectively connected to W node ports, wherein the W node ports respectively belong to the W groups, and the W node ports are located at corresponding positions in respective groups." Chen discloses that each port facing the optical interconnect has a transceiver associated with a fixed and unique wavelength for sending and receiving data, and that the send fiber from each such port is connected to the optical multiplexer that feeds the WSS (Chen, page 501, column 2, section III-B, paragraphs 3 and 4). Because the WSS partitions the thirty-two wavelengths into four groups, each node port of Chen belongs to the one of the four groups containing that port's fixed unique wavelength (Chen, page 500, section III-A (“Building Blocks”), fourth paragraph in column 2; page 500, section II-B (“Optical Networking Technologies”), subsection 2) (“Wavelength Selective Switch (WSS)”), fourth paragraph in column 1). Chen further discloses that the wavelength-to-port association is a design-time decision and that the same set of wavelengths is recycled across top-of-rack switches, so that all wavelengths at a source top-of-rack switch are multiplexed and, after demultiplexing, delivered to individual ports at the destination top-of-rack switches (Chen, page 501, section III-A (“Building Blocks”), first paragraph in column 1). The recited node ports located at corresponding positions in respective groups read on the ports of Chen bearing the same recycled wavelength at each top-of-rack switch, which occupy the same position within their respective groups.
Chen further teaches "wherein the W node ports correspond to respective MEMS ports in the respective groups, and the MEMS ports corresponding to the W node ports are respectively connected to the W second WSS ports." Chen discloses that the four fibers carrying the four wavelength groups are connected to the MEMS via circulators to enable bidirectional communications (Chen, page 500, section III-A (“Building Blocks”), fourth paragraph in column 2), and that a circulator connects the send channel of the transceiver from a top-of-rack switch to the MEMS after the channel has passed through the multiplexer and WSS (Chen, page 501, section III-A (“Building Blocks”), “Efficient Port Usage”, second paragraph in column 1). Chen further discloses that each top-of-rack switch is thereby able to communicate simultaneously with any four other top-of-rack switches, the MEMS configuration determining which set of top-of-rack switches are connected (Chen, page 501, section III-B (“Putting It All Together: OSA-2560”), fifth paragraph in column 1; page 500, section III-A (“Building Blocks”), first paragraph in column 2). The recited MEMS ports corresponding to the W node ports read on the four MEMS ports of Chen to which the four group fibers from the WSS are respectively connected.
Shen and Chen are analogous to the claimed invention as all two are from the same field of endeavor of optical interconnection of computing nodes through a micro-electro-mechanical optical switch. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to connect the ports of the servers of Shen to the ports of the micro-electro-mechanical system of Shen through the wavelength selective switch port arrangement of Chen. The motivation to combine Shen and Chen is that Chen teaches that this port arrangement makes full use of the MEMS ports by rendering each circuit over the MEMS bidirectional (Chen, page 501, section III-A (“Building Blocks”), “Efficient Port Usage”, second paragraph in column 1), which addresses the problem Shen identifies of increasingly frequent data transmission between GPU chips as the accelerator cluster scale grows (Shen, ¶ [0005]).
As to dependent Claim 5, the limitations of claims 1 and 3 from which claim 5 depends are rejected under the same rationale set forth above with respect to claims 1 and 3.
Regarding the additional limitations of claim 5, Shen teaches:
"wherein the first group comprises a first node, and the first node comprises a first processor" — Shen discloses a server array including a server 11, a server 12, a server 13, and a server 14 (Shen, ¶ [0027], FIG. 1), and discloses that the server 11 includes a CPU 111, a CPU 112, a GPU 113, and a GPU 114 (Shen, ¶ [0029], FIG. 1). The recited first node reads on the server 11 of Shen and the recited first processor reads on the GPU 113 of Shen.
"wherein the first processor is configured to perform model training of the first processor to obtain intermediate data of the first processor, and obtain first target data based on the intermediate data of the first processor, wherein the first target data is all or a part of the intermediate data of the first processor" — Shen discloses that the level-2 data subsets and the level-2 data flow sub-diagrams are sent to the corresponding GPUs and that each GPU is instructed to perform calculation on a received level-2 data subset based on a received level-2 data flow sub-diagram (Shen, ¶ [0039], FIG. 3, step S12). Shen further discloses that the training data set and the data flow diagram undertaken by the GPU 113 are referred to as a first training data set and a first data flow diagram, and that the calculation result of the GPU 113 is the training data set undertaken by the GPU 123 (Shen, ¶ [0043]), and that an output result of a previous GPU is used as input data of a next GPU (Shen, ¶ [0030]). The recited intermediate data of the first processor reads on the result of the calculation performed by the GPU 113 on its level-2 data subset, and the recited first target data reads on the calculation result of the GPU 113 that is forwarded, which is the entirety of that result and therefore satisfies the "all" alternative of the recitation.
"wherein the second processor is configured to adjust a parameter for model training in the second processor based on the first target data" — Shen discloses that after receiving the calculation result of the GPU 113, the GPU 123 further performs calculation based on the data flow sub-diagram of the GPU 123 (Shen, ¶ [0051], FIG. 3, step S15). Shen further discloses that the AI training process iteratively performs forward propagation, loss calculation, and backward propagation, the backward propagation propagating the loss backward level by level to obtain the gradients of all parameters, until convergence is formed (Shen, ¶ [0028]), and that the respective learning results of the accelerators are combined and distributed to each accelerator again before a next iteration is performed (Shen, ¶ [0004]).
"wherein the second processor is located in a second node, and the second node is another node other than the first node in the first group, or is a node in any one of the (W-1) extended groups" — Shen discloses that the server 12 includes a GPU 123 and a GPU 124 (Shen, ¶ [0040], FIG. 1), and that the input of the GPU 123 depends on the output of the GPU 113 (Shen, ¶ [0042], FIG. 4). The recited second node reads on the server 12 of Shen, which is another node other than the server 11 in the array. Shen further discloses that data sent by the GPU 113 can reach the GPU 144 of the server 14 after successively passing through the OXC 15, the OXC 16, and the OXC 17 (Shen, ¶ [0031], FIG. 1). The claim recites the two locations of the second node in the alternative, and Shen satisfies the first alternative.
Shen teaches sending the first target data through an optical transmission channel, in that after the GPU 113 completes the calculation, the calculation result may be immediately sent to the GPU 123 by using the optical path 155-156-157 established through the MEMS 15 (Shen, ¶¶ [0051], [0052], FIG. 4).
However, Shen does not teach:
"wherein the first processor is further configured to send the first target data to a second processor sequentially through optical transmission channels respectively established by the WSS and a second MEMS"
"wherein the WSS and the second MEMS are sequentially located between the first node and the second node, and the second MEMS and the second node belong to a same group"
In the same field of endeavor, Chen teaches "wherein the first processor is further configured to send the first target data to a second processor sequentially through optical transmission channels respectively established by the WSS and a second MEMS." Chen discloses that the send fiber from the transceiver of each optical-facing port at a top-of-rack switch is connected to an optical multiplexer, that the multiplexer feeds a 1×4 WSS, that the WSS splits the set of thirty-two wavelengths into four groups each transmitted on its own fiber, and that these fibers are connected to the MEMS via circulators (Chen, page 501, section III-B (“Putting It All Together: OSA-2560”), fourth paragraph in column 1). Chen further discloses that a circulator connects the send channel of the transceiver from a top-of-rack switch to the MEMS after the channel has passed through the multiplexer and WSS (Chen, page 501, section III-A (“Building Blocks”), “Efficient Port Usage”, second paragraph in column 1), and that the MEMS achieves a reconfigurable one-to-one circuit by mechanically adjusting micro mirrors, the MEMS configuration determining which set of top-of-rack switches are connected (Chen, page 500, section II-B (“Optical Networking Technologies”), subsection 3 (“Optical Switching Matrix (OSM)”), fifth paragraph in column 1; page 500, section III-A (“Building Blocks”), first paragraph in column 2). The transmitted data of Chen therefore traverses the channel established by the WSS and then the channel established by the MEMS, in that sequence.
Chen further teaches "wherein the WSS and the second MEMS are sequentially located between the first node and the second node, and the second MEMS and the second node belong to a same group." Chen discloses the ordering of the send path from the top-of-rack switch through the multiplexer, the WSS, and the circulator to the MEMS, and thence to the destination top-of-rack switch, whose receive path passes through a power coupler and a demultiplexer that assigns each incoming wavelength to its associated port (Chen, page 500, section III-A (“Building Blocks”), fourth paragraph in column 2; page 501, section III-A (“Building Blocks”), “Efficient Port Usage”, second paragraph in column 1). Chen further discloses that each of the four group fibers from the WSS terminates at a distinct MEMS port and that the WSS splits the selected wavelengths to the MEMS port that has a circuit to the destination, so that each top-of-rack switch can communicate simultaneously with any four other top-of-rack switches (Chen, page 501, column 2, section III-B, paragraphs 4 and 5). The recited second MEMS and second node belonging to a same group read on the MEMS port of Chen that is associated with the particular wavelength group directed to the destination top-of-rack switch.
Shen and Chen are analogous to the claimed invention as both are from the same field of endeavor of optical interconnection of computing nodes through a micro-electro-mechanical optical switch. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to route the calculation result sent by the GPU of Shen first through the wavelength selective switch of Chen and then through the micro-electro-mechanical system of Shen. The motivation to combine Shen and Chen is that Chen teaches that this ordering permits the capacity of each of the four links to be varied at runtime through WSS configuration (Chen, page 501, section III-B (“Putting It All Together: OSA-2560”), fifth paragraph in column 1), which addresses the problem Shen identifies of increasingly frequent data transmission between GPU chips as the accelerator cluster scale grows (Shen, ¶ [0005]).
As to dependent Claim 6, the limitations of claims 1, 3, and 5 from which claim 6 depends are rejected under the same rationale set forth above with respect to claims 1, 3, and 5.
Regarding the additional limitations of claim 6, Shen teaches converting the data of the first processor into optical form for transmission over the optical transmission channel, in that an electrical signal sent by the GPU 113 is converted into an optical signal and reaches the destination graphics processing unit after passing through a fiber channel 151, a reflector 152, a reflector 153, and a fiber channel 154 (Shen, ¶ [0042], FIG. 4).
However, Shen does not teach:
"wherein the first processor is configured to modulate the first target data to a carrier, wherein a wavelength of the carrier is a preset wavelength corresponding to the group to which the second node belongs"
"wherein the WSS is configured to send the carrier carrying the first target data to the second MEMS based on a mapping relationship between the wavelength of the carrier and the group to which the second node belongs"
In the same field of endeavor, Chen teaches "wherein the first processor is configured to modulate the first target data to a carrier, wherein a wavelength of the carrier is a preset wavelength corresponding to the group to which the second node belongs." Chen discloses that each port facing the optical interconnect has a transceiver associated with a fixed and unique wavelength for sending and receiving data, and that dense wavelength division multiplexing transceivers are used (Chen, page 501, section III-B (“Putting It All Together: OSA-2560”), third paragraph in column 1; page 500, section II-B (“Optical Networking Technologies”), subsection 5), (“Optical Transceivers”), last paragraph in column 1). Chen further discloses that this wavelength-to-port association is a design-time decision, that the same set of wavelengths is recycled across top-of-rack switches, and that each port thus sends and receives traffic at one fixed wavelength (Chen, page 501, section III-A (“Building Blocks”), first paragraph in column 1). Chen further discloses that the WSS splits the set of thirty-two wavelengths into four groups, each group being transmitted on its own fiber to the MEMS port having a circuit to the destination (Chen, page 500, section III-A (“Building Blocks”), fourth paragraph in column 2). The recited preset wavelength corresponding to the group to which the second node belongs reads on the fixed and unique design-time wavelength of Chen that falls within the wavelength group the WSS directs toward the destination top-of-rack switch, and the recited modulating of the first target data to a carrier reads on the transmission of that data by the transceiver of Chen at its assigned wavelength.
Chen further teaches "wherein the WSS is configured to send the carrier carrying the first target data to the second MEMS based on a mapping relationship between the wavelength of the carrier and the group to which the second node belongs." Chen discloses that a WSS consists of one common port and a plurality of wavelength ports, that it partitions, runtime-configurably within a few milliseconds, the set of wavelengths coming in through the common port among the wavelength ports, and that if the common port receives eighty wavelengths it can route wavelengths 1–20 on port 1 and 30–40 on port 2 (Chen, page 500, section II-B (“Optical Networking Technologies”), subsection 2) (“Wavelength Selective Switch (WSS)”), fourth paragraph in column 1). Chen further discloses that the four fibers carrying the four wavelength groups are connected to the MEMS via circulators, and that through WSS configuration the capacity of each of these four links can be varied, the MEMS and WSS configurations being decided by a central manager that estimates traffic demand and pushes the appropriate configurations to the MEMS, WSS units, and top-of-rack switches (Chen, page 501, column 2, section III-B, paragraphs 4 and 5). The recited mapping relationship between the wavelength of the carrier and the group to which the second node belongs reads on the runtime-configurable partition of Chen that assigns each wavelength to the WSS output port whose fiber terminates at the MEMS port bearing a circuit to the destination top-of-rack switch.
Shen and Chen are analogous to the claimed invention as both are from the same field of endeavor of optical interconnection of computing nodes through a micro-electro-mechanical optical switch. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to convert the electrical signal of the GPU of Shen onto a carrier at the destination-associated wavelength of Chen and to forward that carrier to the micro-electro-mechanical system under the wavelength-to-port mapping of Chen. The motivation to combine Shen and Chen is that Chen teaches that this wavelength assignment permits all wavelengths at a source node to be multiplexed onto a single fiber and, after demultiplexing, delivered to individual ports at the destination nodes (Chen, page 501, section III-A (“Building Blocks”), first paragraph in column 1), which addresses the problem Shen identifies of increasingly frequent data transmission between GPU chips as the accelerator cluster scale grows (Shen, ¶ [0005]).
As to dependent Claim 7, the limitations of claims 1 and 3 from which claim 7 depends are rejected under the same rationale set forth above with respect to claims 1 and 3.
Regarding the additional limitations of claim 7, Shen teaches the optical transmission channel constructed between graphics processing units of different servers, in that an electrical signal sent by the GPU 113 is converted into an optical signal and reaches the destination graphics processing unit by way of the reflectors of the MEMS 15 (Shen, ¶ [0042], FIG. 4).
However, Shen does not teach:
"wherein each of the W groups corresponds to two preset wavelengths"
"and W is equal to 1/2 of a total quantity of available wavelengths in the WSS"
In the same field of endeavor, Chen teaches assigning a set quantity of preset wavelengths to each of the groups into which a WSS partitions its wavelengths, and teaches the relationship between the number of groups and the total quantity of available wavelengths in the WSS.
As to "wherein each of the W groups corresponds to two preset wavelengths," Chen discloses that a WSS consists of one common port and a plurality of wavelength ports, and that it partitions, runtime-configurably within a few milliseconds, the set of wavelengths coming in through the common port among the wavelength ports, the illustration given being a common port receiving eighty wavelengths and routing wavelengths 1–20 on port 1 and 30–40 on port 2 (Chen, page 500, section II-B (“Optical Networking Technologies”), subsection 2) (“Wavelength Selective Switch (WSS)”), fourth paragraph in column 1). Chen further discloses that each port facing the optical interconnect has a transceiver associated with a fixed and unique wavelength for sending and receiving data, that this wavelength-to-port association is a design-time decision, and that the same set of wavelengths is recycled across top-of-rack switches (Chen, page 501, section III-B (“Putting It All Together: OSA-2560”), third paragraph in column 1; page 501, section III-A (“Building Blocks”), first paragraph in column 1). Chen further discloses that, in the disclosed instantiation, the 1×4 WSS splits the set of thirty-two wavelengths it sees into four groups, each group being transmitted on its own fiber to the MEMS (Chen, page 501, section III-B (“Putting It All Together: OSA-2560”), fourth paragraph in column 1). The recited preset wavelengths read on the fixed and unique design-time wavelengths of Chen, and the recited correspondence between each group and a quantity of those preset wavelengths reads on the partition of Chen by which each of the four groups carries a defined subset of the thirty-two wavelengths.
As to "and W is equal to 1/2 of a total quantity of available wavelengths in the WSS," Chen discloses that within the C-band and with 100 GHz dense wavelength division multiplexing channel spacing, typically forty or more wavelength channels can be transmitted over a single optical fiber (Chen, page 500, section II-B (“Optical Networking Technologies”), subsection 1 (“Wavelength Division Multiplexing (WDM)”), third paragraph in column 1). Chen further discloses a wavelength-assignment procedure in which the number of wavelengths required per link is determined, each link is replaced with that many parallel links, edge-coloring is applied with the wavelengths as the colors, and the least-used color is removed when the coloring exceeds the available wavelength count (Chen, page 503, section IV). Chen expressly recognizes that the number of groups is a variable subject to optimization, disclosing that a larger node degree enables one top-of-rack switch to connect to more other top-of-rack switches simultaneously and thereby achieve higher performance, but that, given a 320-port MEMS, a larger degree also means that fewer top-of-rack switches can be supported, and that the disclosed degree of four is chosen as a tradeoff between network size and performance (Chen, page 501, section III-B (“Putting It All Together: OSA-2560”), fifth paragraph in column 1, and paragraph 6).
Chen therefore teaches that the quantity of preset wavelengths assigned to each group, and correspondingly the ratio of the number of groups W to the total quantity of available wavelengths in the WSS, are variables that the artisan sets according to the desired tradeoff between node degree and per-link capacity. The particular values recited in claim 7, namely two preset wavelengths per group and W equal to one-half of the total quantity of available wavelengths, are the result of routine optimization of those variables. It has been held that where the general conditions of a claim are disclosed in the prior art, discovering the optimum or workable ranges of a result-effective variable involves only routine skill in the art. In re Aller, 220 F.2d 454, 456, 105 USPQ 233, 235 (CCPA 1955); see MPEP § 2144.05(II). The recited values are further consistent with the bidirectional operation taught by Chen, which requires that each circuit over the MEMS be bidirectional and employs optical circulators between the node and the MEMS ports for that purpose (Chen, page 501, section III-A (“Building Blocks”), “Efficient Port Usage”, second paragraph in column 1).
Shen and Chen are analogous to the claimed invention as both are from the same field of endeavor of optical interconnection of computing nodes through a micro-electro-mechanical optical switch. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to configure the wavelength selective switch of Chen, as applied to the AI training network of Shen, such that each group corresponds to two preset wavelengths and W equals one-half of the total quantity of available wavelengths. The motivation to combine Shen and Chen is that Chen teaches that this partition is a tradeoff between network size and per-link performance that the artisan selects to suit the deployment (Chen, page 501, section III-B (“Putting It All Together: OSA-2560”), last paragraph in column 1, continuing to the first paragraph in column 2), which addresses the problem Shen identifies of increasingly frequent data transmission between GPU chips as the accelerator cluster scale grows (Shen, ¶ [0005]).
As to dependent Claim 8, the limitations of claim 1 from which claim 8 depends are rejected under the same rationale set forth above with respect to claim 1.
Claim 8 recites two alternative configurations joined by the conjunction "or." Under the broadest reasonable interpretation, disclosure of either alternative in the prior art is sufficient to render the claim unpatentable. The second alternative is addressed below.
Regarding the additional limitations of claim 8, Shen teaches:
"wherein training data in any two of the S×C processors is different" — Shen discloses that the master server 11 splits the training data set and the data flow diagram into several parts and separately sends the several parts to the slave server 12, the slave server 13, and the slave server 14, so that each server shares a part of the training task (Shen, ¶ [0037], FIG. 3, step S11). Shen further discloses that, with four servers and an even allocation, each server processes a one-quarter training data set and the corresponding one-quarter data flow diagram, designated a level-1 data subset (Shen, ¶ [0038]). Shen further discloses that the CPUs of the slave servers split the level-1 training data subsets again based on a quantity of GPUs into a plurality of level-2 data subsets, that the level-2 data subsets are sent to the corresponding GPUs, and that each GPU is instructed to perform calculation on a received level-2 data subset (Shen, ¶ [0039], FIG. 3, step S12), the GPU 123 and the GPU 124 each undertaking a one-eighth training data set (Shen, ¶ [0040]). Each of the graphics processing units of Shen therefore operates on a distinct portion of the training data set.
Shen teaches a collective communication manner between the processors, in that the respective learning results of the plurality of accelerators are combined and distributed to each accelerator again before a next iteration is performed (Shen, ¶ [0004]).
However, Shen does not teach "and a collective communication manner between the S×C processors is allreduce."
In the same field of endeavor, Lu teaches "and a collective communication manner between the S×C processors is allreduce." Lu discloses that in data parallelism massive amounts of data are distributed to different computing devices, each containing the same training model, and that after calculating its own data each node must exchange its local gradients with the other nodes, for which purpose several synchronization algorithms have been proposed (Lu, page 4248, Section II-A (“Data Parallelism”), first paragraph in column 1). Lu further discloses the ring-based synchronization algorithm, identified as ring allreduce, in which each device functions as both parameter server and worker, in which at each iteration a device passes its local gradients to the next device in the ring and receives the gradients sent by the previous device, and in which the model parameters can be updated after each device has received the gradients from all other devices (Lu, page 4248, Section II-A (“Data Parallelism”), subsection 2) (“Ring-based Synchronization Algorithm, last paragraph in column 1, continuing to the first paragraph in column 2). Lu further discloses the torus-based synchronization algorithm, in which the synchronization algorithm divides the data in the device into equal parts according to the number of dimensions and different parts of the data perform ring allreduce in each dimension simultaneously (Lu, page 4248, Section II-A (“Data Parallelism”), subsection 3) (“Torus-Based Synchronization Algorithm”), second paragraph in column 2). The recited collective communication manner of allreduce reads on the ring allreduce synchronization algorithm of Lu, which is applied among computing devices that each hold the same training model and different training data.
Shen and Lu are analogous to the claimed invention as both are from the same field of endeavor of interconnecting computing nodes that exchange training results in a distributed machine learning system by means of a micro-electro-mechanical optical switch. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to synchronize the graphics processing units of Shen using the ring allreduce synchronization algorithm of Lu. The motivation to combine Shen and Lu is that Lu teaches that ring allreduce decentralizes synchronization such that no central device is required to aggregate the gradients of all workers and the bandwidth of each device is fully utilized, thereby avoiding the network performance bottleneck created by centralized processing of gradients at a parameter server (Lu, page 4248, Section II-A (“Data Parallelism”), subsection 1) (“Parameter Server-based Synchronization Algorithm”), and subsection 2) (“Ring-based Synchronization Algorithm, last paragraph in column 1 continuing to the first paragraph in column 2), which addresses the problem Shen identifies of increasingly frequent data transmission between GPU chips as the accelerator cluster scale grows (Shen, ¶ [0005]).
Claim 14 is rejected under 35 U.S.C. 103 as being unpatentable over Shen et al. (Shen), WO 2020/199560 A1, "AI Training Network and Method," published October 8, 2020, an English translation of which is available as US 2022/0012590 A1, published January 13, 2022, in view of Lu et al. (Lu), Non-Patent Literature, "X-NEST: A Scalable, Flexible, and High-Performance Network Architecture for Distributed Machine Learning," Journal of Lightwave Technology, vol. 39, no. 13, pp. 4247–4254, published July 1, 2021, cited in the Information Disclosure Statement filed February 26, 2025.
As to dependent Claim 14, the limitations of claim 10 from which claim 14 depends are rejected under the same rationale set forth above with respect to claim 10.
Claim 14 recites two alternative configurations joined by the conjunction "or." Under the broadest reasonable interpretation, disclosure of either alternative in the prior art is sufficient to render the claim unpatentable. The second alternative is addressed below.
Regarding the additional limitations of claim 14, Shen teaches:
"performing, by the first processor, the model training in the first processor to obtain intermediate data of the first processor" — Shen discloses that the level-2 data subsets and the level-2 data flow sub-diagrams are sent to the corresponding GPUs and that each GPU is instructed to perform calculation on a received level-2 data subset based on a received level-2 data flow sub-diagram (Shen, ¶ [0039], FIG. 3, step S12). Shesn further discloses that the first graphics processing unit performs AI training calculation on a first dataset based on a first data flow diagram (Shen, ¶ [0006]), and that the training data set and the data flow diagram undertaken by the GPU 113 are referred to as a first training data set and a first data flow diagram (Shen, ¶ [0043]). The recited intermediate data of the first processor reads on the result of the calculation performed by the GPU 113 of Shen on its level-2 data subset.
"training data in the first processor and the second processor is different" — Shen discloses that the master server 11 splits the training data set into several parts which are separately sent to the slave servers so that each server shares a part of the training task (Shen, ¶ [0037], FIG. 3, step S11), that each server processes a one-quarter training data set (Shen, ¶ [0038]), that the CPUs of the slave servers split the level-1 training data subsets again based on a quantity of GPUs into level-2 data subsets that are sent to the corresponding GPUs (Shen, ¶ [0039], FIG. 3, step S12), and that the GPU 123 and the GPU 124 each undertake a one-eighth training data set (Shen, ¶ [0040]). The GPU 113 and the GPU 123 of Shen therefore operate on distinct portions of the training data set.
Shen teaches determining the first target data based on a collective communication manner, in that the respective learning results of the plurality of accelerators are combined and distributed to each accelerator again before a next iteration is performed (Shen, ¶ [0004]).
However, Shen does not teach:
"determining, by the first processor, the first target data based on a collective communication manner and the intermediate data of the first processor, wherein the first target data is all or a part of the intermediate data of the first processor"
"and the collective communication manner is allreduce"
In the same field of endeavor, Lu teaches "determining, by the first processor, the first target data based on a collective communication manner and the intermediate data of the first processor, wherein the first target data is all or a part of the intermediate data of the first processor." Lu discloses that in data parallelism each node, after calculating its own data, must exchange its local gradients with other nodes, and that several synchronization algorithms have been proposed to govern that exchange (Lu, page 4248, Section II-A (“Data Parallelism”), first paragraph in column 1). Lu further discloses the torus-based synchronization algorithm, in which the synchronization algorithm divides the data in the device into equal parts according to the number of dimensions, different parts of the data then perform ring allreduce in each dimension simultaneously, and, after the data have completed parameter synchronization in one dimension, the same operation is performed in another dimension until every part of the data has completed parameter synchronization in all dimensions (Lu, page 4248, Section II-A (“Data Parallelism”), subsection 3) (“Torus-Based Synchronization Algorithm”), second paragraph in column 2). The recited determining of the first target data based on a collective communication manner and the intermediate data reads on the division of a device's locally computed data into parts under the governing synchronization algorithm of Lu, and the recited first target data being a part of the intermediate data reads on the individual part of that divided data which the device transmits in a given dimension.
Lu further teaches "and the collective communication manner is allreduce." Lu discloses the ring-based synchronization algorithm, identified as ring allreduce, in which each device functions as both parameter server and worker, in which at each iteration a device passes its local gradients to the next device in the ring and receives the gradients sent by the previous device, and in which the model parameters can be updated after each device has received the gradients from all other devices (Lu, page 4248, Section II-A (“Data Parallelism”), subsection 2) (“Ring-based Synchronization Algorithm, last paragraph in column 1, continuing to the first paragraph in column 2).
Shen and Lu are analogous to the claimed invention as both are from the same field of endeavor of interconnecting computing nodes that exchange training results in a distributed machine learning system by means of a micro-electro-mechanical optical switch. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to determine the calculation result that the GPU 113 of Shen sends to the GPU 123 according to the allreduce synchronization algorithm of Lu. The motivation to combine Shen and Lu is that Lu teaches that ring allreduce decentralizes synchronization such that no central device is required to aggregate the gradients of all workers and the bandwidth of each device is fully utilized, thereby avoiding the network performance bottleneck created by centralized processing of gradients at a parameter server (Lu, page 4248, Section II-A (“Data Parallelism”), subsection 1) (“Parameter Server-based Synchronization Algorithm”), and subsection 2) (“Ring-based Synchronization Algorithm, last paragraph in column 1 continuing to the first paragraph in column 2), which addresses the problem Shen identifies of increasingly frequent data transmission between GPU chips as the accelerator cluster scale grows (Shen, ¶ [0005]).
Claim 15-17 is rejected under 35 U.S.C. 103 as being unpatentable over Shen et al. (Shen), WO 2020/199560 A1, "AI Training Network and Method," published October 8, 2020, an English translation of which is available as US 2022/0012590 A1, published January 13, 2022, in view of Lepikhin et al. (Lepikhin), Non-Patent Literature, "GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding," arXiv:2006.16668v1, published June 30, 2020, cited in the Information Disclosure Statement filed January 13, 2025, and further in view of Lu et al. (Lu), Non-Patent Literature, "X-NEST: A Scalable, Flexible, and High-Performance Network Architecture for Distributed Machine Learning," Journal of Lightwave Technology, vol. 39, no. 13, pp. 4247–4254, published July 1, 2021, cited in the Information Disclosure Statement filed February 26, 2025.
As to dependent Claim 15, the limitations of claims 10 and 14 from which claim 15 depends are rejected under the same rationale set forth above, the first alternative of claim 14 being relied upon for this claim.
Claim 15 recites "the alltoall," and accordingly the first alternative of claim 14 is relied upon for this claim.
Regarding the additional limitations of claim 15, Shen teaches determining the first target data for transmission to a particular destination processor, in that a sub-data flow diagram received by a GPU includes a sending operator where the GPU needs to send data and a receiving operator where the GPU needs to receive data, and the recorded transmission information identifies a source GPU and a destination GPU (Shen, ¶ [0055]).
However, Shen does not teach:
"the collective communication manner is alltoall"
"dividing, by the first processor, the intermediate data of the first processor based on the alltoall and a total quantity of processors corresponding to the alltoall"
"wherein the processors corresponding to the alltoall comprise the first processor and the second processor"
"wherein a quantity of data parts after division is equal to the total quantity of processors, and the data after division comprises the first target data corresponding to the second processor"
In the same field of endeavor, Lepikhin teaches "the collective communication manner is alltoall." Lepikhin discloses that AllReduce and AllToAll are the two critical collective communication operators in the Mixture-of-Experts model, that AllToAll is used in resharding, and that the dispatch and combine operations of the model are AllToAll operators (Lepikhin, page 22, section 5.3 (“Communication Microbenchmarks and Per-Operator Scalability”). Lepikhin further discloses that the experts in the MoE layer cannot be replicated across all the devices because of their size and that the only viable strategy is to shard the experts into many devices, while the attention layer is parallelized by splitting along the batch dimension (Lepikhin, page 6, section 3 (“Highly Parallel Implementation using GShard”), the paragraph beginning “Next, we annotate the linear algebra computation”), so that the devices participating in the AllToAll hold both different model shards and different data.
Lepikhin further teaches "dividing, by the first processor, the intermediate data of the first processor based on the alltoall and a total quantity of processors corresponding to the alltoall." Lepikhin discloses that the AllToAll operator logically splits the input of each participant along one dimension and then sends each piece to a different participant, and that on receiving the data pieces from the others each participant concatenates the pieces to produce its result (Lepikhin, page 11, section 3.3.1 (“Communication Primitives”), the “AllToAll” paragraph). Lepikhin further discloses that the input tensor is partitioned along the group dimension across D devices by the annotation split(inputs, 0, D), where D is the device count, and that after the dispatched expert inputs are computed a further split changes the sharding from the group dimension to the expert dimension (Lepikhin, page 9, section 3.2 (“GShard Annotation API for Parallel Execution”), the text following Algorithm 2 and accompanying code listing). The recited dividing based on a total quantity of processors corresponding to the alltoall reads on the splitting of the tensor of Lepikhin along a dimension whose extent is the device count D.
Lepikhin further teaches "wherein the processors corresponding to the alltoall comprise the first processor and the second processor." Lepikhin discloses that the AllToAll operator operates on the input of each participant and sends each piece to a different participant (Lepikhin, page 11, section 3.3.1), and that the model is trained on 2048 TPU v3 accelerators (Lepikhin, page 1, Abstract). The recited processors corresponding to the alltoall read on the participating devices of Lepikhin, which include both the sending device and each receiving device.
Lepikhin further teaches "wherein a quantity of data parts after division is equal to the total quantity of processors, and the data after division comprises the first target data corresponding to the second processor." Lepikhin discloses that AllToAll splits the input of each participant along one dimension and sends each piece to a different participant, so that one piece is produced for and addressed to each participant (Lepikhin, page 11, section 3.3.1), and illustrates this resharding by all-to-all in which each partition contributes a distinct piece to every other partition (Lepikhin, page 12, Figure 4). Lepikhin further discloses that the expert capacity is set to O(N/E) for N tokens and E experts and that the gating function keeps a running counter of how many tokens are dispatched to each expert, so that the dispatched portion corresponding to a given destination is identified (Lepikhin, pages 6–7, section 2.2 (“Position-wise Mixture-of-Experts Layer”), "Expert capacity" and “Local group dispatching” bullets). The recited first target data corresponding to the second processor reads on the piece of Lepikhin that is addressed to the particular destination device.
Shen and Lepikhin are analogous to the claimed invention as both are from the same field of endeavor of distributing a neural network model and its training data across a plurality of processors that exchange training results during training. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to divide the calculation result of the GPU 113 of Shen by the device count and to exchange the resulting pieces among the graphics processing units using the AllToAll operator of Lepikhin. The motivation to combine Shen and Lepikhin is that Lepikhin teaches that sharding the model across many devices with AllToAll resharding is the only viable strategy for models too large to replicate, and that the cost of AllToAll increases only sublinearly as the number of partitions grows (Lepikhin, page 6, section 3 (“Highly Parallel Implementation using GShard”), the paragraph beginning “Next, we annotate the linear algebra computation”; page 11, section 3.3.1), which addresses the problem Shen identifies of increasingly frequent data transmission between GPU chips as the accelerator cluster scale grows (Shen, ¶ [0005]).
As to dependent Claim 16, the limitations of claims 10, 14, and 15 from which claim 16 depends are rejected under the same rationale set forth above with respect to claims 10, 14, and 15.
Regarding the additional limitations of claim 16, Shen teaches:
"wherein the second processor comprises C processors" — Shen discloses that a server includes a central processing unit and a graphics processing unit, that the server 11 includes a CPU 111, a CPU 112, a GPU 113, and a GPU 114 (Shen, ¶ [0029], FIG. 1), and that the server 12 includes a GPU 123 and a GPU 124, which separately undertake a one-eighth training data set and a one-eighth data flow diagram (Shen, ¶ [0040], FIG. 1). Shen further discloses that the CPUs of the slave servers split the level-1 training data subsets and the level-1 data flow sub-diagrams again based on a quantity of GPUs (Shen, ¶ [0039], FIG. 3, step S12). The recited C processors read on the two graphics processing units contained in the server 12 of Shen.
Shen teaches identifying the source and the destination of a transmission among the processors of the array, in that the recorded transmission information comprises a four-tuple in which one element describes a source GPU that sends the packet and another describes a destination GPU that receives the packet, and in that a model is trained for each pair of specific source and destination GPUs (Shen, ¶¶ [0055], [0057], [0058]).
However, Shen does not teach:
"wherein the alltoall corresponds to S nodes, the first node is an s1th node in the S nodes, and the second node is an s2th node in the S nodes"
"wherein s1 and s2 are set to every integer in [0, S], and s1 is less than s2"
"and the first target data is (s2×C)th data to (s2×C+C−1)th data in S×C pieces of data after the division"
In the same field of endeavor, Lepikhin teaches "wherein the alltoall corresponds to S nodes, the first node is an s1th node in the S nodes, and the second node is an s2th node in the S nodes." Lepikhin discloses that the AllToAll operator logically splits the input of each participant along one dimension and then sends each piece to a different participant, and that on receiving the data pieces from the others each participant concatenates the pieces to produce its result (Lepikhin, page 11, section 3.3.1 (“Communication Primitives”), the “AllToAll” paragraph). Lepikhin further discloses that the input tensor is partitioned along the group dimension across D devices by the annotation split(inputs, 0, D), where D is the device count (Lepikhin, page 9, section 3.2 (“GShard Annotation API for Parallel Execution”), the text following Algorithm 2 and accompanying code listing), and illustrates the resharding by all-to-all in which each indexed partition contributes a distinct piece to every other indexed partition (Lepikhin, page 12, Figure 4). The recited S nodes read on the D participating devices of Lepikhin, and the recited s1th node and s2th node read on the indexed source and destination partitions of Lepikhin.
Lepikhin further teaches "wherein s1 and s2 are set to every integer in [0, S], and s1 is less than s2." Lepikhin discloses that AllToAll sends each piece of the split input to a different participant and that each participant concatenates the pieces received from the others (Lepikhin, page 11, section 3.3.1), so that the operator is performed over the complete range of participant indices. Lepikhin further discloses that the sharding annotations apply to all dimensions in the same way and that the user works with full shapes without needing to address uneven partitioning, the partitioner assigning the shards across the device index range (Lepikhin, page 9, section 3.2, paragraphs preceding and following the code listing).
Lepikhin further teaches "and the first target data is (s2×C)th data to (s2×C+C−1)th data in S×C pieces of data after the division." Lepikhin discloses that after the dispatched expert inputs are computed a further split changes the sharding from the group dimension to the expert dimension, so that the pieces are addressed to destinations by an index computed from the sharded dimensions (Lepikhin, page 9, section 3.2 (“GShard Annotation API for Parallel Execution”), the text following Algorithm 2). Lepikhin further discloses that the expert capacity is set to O(N/E) for N tokens and E experts and that the gating function keeps a running counter of how many tokens are dispatched to each expert, so that a determinate contiguous allocation of the divided data corresponds to each destination (Lepikhin, pages 6–7, section 2.2 (“Position-wise Mixture-of-Experts Layer”), "Expert capacity" and “Local group dispatching” bullets). To the extent that the particular index arithmetic recited in claim 16 is not expressly disclosed, the assignment of the (s2×C)th through (s2×C+C−1)th of S×C divided pieces to the C processors of the s2th node is the ordinary and predictable consequence of indexing the participating processors first by node and then by position within the node, and is a matter of obvious design choice. See MPEP § 2144.04(VI).
Shen and Lepikhin are analogous to the claimed invention as both are from the same field of endeavor of distributing a neural network model and its training data across a plurality of processors that exchange training results during training. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to index the servers and the graphics processing units of Shen and to address each piece of the divided calculation result to the corresponding indexed destination in the manner taught by Lepikhin. The motivation to combine Shen and Lepikhin is that Lepikhin teaches that indexed sharding across the device count frees the developer from system optimization effort and avoids baking low-level parallel implementation details into the model code (Lepikhin, page 6, section 3 (“Highly Parallel Implementation using GShard”), the paragraph beginning “Next, we annotate the linear algebra computation”), which addresses the problem Shen identifies of increasingly frequent data transmission between GPU chips as the accelerator cluster scale grows (Shen, ¶ [0005]).
As to dependent Claim 17, the limitations of claims 10 and 14 from which claim 17 depends are rejected under the same rationale set forth above with respect to claims 10 and 14. Claim 17 recites "the alltoall," and accordingly the first alternative of claim 14 is relied upon for this claim.
Regarding the additional limitations of claim 17, Shen teaches identifying the source and the destination of a transmission among the processors of the array, in that the recorded transmission information comprises a four-tuple in which one element describes a source GPU that sends the packet and another describes a destination GPU that receives the packet, and in that a model is trained for each pair of specific source and destination GPUs (Shen, ¶¶ [0055], [0057], [0058]).
However, Shen does not teach:
"wherein the alltoall corresponds to W groups"
"the first node is an s1th node of a w1th group in the W groups, and the second node is an s2th node of a w2th group in the W groups"
"wherein w1 is set to every integer in [0, W−1], and w2=w1+offset, wherein offset=((s2%W)−(s1%W))%W"
In the same field of endeavor, Lepikhin teaches "wherein the alltoall corresponds to W groups." Lepikhin discloses that the AllToAll operator logically splits the input of each participant along one dimension and then sends each piece to a different participant, and that on receiving the data pieces from the others each participant concatenates the pieces to produce its result (Lepikhin, page 11, section 3.3.1 (“Communication Primitives”), the “AllToAll” paragraph). Lepikhin further discloses that the gating function partitions all tokens in a training batch evenly into G groups, and that the input tensor is partitioned along that group dimension across D devices by the annotation split(inputs, 0, D), where D is the device count, a further split thereafter changing the sharding from the group dimension to the expert dimension (Lepikhin, pages 6–7, section 2.2, "Local group dispatching"; page 9, section 3.2 (“GShard Annotation API for Parallel Execution”), the text following Algorithm 2 and accompanying code listing). The recited W groups read on the groups of Lepikhin across which the all-to-all resharding is performed.
Lepikhin further teaches "the first node is an s1th node of a w1th group in the W groups, and the second node is an s2th node of a w2th group in the W groups." Lepikhin discloses that the sharding annotations apply to all dimensions in the same way, that different parts of the model may be partitioned in different ways, and that the partitioner assigns shards across the indexed dimensions so that a given piece is addressed by its position within the group dimension and within the expert dimension (Lepikhin, page 9, section 3.2, paragraphs preceding and following the code listing). Lepikhin further illustrates the resharding by all-to-all in which each indexed partition contributes a distinct piece to every other indexed partition (Lepikhin, page 12, Figure 4). The recited two-index addressing of a node by its group index and by its position within that group reads on the two-dimensional shard addressing of Lepikhin.
Lepikhin further teaches "wherein w1 is set to every integer in [0, W−1], and w2=w1+offset, wherein offset=((s2%W)−(s1%W))%W." Lepikhin discloses that AllToAll sends each piece of the split input to a different participant and that each participant concatenates the pieces received from the others, so that the operator is performed across the complete range of group indices (Lepikhin, page 11, section 3.3.1), and that the change of sharding from the group dimension to the expert dimension determines, for each source shard, the destination shard to which its piece is routed (Lepikhin, page 9, section 3.2 (“GShard Annotation API for Parallel Execution”), the text following Algorithm 2). To the extent that the particular modulo arithmetic recited in claim 17 is not expressly disclosed, the scheduling of a group-wise all-to-all by a per-round offset computed modulo the group count, so that in each round every group exchanges with a distinct partner group and no two groups contend for the same partner, is the ordinary and predictable consequence of ordering an all-to-all exchange among indexed groups, and is a matter of obvious design choice. See MPEP § 2144.04(VI).
Shen and Lepikhin are analogous to the claimed invention as both are from the same field of endeavor of distributing a neural network model and its training data across a plurality of processors that exchange training results during training. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to schedule the all-to-all exchange of the divided calculation results among the graphics processing units of Shen by the group-indexed offset addressing taught by Lepikhin. The motivation to combine Shen and Lepikhin is that Lepikhin teaches that indexed sharding across the group and device dimensions frees the developer from system optimization effort and avoids baking low-level parallel implementation details into the model code (Lepikhin, page 6, section 3 (“Highly Parallel Implementation using GShard”), the paragraph beginning “Next, we annotate the linear algebra computation”), which addresses the problem Shen identifies of increasingly frequent data transmission between GPU chips as the accelerator cluster scale grows (Shen, ¶ [0005]).
The combination of Shen and Lepikhin, however, does not teach that the nodes participating in the exchange are physically organized into the W groups such that each node is identified by a group index and a position within that group.
In the same field of endeavor, Lu teaches that organization. Lu discloses that a scalable topology is achieved by a hierarchical network with three levels, namely group, pod, and system, that there are k pods in the network, that each pod has k groups, and that each group has a top-of-rack switch connected to k computing nodes or servers, giving a total of k² top-of-rack switches and k³ computing nodes (Lu, page 4248, section III-A (“Topology”), first paragraph in column 2, continuing to page 4249, column 1). Lu further discloses that each group has a coordinate G (x, y), where x represents the pod identifier to which the group belongs and y represents the group identifier within the pod, and that pods are connected through MEMS switches with 100 Gbps optical links, top-of-rack switches having the same x coordinate being connected to 2m different MEMS switches and those having the same y coordinate being connected to the same MEMS switch (Lu, page 4248, section III-A (“Topology”), first paragraph in column 2, continuing to page 4249, column 1; page 4249, FIG. 1 and caption). The recited W groups read on the k pods of Lu, and the recited identification of a node by a group index and a position within that group reads on the coordinate addressing of Lu.
Shen, Lepikhin, and Lu are analogous to the claimed invention as all three are from the same field of endeavor of distributing a neural network model and its training data across a plurality of interconnected processors that exchange training results during training. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to organize the servers of Shen into the coordinate-addressed groups taught by Lu and to schedule the all-to-all exchange of Lepikhin across those groups. The motivation to combine Shen, Lepikhin, and Lu is that Lu teaches that the coordinate-addressed hierarchical group and pod arrangement, with the groups interconnected through MEMS switches, yields a scalable topology reaching k³ computing nodes (Lu, page 4248, column 2, section III-A, paragraph 1), which addresses the problem Shen identifies of an accelerator cluster of increasing scale (Shen, ¶ [0005]).
Claims 18 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Shen et al. (Shen), WO 2020/199560 A1, "AI Training Network and Method," published October 8, 2020, an English translation of which is available as US 2022/0012590 A1, published January 13, 2022, in view of Lu et al. (Lu), Non-Patent Literature, "X-NEST: A Scalable, Flexible, and High-Performance Network Architecture for Distributed Machine Learning," Journal of Lightwave Technology, vol. 39, no. 13, pp. 4247–4254, published July 1, 2021, cited in the Information Disclosure Statement filed February 26, 2025, and further in view of Cevolani et al. (Cevolani), US 2021/0311808 A1, "Control of Data Transfer Between Processing Nodes," published October 7, 2021.
As to dependent Claim 18, the limitations of claims 10 and 14 from which claim 18 depends are rejected under the same rationale set forth above with respect to claims 10 and 14. Claim 18 recites "the allreduce," and accordingly the second alternative of claim 14 is relied upon for this claim.
Regarding the additional limitations of claim 18, Shen teaches:
"wherein the first processor is an ith processor in the first node, and the second processor is an ith processor in the second node" — Shen discloses that the server 11 includes a CPU 111, a CPU 112, a GPU 113, and a GPU 114, and that the server 12 includes a CPU 121, a CPU 122, a GPU 123, and a GPU 124 (Shen, ¶¶ [0029], [0040], FIG. 1). Shen further discloses that the CPUs of the slave servers split the level-1 training data subsets and the level-1 data flow sub-diagrams again based on a quantity of GPUs, that the level-2 subsets are sent to corresponding GPUs, and that the GPU 123 and the GPU 124 separately undertake a one-eighth training data set and a one-eighth data flow diagram (Shen, ¶¶ [0039], [0040], FIG. 3, step S12). Shen further discloses that the input of the GPU 123 of the server 12 depends on the output of the GPU 113 of the server 11 (Shen, ¶ [0042], FIG. 4). The GPU 113 and the GPU 123 of Shen occupy corresponding ordinal positions among the graphics processing units of their respective servers.
Shen teaches obtaining data of other processors in the first node through a channel within that node, in that where the two GPUs that communicate with each other belong to the same server the communication may be performed by using a bus inside the server (Shen, ¶ [0031]), and in that communication between the CPU and the GPU within a server is performed by using a peripheral component interconnect express bus (Shen, ¶ [0029]).
However, Shen does not teach:
"the collective communication manner is allreduce"
"dividing, by the first processor, the intermediate data of the first processor based on the allreduce and a total quantity C of processors in the first node, to obtain C pieces of data"
"obtaining, by the first processor, ith data of other (C−1) processors in the first node through the intra-node channel of the first node"
"obtaining the first target data after the first processor performs summation on ith data in the C pieces of data and the ith data of the other (C−1) processors in the first node"
In the same field of endeavor, Lu teaches "the collective communication manner is allreduce." Lu discloses the ring-based synchronization algorithm, identified as ring allreduce, in which each device functions as both parameter server and worker, in which at each iteration a device passes its local gradients to the next device in the ring and receives the gradients sent by the previous device, and in which the model parameters can be updated after each device has received the gradients from all other devices (Lu, page 4248, Section II-A (“Data Parallelism”), subsection 2) (“Ring-based Synchronization Algorithm, last paragraph in column 1, continuing to the first paragraph in column 2).
Lu further teaches dividing the locally computed data into parts under the allreduce, disclosing the torus-based synchronization algorithm in which the synchronization algorithm divides the data in the device into equal parts according to the number of dimensions, different parts of the data then perform ring allreduce in each dimension simultaneously, and, after the data have completed parameter synchronization in one dimension, the same operation is performed in another dimension until every part of the data has completed parameter synchronization in all dimensions (Lu, page 4248, Section II-A (“Data Parallelism”), subsection 3) (“Torus-Based Synchronization Algorithm”), second paragraph in column 2). Lu further discloses that this algorithm has better parallelism and fault tolerance than the ring-based algorithm (Lu, page 4248, Section II-A (“Data Parallelism”), subsection 3) (“Torus-Based Synchronization Algorithm”), second paragraph in column 2).
Shen and Lu are analogous to the claimed invention as both are from the same field of endeavor of interconnecting computing nodes that exchange training results in a distributed machine learning system by means of a micro-electro-mechanical optical switch. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to synchronize the graphics processing units of Shen using the allreduce synchronization algorithm of Lu, in which each device divides its locally computed data into parts that are synchronized dimension by dimension. The motivation to combine Shen and Lu is that Lu teaches that ring allreduce decentralizes synchronization such that no central device is required to aggregate the gradients of all workers and the bandwidth of each device is fully utilized, and that dividing the data and synchronizing per dimension further improves parallelism and fault tolerance (Lu, page 4248, column 1, section II-A, subsection 2, and column 2, subsection 3), which addresses the problem Shen identifies of increasingly frequent data transmission between GPU chips as the accelerator cluster scale grows (Shen, ¶ [0005]).
The combination of Shen and Lu, however, does not teach "dividing, by the first processor, the intermediate data of the first processor based on the allreduce and a total quantity C of processors in the first node, to obtain C pieces of data"; "obtaining, by the first processor, ith data of other (C−1) processors in the first node through the intra-node channel of the first node"; and "obtaining the first target data after the first processor performs summation on ith data in the C pieces of data and the ith data of the other (C−1) processors in the first node."
In the same field of endeavor, Cevolani teaches "dividing, by the first processor, the intermediate data of the first processor based on the allreduce and a total quantity C of processors in the first node, to obtain C pieces of data." Cevolani discloses a data processing system comprising a plurality of processing nodes, each comprising at least one memory configured to store an array of data items, wherein each of the plurality of first processing nodes belongs to at least two different sets of processing nodes, and wherein each processing node takes part in one or more reduce scatter collectives using the respective array of data items to obtain a reduced subset of an array of data items, each reduce scatter collective being performed between the processing nodes of a different one of those sets (Cevolani, ¶ [0009]). Cevolani further discloses that by first performing each of the reduce-scatter collectives between processing nodes belonging to a set of processing nodes, each of the processing nodes in a given set receives a different subset of the array of data (Cevolani, ¶ [0010]). Cevolani further discloses that each stored element of the array comprises weight updates for a neural network (Cevolani, ¶ [0023]) and that the processing nodes in a rack each store a different set of the full gradients for a machine learning model (Cevolani, ¶ [0064], FIG. 24). The recited dividing based on a total quantity C of processors in the first node to obtain C pieces of data reads on the reduce-scatter of Cevolani performed within a set of processing nodes, which apportions the array into one subset per member of that set.
Cevolani further teaches "obtaining, by the first processor, ith data of other (C−1) processors in the first node through the intra-node channel of the first node." Cevolani discloses that the reduce scatter collective is performed between processing nodes of a first of the respective two sets of processing nodes, prior to the all-reduce collective performed between processing nodes of the second set (Cevolani, ¶ [0024]). Cevolani further discloses processing nodes in a rack, each of which stores a different set of the full gradients for a machine learning model, and thereafter each of which stores a subset of reduced gradients for that model, the all-reduce across racks being performed subsequently between gateways in different racks (Cevolani, ¶¶ [0064]–[0066], FIGS. 24, 25, and 26). The recited obtaining of the ith data of the other (C−1) processors through the intra-node channel reads on the intra-set exchange of Cevolani, in which each member of the set receives from every other member of that set the portion of the array corresponding to its own index position.
Cevolani further teaches "obtaining the first target data after the first processor performs summation on ith data in the C pieces of data and the ith data of the other (C−1) processors in the first node." Cevolani discloses that the operator performs elementwise reduction over the inputs from all participants and that, following the reduce-scatter collective of the all-reduce collective and prior to the all-gather collective, operations are performed on each stored element of the array of data items stored by the respective first processing node to modify the data of those stored elements (Cevolani, ¶ [0022]). Cevolani further discloses that the modifying operations comprise providing updated weights of the neural network using the weight updates (Cevolani, ¶ [0023]), that an all-reduce function is implemented by a reduce-scatter step followed by an all-gather step (Cevolani, ¶¶ [0053]–[0054], FIGS. 15 and 16A), and that each processing node in the rack thereafter stores a subset of reduced gradients (Cevolani, ¶ [0065], FIG. 25). The recited first target data obtained after summation reads on the reduced subset of Cevolani held by the processing node following the intra-set reduce-scatter.
Shen, Lu, and Cevolani are analogous to the claimed invention as all three are from the same field of endeavor of exchanging and aggregating training results among a plurality of interconnected processors in a distributed machine learning system. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to structure the allreduce of Lu, as applied to the graphics processing units of Shen, as the intra-set reduce-scatter of Cevolani in which each processor of a server obtains and sums the correspondingly indexed portion held by the other processors of that same server. The motivation to combine Shen, Lu, and Cevolani is that Cevolani teaches that performing the reduce-scatter first within a set causes each node to hold only a subset of the array, so that the subsequent data exchange involves a smaller amount of data and results in lower bandwidth utilization (Cevolani, ¶ [0010]), which addresses the problem Shen identifies of increasingly frequent data transmission between GPU chips as the accelerator cluster scale grows (Shen, ¶ [0005]).
As to dependent Claim 19, the limitations of claims 10 and 14 from which claim 19 depends are rejected under the same rationale set forth above with respect to claims 10 and 14. Claim 19 recites "the allreduce," and accordingly the second alternative of claim 14 is relied upon for this claim.
Regarding the additional limitations of claim 19, Shen teaches:
"one node comprises C processors" — Shen discloses that a server includes a central processing unit and a graphics processing unit, that the server 11 includes a CPU 111, a CPU 112, a GPU 113, and a GPU 114 (Shen, ¶ [0029], FIG. 1), and that the server 12 includes a GPU 123 and a GPU 124 which separately undertake a one-eighth training data set and a one-eighth data flow diagram (Shen, ¶ [0040], FIG. 1). Shen further discloses that the CPUs of the slave servers split the level-1 training data subsets and the level-1 data flow sub-diagrams again based on a quantity of GPUs (Shen, ¶ [0039], FIG. 3, step S12).
"optical transmission channels that are between the first node and other (S−1) nodes in the group to which the first processor belongs and that are established by the MEMS" — Shen discloses that where the two GPUs that communicate with each other belong to different servers the communication needs to be performed by using a communication channel outside the servers, in other words by using the optical cross-connects of FIG. 1, and that data sent by the GPU 113 can reach the GPU 144 after successively passing through the OXC 15, the OXC 16, and the OXC 17 (Shen, ¶ [0031], FIG. 1). Shen further discloses that the optical cross-connect is a micro-electro-mechanical system (Shen, ¶ [0012]), that the MEMS 15 includes a micro-electro-mechanical controller 150 and two reflection lens arrays whose reflectors are physically adjustable in deflection angle, and that after the reflection angle of the reflector 153 is adjusted to 30° a channel between the GPU 113 and the GPU 123 is successfully established (Shen, ¶ [0042], FIG. 4), the reflection path 155-156-157 being thereby established (Shen, ¶ [0050]).
Shen teaches obtaining data of other processors within the first node, in that where the two GPUs that communicate with each other belong to the same server the communication may be performed by using a bus inside the server (Shen, ¶ [0031]).
However, Shen does not teach:
"wherein the allreduce corresponds to W groups, one group comprises S nodes"
"the first processor is an ith processor in a group to which the first processor belongs, and the second processor is an ith processor in a group to which the second processor belongs"
"dividing, by the first processor, the intermediate data of the first processor based on the allreduce and a total quantity S×C of processors in the group, to obtain S×C pieces of data"
"obtaining, by the first processor through the intra-node channel of the first node and/or optical transmission channels that are between the first node and other (S−1) nodes in the group to which the first processor belongs and that are established by the MEMS, ith data of other (S×C−1) processors in the group to which the first processor belongs"
"obtaining the first target data after the first processor performs summation on ith data in the S×C pieces of data and the ith data of the other (S×C−1) processors in the group to which the first processor belongs"
In the same field of endeavor, Lu teaches "wherein the allreduce corresponds to W groups, one group comprises S nodes." Lu discloses the ring-based synchronization algorithm, identified as ring allreduce, in which each device passes its local gradients to the next device in the ring and receives the gradients sent by the previous device, and in which the model parameters can be updated after each device has received the gradients from all other devices (Lu, page 4248, Section II-A (“Data Parallelism”), subsection 2) (“Ring-based Synchronization Algorithm, last paragraph in column 1, continuing to the first paragraph in column 2). Lu further discloses that a scalable topology is achieved by a hierarchical network with three levels, namely group, pod, and system, that there are k pods in the network, that each pod has k groups, and that each group has a top-of-rack switch connected to k computing nodes or servers, giving a total of k² top-of-rack switches and k³ computing nodes (Lu, page 4248, section III-A (“Topology”), first paragraph in column 2, continuing to page 4249, column 1). Lu further discloses that each group has a coordinate G (x, y), where x represents the pod identifier to which the group belongs and y represents the group identifier within the pod, that groups in the same pod are connected through a fully interconnected network over electrical links, and that pods are connected through MEMS switches with 100 Gbps optical links (Lu, page 4248, section III-A (“Topology”), first paragraph in column 2, continuing to page 4249, column 1; page 4249, FIG. 1 and caption). The recited W groups read on the pods of Lu, and the recited S nodes of one group read on the computing nodes connected to a top-of-rack switch within a group of Lu.
Lu further teaches "the first processor is an ith processor in a group to which the first processor belongs, and the second processor is an ith processor in a group to which the second processor belongs." Lu discloses the coordinate addressing G (x, y) by which each group is identified by its pod and by its position within that pod, and discloses that top-of-rack switches having the same y coordinate are connected to the same MEMS switch (Lu, page 4248, section III-A (“Topology”), first paragraph in column 2, continuing to page 4249, column 1). Under this addressing, the devices occupying corresponding positions in different groups are those joined through the same MEMS switch.
Lu further teaches dividing the locally computed data for the allreduce, disclosing the torus-based synchronization algorithm in which the synchronization algorithm divides the data in the device into equal parts according to the number of dimensions, different parts of the data then perform ring allreduce in each dimension simultaneously, and, after the data have completed parameter synchronization in one dimension, the same operation is performed in another dimension until every part of the data has completed parameter synchronization in all dimensions (Lu, page 4248, Section II-A (“Data Parallelism”), subsection 3) (“Torus-Based Synchronization Algorithm”), second paragraph in column 2).
Shen and Lu are analogous to the claimed invention as both are from the same field of endeavor of interconnecting computing nodes that exchange training results in a distributed machine learning system by means of a micro-electro-mechanical optical switch. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to organize the servers of Shen into the MEMS-interconnected groups of Lu and to synchronize the graphics processing units within each such group using the allreduce synchronization algorithm of Lu. The motivation to combine Shen and Lu is that Lu teaches that the hierarchical group and pod arrangement yields a scalable topology reaching k³ computing nodes and that ring allreduce decentralizes synchronization so that the bandwidth of each device is fully utilized without a parameter-server bottleneck (Lu, page 4248, Section II-A (“Data Parallelism”), subsection 1) (“Parameter Server-based Synchronization Algorithm”), and subsection 2) (“Ring-based Synchronization Algorithm, last paragraph in column 1 continuing to the first paragraph in column 2, and column 2, section III-A, paragraph 1), which addresses the problem Shen identifies of increasingly frequent data transmission between GPU chips as the accelerator cluster scale grows (Shen, ¶ [0005]).
The combination of Shen and Lu, however, does not teach "dividing, by the first processor, the intermediate data of the first processor based on the allreduce and a total quantity S×C of processors in the group, to obtain S×C pieces of data"; "obtaining, by the first processor through the intra-node channel of the first node and/or optical transmission channels that are between the first node and other (S−1) nodes in the group to which the first processor belongs and that are established by the MEMS, ith data of other (S×C−1) processors in the group to which the first processor belongs"; and "obtaining the first target data after the first processor performs summation on ith data in the S×C pieces of data and the ith data of the other (S×C−1) processors in the group to which the first processor belongs."
In the same field of endeavor, Cevolani teaches "dividing, by the first processor, the intermediate data of the first processor based on the allreduce and a total quantity S×C of processors in the group, to obtain S×C pieces of data." Cevolani discloses that each of a plurality of first processing nodes stores an array of data items and takes part in one or more reduce scatter collectives using that array to obtain a reduced subset, each reduce scatter collective being performed between the processing nodes of a different one of at least two different sets to which the node belongs (Cevolani, ¶ [0009]). Cevolani further discloses that by first performing the reduce-scatter between the processing nodes belonging to a set, each processing node in that set receives a different subset of the array of data (Cevolani, ¶ [0010]), and that where the respective sets comprise more than two sets of processing nodes the reduce scatter collectives comprise a plurality of reduce scatter collectives and the all-gather collectives comprise a plurality of all-gather collectives (Cevolani, ¶ [0025]). Cevolani further discloses that each stored element of the array comprises weight updates for a neural network (Cevolani, ¶ [0023]). The recited dividing based on the total quantity S×C of processors in the group reads on the successive apportionment of Cevolani across the members of the nested sets to which a processing node belongs.
Cevolani further teaches "obtaining, by the first processor through the intra-node channel of the first node and/or optical transmission channels that are between the first node and other (S−1) nodes in the group to which the first processor belongs and that are established by the MEMS, ith data of other (S×C−1) processors in the group to which the first processor belongs." Cevolani discloses that the reduce scatter collective is performed between processing nodes of a first of the respective two sets, that the all-reduce collective is thereafter performed between processing nodes of a second of the respective two sets, and that the all-gather collective is then performed between processing nodes of the first set (Cevolani, ¶ [0024]). Cevolani further discloses processing nodes in a rack, each of which stores a different set of the full gradients for a machine learning model and thereafter a subset of reduced gradients, and gateways in different racks between which an all-reduce is performed in different directions in a ring (Cevolani, ¶¶ [0064]–[0066], FIGS. 24, 25, and 26). The recited obtaining of the ith data of the other (S×C−1) processors through the intra-node channel and through the MEMS-established inter-node channels reads on the two-stage exchange of Cevolani, in which the correspondingly indexed portion is first gathered within the rack and thereafter across racks.
Cevolani further teaches "obtaining the first target data after the first processor performs summation on ith data in the S×C pieces of data and the ith data of the other (S×C−1) processors in the group to which the first processor belongs." Cevolani discloses that following the reduce-scatter collective of the all-reduce collective and prior to the all-gather collective, operations are performed on each stored element of the array of data items stored by the respective first processing node to modify the data of those stored elements (Cevolani, ¶ [0022]), and that those operations comprise providing updated weights of the neural network using the weight updates (Cevolani, ¶ [0023]). Cevolani further discloses that an all-reduce function is implemented by a reduce-scatter step followed by an all-gather step (Cevolani, ¶¶ [0053]–[0054], FIGS. 15 and 16A), and that each processing node in the rack thereafter stores a subset of reduced gradients and subsequently a subset of the updated weights for the machine learning model (Cevolani, ¶¶ [0065], [0067], FIGS. 25 and 27).
Shen, Lu, and Cevolani are analogous to the claimed invention as all three are from the same field of endeavor of exchanging and aggregating training results among a plurality of interconnected processors in a distributed machine learning system. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to structure the group-wide allreduce of Lu, as applied to the graphics processing units of Shen, as the nested reduce-scatter and all-reduce of Cevolani, in which the correspondingly indexed portion is gathered and summed first within a server over the intra-node channel and thereafter across the servers of the group over the MEMS-established optical channels. The motivation to combine Shen, Lu, and Cevolani is that Cevolani teaches that performing the reduce-scatter first within a set causes each node to hold only a subset of the array, so that the subsequent exchange involves a smaller amount of data and results in lower bandwidth utilization (Cevolani, ¶ [0010]), which addresses the problem Shen identifies of increasingly frequent data transmission between GPU chips as the accelerator cluster scale grows (Shen, ¶ [0005]).
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to HUNG VAN LE whose telephone number is (571)270-0164. The examiner can normally be reached 8 a.m. - 5 p.m..
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Cesar Paula can be reached at (571) 272-4128. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/HUNG VAN LE/Examiner, Art Unit 2145
/CESAR B PAULA/Supervisory Patent Examiner, Art Unit 2145