Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 1-3 and 5-18 are rejected under 35 U.S.C. 103 as being unpatentable over Sridharan et al. – (US 2019/0205745 A1 – hereinafter after Sridharan.) and further in view of Xu et al. – (“EdgeMesh: A Hybrid Distributed Training Mechanism for Heterogeneous Edge Devices” – hereinafter Xu).
In regards to claim 1, Sridharan disclose an apparatus comprising: a processing device configured to:
(Xu page 11 second paragraph teaches Nvidia Jetson TX2 (TX2), Nvidia Jetson Nana (Nano), and PC, all of which have processors.) (Sridharan para. [0381] teaches a processor wherein it cites “a processor to execute instructions provided by the library, the instructions to cause the processor to create one or more groups of the worker nodes,…”)
provide an initial artificial neural network (ANN) model to a plurality of groups of node devices; (Sridharan para. [0199] teaches splitting a model between different nodes wherein it cites “Model parallelism uses the same data for each compute node, with the model split among compute nodes.”, this analogous to the providing an ANN model to plurality of devices as the instant specification in para. [0032] cites “The central server 202 can provide the initial ANN model 205-2 by providing portions of the ANN model 205-2 to the groups 222.”, which means providing parts of the model to different groups, nodes or devices for processing.)
provide an input to a group of nodes from the plurality of groups of nodes; (Sridharan para. [0203-0204] teaches the input data is also split across nodes, and para. [0222-0223] teaches grouping worker nodes into groups.)
responsive to providing the input, receive activation signals from a first portion of the plurality of groups; (Sridharan para. [0201] teaches using input data to get activations and partial activations wherein it cites “In each of FIGS. 20A-20E, input data 2002 is processed by a machine learning model having a set of weights 2004 to generate a set of activations 2008 or partial activations 2006.”. Para. [0204] teaches Nodes 0-3 each generating its own respective partial activation. Table 1 teaches gather function which gathers data from multiple group processes into one process and reduce which the results across group processes is collected in one specified process. Also see para. [0205] which takes the partial activations form multiple nodes generated in layer N-1, reduces them, and then sends it’s the nodes in the next layer (Layer N) as activations.)
provide the activation signals to a second portion of the plurality of groups; (Sridharan see para. [0205] which takes the partial activations form multiple nodes generated in layer N-1, reduces them, and then sends it’s the nodes in the next layer (Layer N) as activations; and para. [0203] teaches summing activations in one layer to send to the next layer wherein it cites “Using the illustrated approach, a reduce operation is performed to sum up the activations to obtain the actual output and then scatter the activations for use in computing activations for the next layer.”)
provide commands to the plurality of groups nodes to train the initial ANN model (Sridharan para. [0215] discloses layer-to-layer communications and “intelligent message scheduling across the defined neural network layers,…”; para. [0241] teaches that a list of communication operations and associated nodes is specified in an advance and repeatedly executed as forward propagation, backpropagation, and gradient distribution are performed.; and para. [0381] cites “the instructions to cause the processor to create one or more groups of the worker nodes, the one or more groups of worker nodes to be created based on a communication pattern for messages to be transmitted between the worker nodes during distributed training of the neural network.”, thus the above operations are performed by way of commands/instructions.)
to generate a trained ANN model based on training feedback generated using different activation signals received from the second portion of the plurality of groups; and (Sridharan para. [0206] discloses successive forward operations propagating activation data through successive layers. The second portion receivers use the first activation data as input and perform the successive layer operations, thereby generating different, downstream activation data. Para. [0173] teaches that the network output is compared with a desired output using a loss function, this produces error values that are backpropagated to update the weights using Stochastic gradient descent (SGD), this is also taught in para. [0162]. Para. [0206] also teaches backpropagation, using SGD to update weights, followed by Allreduce operations that update each layer’s weights. Then para. [0184] teaches the repeated weight adjustments produce a trained neural network. Examiner’s note: the feedback is the backpropagated errors as the instant specification in para. [0024] cites “The central server 102 can utilize the output to generate corrections (e.g., training feedback) for the initial ANN model 105-2. The corrections can be used to modify the weights, biases, and/or activation functions of the initial ANN model 105-2.”.)
receive the trained ANN model from the plurality of groups of nodes. (Sridharan para. [0190] teaches combining the results from various nodes and transferring parameters or model updates to a server that maintains global parameters/updates. Also para. [0184] teaches that the repeated updates results in a trained neural network.)
However, Sridharan does not explicitly disclose a plurality of edge devices wherein the edge devices received the neural network.
Xu disclose a plurality of edge device wherein the edge devices receive the neural network. (Xu section 3.2 teaches a server connected to multi-edge devices wherein it cites “This article extends the original MTF framework and proposes a parameter server framework for multi-edge device model parallelism. Unlike the previous hybrid multiprocessor architecture, this framework allows multiple devices to communicate to each other and provides mesh computing between nodes.” It also teaches wherein each node cluster maintains a convolution layer the model wherein it cites “Each convolution layer of the model is maintained by an edge node cluster that implements model parallelism,…”. Page 3.4 page 10 first paragraph each device obtains a convolution filter partition block.)
It would have been obvious to one of ordinary skill in the art before the earliest effective filing date of the claimed invention to modify the teachings of the Sridharan with that of Xu in order to allow for using sever-edge device architecture as both reference deal with distributed machine learning. The benefit of using edge devices in distributed learning is allow for more efficient processing by processing he data on individual devices wherein the data is collected and the sharing the important features, values and update with the server to create global model that is shared with other devices.
In regards to claim 2, Sridharan in view of Xu disclose the apparatus of claim 1, wherein the processing device is further configured to generate the training feedback by performing a loss calculation using the activation signals and the different activation signals. (Sridharan para. [0206] discloses successive forward operations propagating activation data through successive layers. The second portion receivers use the first activation data as input and perform the successive layer operations, thereby generating different, downstream activation data. Para. [0173] teaches that the network output is compared with a desired output using a loss function, this produces error values that are backpropagated to update the weights using Stochastic gradient descent (SGD), this is also taught in para. [0162]. Para. [0206] also teaches backpropagation, using SGD to update weights, followed by Allreduce operations that update each layer’s weights.)
In regards to claim 3, Sridharan in view of Xu discloses the apparatus of claim 1, wherein the processing device is further configured to provide the initial ANN model to the plurality of groups by providing a different portion of the initial ANN model to each of the plurality of groups of edge devices. (Sridharan para. [0199] teaches splitting a model between different nodes wherein it cites “Model parallelism uses the same data for each compute node, with the model split among compute nodes.”, this analogous to the providing an ANN model to plurality of devices as the instant specification in para. [0032] cites “The central server 202 can provide the initial ANN model 205-2 by providing portions of the ANN model 205-2 to the groups 222.”, which means providing parts of the model to different groups, nodes or devices for processing.)
In regards to claim 5, Sridharan in view of Xu discloses the apparatus of claim 1, wherein the processing device is further configured to provide the initial ANN model to the plurality of groups by providing a different layer of the initial ANN model to each of the plurality of groups of edge devices. (Sridharan para. [0189] teaches assigning different layers of a single neural network to different processing nodes wherein it cites, “different computational nodes in a distributed system can perform training computations for different parts of a single network. For example, each layer of a neural network can be trained by a different processing node of the distributed system.”. Sridharan para. [0199] further teaches that, in model parallelism, the model is split among the compute nodes. Sridharan paras. [0222–0223] teach organizing the worker nodes into multiple groups and using parameter servers to provide inter-group communication. Xu teaches implementing a neural-network layer using a group or cluster of edge devices rather than an individual processing node in Xu section 3.2, pages 5–6, wherein it cites “Each convolution layer of the model is maintained by an edge node cluster that implements model parallelism, and multiple clusters replace the single node of the original parameter server.” Xu further teaches that, within a processor cluster, multiple edge devices “jointly undertake the convolutional layer filters to perform the model parallelism.” Thus, Xu’s edge-device cluster collectively performs the function assigned to an individual layer-processing node in Sridharan.)
In regards to claim 6, Sridharan in view of Xu disclose the apparatus of claim 1, wherein the processing device configured to receive the trained ANN model is further configured to receive a plurality of same layers of the trained ANN model from each of the plurality of groups of edge devices. (Sridharan para. [0190] teaches receiving data from devices and averaging it, wherein the same layers are averaged together data wherein it cites “In data parallelism 1904 , the different nodes of the distributed network have a complete instance of the model and each node receives a different portion of the data . The results from the different nodes are then combined . While different approaches to data parallelism are possible , data parallel training approaches all require a technique of combining results and synchronizing the model parameters between each node . Exemplary approaches to combining data include parameter averaging and update based data parallelism . Parameter averaging trains each node on a subset of the training data and sets the global parameters (e.g., weights, biases) to the average of the parameters from each node . Parameter averaging uses a central parameter server that maintains the parameter data.” Sridharan para. [0202] similarly teaches that the same model is replicated across multiple nodes, that the gradients are averaged, and that an Allgather operation synchronizes the weights across the nodes. Then Xu section 3.2, pages 5–6, teaches implementing combined data parallelism and model parallelism using multiple groups or clusters of edge devices and an edge-gateway parameter server. Therefore, when Sridharan’s corresponding layer-processing nodes are implemented using Xu’s grouped edge devices, the parameter server receives the trained parameters or updates from multiple instances of the same layer from the respective groups of edge devices. )
In regards to claim 7, Sridharan in view of Xu disclose the apparatus of claim 6, wherein the processing device is further configured to perform weight aggregation on each of the plurality of same layers received from each of the plurality of groups to generate a plurality of layers that comprise the trained ANN model. (Sridharan para. [0190] teaches parameter averaging wherein the parameters of corresponding model instances are combined to generate global model parameters wherein it cites, “Parameter averaging trains each node on a subset of the training data and sets the global parameters (e.g., weights, biases) to the average of the parameters from each node.” Sridharan further teaches that a “central parameter server” maintains the aggregated parameter data. Sridharan para. [0202] teaches averaging the gradients generated by the replicated model instances and using an Allreduce operation “to update the weights of each layer for the next forward pass.” and further teaches using an Allgather operation after stochastic gradient descent “to synchronize weights across nodes.” Sridharan para. [0206] teaches performing separate Allreduce operations for Layer N and the other layers of the ANN “to update the weights of each layer for the next forward pass.” Thus, corresponding weights from multiple instances of the same layer are aggregated to generate one updated layer, and the separately aggregated layers collectively comprise the trained ANN model.)
In regards to claim 8, Sridharan in view of Xu disclose the apparatus of claim 6, wherein the processing device is further configured to: receive a quantity of weights from each edge device in each of the plurality of groups of edge devices; and aggregate the quantity of weights received from edge devices in each of the plurality of groups to generate a layer from the plurality of layers, for each of the plurality of groups, of the trained ANN model. (Sridharan para. [0190] teaches receiving model parameters from each participating node at a central parameter server and aggregating those parameters wherein it cites, “Parameter averaging trains each node on a subset of the training data and sets the global parameters (e.g., weights, biases) to the average of the parameters from each node.” Para. [0190] further teaches that parameter averaging uses “a central parameter server that maintains the parameter data” and alternatively teaches transferring model updates from the nodes to the parameter server. Sridharan para. [0202] teaches that each replicated model instance computes gradients with respect to the model parameters and that the gradients are averaged before stochastic gradient descent is performed. Sridharan para. [0206] teaches performing an individual Allreduce operation for each layer to update that layer’s weights. Therefore, for each group of nodes assigned to a corresponding layer, the server receives the weights or weight updates from the individual nodes and aggregates those weights to generate the trained layer corresponding to that group.)
In regards to claim 9, Sridharan in view of Xu discloses the apparatus of claim 1, wherein the processing device configured to provide the activation signal to the second portion of the plurality of groups is further configured to provide a different instance of the activation signal to each edge device in the second portion of the plurality of groups. (Sridharan para. [0200], Table 1, teaches an Allgather operation wherein “all processes receive the gather result” and an Allreduce operation wherein the result of the reduce operation “is broadcast to all processes within a group.” Thus, a separate instance of the same aggregated activation result is provided to each receiving process in the group. Sridharan para. [0205] further teaches reducing the partial activations generated by multiple Layer N−1 nodes and scattering the resulting activations to multiple Layer N nodes. Sridharan para. [0206] teaches using an Alltoall operation to transfer activation data “to all available receivers,” which use the received activation data as input for operations on successive layers. Xu section 3.2, page 6, teaches that input tensors are replicated across the processors of an edge-device cluster and that the devices jointly process the input or feature map using their corresponding convolutional filters. This teaches each edge device in the receiving group receives its own instance of the activation signal for processing. Examiner’s note: Instant specification in para. [0063] teaches that providing a different instance can include providing the same activation-signal matrix to each device in the group, indicating each device get is own copy of the instance to be a different instance.)
In regards to claim 10, Sridharan in view of Xu discloses the apparatus of claim 1, wherein the processing device is further configured to generate a single set of the activation signals prior to providing the activation signals to the second portion. (Sridharan para. [0203] teaches generating a single activation result from multiple partial activations before distributing that result to the next-layer nodes wherein it cites, “a reduce operation is performed to sum up the activations to obtain the actual output and then scatter the activations for use in computing activations for the next layer.”.)
In regards to claim 11, Sridharan in view of Xu discloses the apparatus of claim 10, wherein the activation signals comprise multiple sets, each set corresponding to a different memory device from the first portion of the plurality of groups of memory devices. (Sridharan para. [0204] teaches multiple sets of partial activation signals, each generated by a different processing node. Specifically, Node 0 generates first partial activation 2006A, Node 1 generates second partial activation 2006B, Node 2 generates third partial activation 2006C, and Node 3 generates fourth partial activation 2006D. Each partial activation set therefore corresponds to a particular node that generated and stored that activation data. Sridharan paras. [0051–0052] teach that the processing systems include corresponding memory devices used to store data and instructions. Sridharan para. [0205] then teaches reducing the multiple node-specific partial activation sets to generate the activation signals provided to the next layer.)
In regards to claim 12, Sridharan discloses an apparatus comprising: a processing device configured to: (Sridharan para. [0188] teaches “The distributed computational nodes can each include one or more host processors and one or more general-purpose processing nodes.”.)
receive a portion of an initial artificial neural network (ANN) model from a central server, wherein the apparatus is part of a first group of node devices that receive the portion of the initial ANN model; (Sridharan para. [0199] teaches model parallelism wherein the model is split among multiple compute nodes. Para. [0204] teaches that each node receives a corresponding block of weight data representing a portion of the model. Para. [0190] teaches a central parameter server that maintains the model parameters. Paras. [0202] and [0191] teach combining data and model parallelism, wherein the model is replicated across nodes while corresponding model portions are processed by separate processing devices. Paras. [0222–0223] teach organizing worker nodes into groups and using parameter servers to provide inter-group communication. Thus, the corresponding model portion can be provided from the central parameter server to the nodes assigned to that portion.)
receive first activation signals from the central server, wherein the first activation signals are generated by a second group of node devices; (Sridharan para. [0205] teaches that nodes operating Layer N−1 generate partial activation signals and that the reduced activation signals are provided to a different set of nodes operating Layer N. Para. [0200], Table 1, teaches a REDUCE operation wherein the results generated by multiple processes in a group are collected in one specified process and a REDUCE_SCATTER operation wherein the reduced result is distributed to other processes. Para. [0223] teaches using a parameter server to bridge different groups of worker nodes for inter-group communication. Thus, activation signals generated by the Layer N−1 group are collected through the parameter-server communication framework and provided to the apparatus operating as part of the Layer N group.)
process the first activation signals utilizing the portion of the initial ANN model to generate second activation signals; (Sridharan para. [0205] teaches that the Layer N nodes receive the activation signals generated by the Layer N−1 nodes and process those activation signals using their corresponding model weights. Para. [0206] teaches that the receivers “use the activation data as input data for operations on successive layers,” thereby generating different activation data for the successive layer.)
provide the second activation signals to the central server for providing to a third group of node devices; (Sridharan para. [0203] teaches reducing the partial activations generated for a layer to obtain the output and then scattering that output for use in calculating the activations of the next layer. Para. [0206] teaches successive forward operations wherein activation data is transferred from the nodes that generate the activation data to receivers operating subsequent layers. Para. [0200] teaches collecting the activation data in a specified process before distributing the result, while para. [0223] teaches using a parameter server to provide inter-group communication. Thus, the second activation signals generated by the first group are collected through the parameter-server communication framework and provided to a third group operating the next layer.)
receive feedback from the central server, wherein the feedback is at least partially based on the second activation signals; (Sridharan para. [0173] teaches that the network output is compared with the desired output using a loss function, which produces error values that are backpropagated to update the model weights using stochastic gradient descent. Para. [0206] applies backpropagation to the distributed nodes and teaches using AllReduce, AllGather, and AlltoAll operations to communicate gradients and updated weight data among the processes. Para. [0200] teaches that an Allreduce operation broadcasts the reduced result to all processes in the communication group.)
update weights of the portion of the initial ANN model utilizing the feedback to generate an updated portion of the ANN model; and (Sridharan para. [0173] teaches using the backpropagated error values and stochastic gradient descent to update the model weights.)
provide the updated portion of the ANN model to the central server. (Sridharan para. [0190] teaches transferring the model parameters or model updates from the distributed nodes to the central parameter server.)
However, Sridharan does not explicitly disclose that edge devices organized into groups of edge devices.
Xu discloses groups of edge devices connected with a central server. (Xu section 3.2, pages 5–6, teaches “a parameter server framework for multi-edge device model parallelism” wherein the framework allows multiple edge devices to communicate with each other. Xu section 3.2 page 5-6 furhter teaches that each convolutional layer is maintained by an edge-node cluster, that the edge devices jointly perform the layer calculations, and that the edge-gateway node deploys a parameter server that “internally maintains a global parameter to share model parameters.”)
It would have been obvious to one of ordinary skill in the art before the earliest effective filing date of the claimed invention to modify the teachings of the Sridharan with that of Xu in order to allow for using sever-edge device architecture as both reference deal with distributed machine learning. The benefit of using edge devices in distributed learning is allow for more efficient processing by processing the data on individual devices wherein the data is collected and the sharing the important features, values and update with the server to create global model that is shared with other devices.
In regards to claim 13, Sridharan in view of Xu disclose the apparatus of claim 12, wherein the processing device is further configured to process the first activation signals utilizing the portion of the initial ANN model to generate the second activation signals, wherein each edge device in first group of edge devices processes the first activation signals concurrently. (Sridharan para. [0203] teaches model-parallel processing wherein different portions of a model’s computation are performed simultaneously on different nodes using the same batch of input data wherein it cites, “model parallelism performs different portions of a model’s computation … simultaneous on different nodes for the same batch of examples.” Para. [0204] further teaches that Nodes 0–3 process their respective input-data and weight blocks to generate corresponding partial activations 2006A–2006D. Para. [0206] teaches that the receiving nodes use the received activation data as input for operations on successive layers. Xu section 3.2, pages 5–6, teaches that the input tensors are replicated across the processors of an edge-device cluster, that the devices in the cluster “jointly undertake the convolutional layer filters,” and that “the hidden layer is computed by nodes in parallel.” Xu section 3.4, page 10, similarly teaches that each edge device processes the input data or feature map using its assigned convolution-filter partition and waits for the other nodes to complete their calculations before the results are combined using Allreduce. Thus, each edge device in the first group concurrently processes its instance of the first activation signals using its assigned ANN-model portion to generate corresponding second activation signals.)
In regards to claim 14, Sridharan in view of Xu disclose the apparatus of claim 13, wherein the processing device is configured to receive the first activation signals concurrently with a receipt of the first activation signals by the edge devices of the first group. (Sridharan para. [0200], Table 1, teaches an Allgather operation wherein “all processes receive the gather result” and an Allreduce operation wherein the reduced result “is broadcast to all processes within a group.” Para. [0206] teaches using an Alltoall operation to transfer activation data “to all available receivers,” which use the activation data as input for successive-layer operations. Thus, this teaches that the collective communication operation provides the activation signals to the receiving processes of the group concurrently. Xu section 3.2, page 6, teaches that the input tensors are replicated across the processors of an edge-device cluster and that the processors jointly and in parallel process the input tensor or feature map. Thus this also teacehs the apparatus receives its instance of the first activation signals concurrently with the other edge devices of the first group.)
In regards to claim 15, Sridharan in view of Xu discloses the apparatus of claim 12, wherein the processing device configured to provide the updated portion of the ANN model is further configured to provide the updated portion of the ANN model concurrently with a providing of different updated portion of the ANN model by the edge devices of the first group. (Sridharan para. [0203] teaches that different portions of the model’s computations are performed simultaneously by multiple nodes. Para. [0206] teaches that, during backpropagation, distributed stochastic gradient descent generates updated weight data and that Allreduce operations are performed for the respective layers to communicate and synchronize the updated weights among all processes in the communication group. Para. [0200], Table 1, explains that an Allreduce operation performs a reduction across all participating processes and broadcasts the result to those processes. Sridharan para. [0190] further teaches transferring model parameters or model updates from the distributed nodes to the central parameter server. Xu section 3.2, pages 5–6, teaches that the devices within an edge cluster jointly train their corresponding model portions in parallel and that the gateway parameter server maintains the global parameters. Xu section 3.4, page 10, teaches that each edge device maintains a different convolution-filter partition and participates in the collective Allreduce operation after completing its corresponding calculation. This means that each device or node provides the updated weights for its assigned model portion during the same collective training operation in which the other edge devices provide the different updated portions assigned to those devices.)
In regards to claim 16, Sridharan in view of Xu discloses the apparatus of claim 12, wherein the feedback is based, at least partially, on the first activation signals and the second activation signals. (Sridharan para. [0204] teaches that input data or activation block and a corresponding weight block are processed to generate a partial output activation. Para. [0206] teaches successively propagating the resulting activation data through the network and then performing backpropagation and distributed stochastic gradient descent to generate updated weight data. Xu section 3.1, page 4, Equations 3–5, connects the input and output activations to the weight gradient. Xi is the input processed by a model portion and Yi is its output activation, calculates the output error term ΔY, and calculates the weight gradient using ΔYXi. Thus, the gradient constituting the training feedback is calculated using both the first activation signals supplied as the input Xi and the second activation signals represented by the output Yi and its corresponding error term ΔY).
In regards to claim 17, Sridharan in view of Xu discloses the apparatus of claim 12, wherein the feedback is based on an output that is generated utilizing the second activation signals. (Sridharan para. [0206] teaches that the activation data generated by one layer is successively propagated through the subsequent layers to generate the neural-network output. Sridharan para. [0173] teaches that the network output is compared with a desired output using a loss function, thereby generating error values that are backpropagated to update the model weights.)
In regards to claim 18, Sridharan in view of Xu discloses the apparatus of claim 17, wherein the output is generated by a different group of edge devices. (Sridharan para. [0189] teaches that different layers of a neural network can be trained and executed by different processing nodes. Paras. [0205–0206] teach transferring activation signals from a first set of nodes operating one layer to a different set of nodes operating a successive layer and continuing this process through Layer N. The nodes operating the final layer therefore generate the ANN output from activation signals generated by the preceding layer-processing nodes.)
Claims 4 and 19-23 are rejected under 35 U.S.C. 103 as being unpatentable over Sridharan et al. – (US 2019/0205745 A1 – hereinafter after Sridharan.) in view of Xu et al. – (“EdgeMesh: A Hybrid Distributed Training Mechanism for Heterogeneous Edge Devices” – hereinafter Xu) and further in view of Huang et al. (“GPipe: Easy Scaling with Micro-Batch Pipeline Parallelism” – hereinafter Huang).
In regards to claim 4, Sridharan in view of Xu discloses the apparatus of claim 3, but does not explicitly disclose wherein each of the different portions of the initial ANN model comprises a plurality of layers of the initial ANN model.
Huang discloses wherein each of the different portions of the initial ANN model comprises a plurality of layers of the initial ANN model. (Huang page 2 teaches wherein a model is made up of a sequence of layers, the consecutive layers are grouped and assigned to different cells, wherein it cites “With GPipe, each model can be specified as a sequence of layers, and consecutive groups of layers can be partitioned into cells. Each cell is then placed on a separate accelerator.”)
It would have been obvious to one of ordinary skill in the art before the earliest effective filing date of the claimed invention to modify the teachings of Sridharan in view of Xu with the teachings of Huang in order to allow for partitioning a model into groups of layers and then assigning each group to a device or cell as all the references deal with distributed processing. The benefit of doing so is it allows low idle time by scheduling backpropagation as cited in Huang section 2.3 and it also allows for faster or speedup processing with partitioning data across more accelerators as cited on page 5 third paragraph of Huang.
In regards to claim 19, Sridharan discloses a method comprising:
providing an initial artificial neural network (ANN) model to a plurality of groups of node devices; (Sridharan para. [0199] teaches splitting a model between different nodes wherein it cites “Model parallelism uses the same data for each compute node, with the model split among compute nodes.”, this analogous to the providing an ANN model to plurality of devices as the instant specification in para. [0032] cites “The central server 202 can provide the initial ANN model 205-2 by providing portions of the ANN model 205-2 to the groups 222.”, which means providing parts of the model to different groups, nodes or devices for processing.)
providing a first input to a group of node devices from the plurality of groups of node devices; (Sridharan para. [0203-0204] teaches the input data is also split across nodes, and para. [0222-0223] teaches grouping worker nodes into groups.)
receiving activation signals from the group, wherein the activation signals are generated by the group utilizing the initial ANN model; (Sridharan para. [0201] teaches using input data to get activations and partial activations wherein it cites “In each of FIGS. 20A-20E, input data 2002 is processed by a machine learning model having a set of weights 2004 to generate a set of activations 2008 or partial activations 2006.”. Para. [0204] teaches Nodes 0-3 each generating its own respective partial activation. Table 1 teaches gather function which gathers data from multiple group processes into one process and reduce which the results across group processes is collected in one specified process. Also see para. [0205] which takes the partial activations form multiple nodes generated in layer N-1, reduces them, and then sends it’s the nodes in the next layer (Layer N) as activations.)
providing a second input to the group of node devices; (Sridharan para. [0184] teaches repeatedly processing training inputs and adjusting the weights until the neural network reaches the desired accuracy. Para. [0202] teaches dividing the training data along a mini-batch dimension and independently performing forward propagation for the different samples. Para. [0241] teaches repeatedly executing the predetermined communication and training operations as forward propagation and backpropagation calculations are performed. Thus, after processing a first training input or mini-batch, the group receives and processes a subsequent second training input or mini-batch.)
providing the activation signals to a portion of the plurality of groups,; (Sridharan para. [0205] teaches reducing the partial activations generated by Layer N−1 nodes and providing the resulting activations to nodes operating Layer N. Para. [0206] teaches transferring activation data from generating nodes to receivers that use the activation data as input for operations on successive layers. Para. [0203] teaches that different portions of the model’s computation are performed simultaneously by multiple nodes, and para. [0191] teaches combining model parallelism and data parallelism.)
receiving different activation signals from the portion of the plurality of groups; (Sridharan para. [0206] teaches that the successive-layer receivers process the received activation data to perform the successive-layer operations, thereby generating different downstream activation data. Paras. [0200] and [0203] teach collecting or reducing the activation results generated by the receiving nodes.)
generating training feedback utilizing the different activation signals; (Sridharan para. [0206] discloses successive forward operations propagating activation data through successive layers. The second portion receivers use the first activation data as input and perform the successive layer operations, thereby generating different, downstream activation data. Para. [0173] teaches that the network output is compared with a desired output using a loss function, this produces error values that are backpropagated to update the weights using Stochastic gradient descent (SGD), this is also taught in para. [0162]. Para. [0206] also teaches backpropagation, using SGD to update weights, followed by Allreduce operations that update each layer’s weights. Then para. [0184] teaches the repeated weight adjustments produce a trained neural network. Examiner’s note: the feedback is the backpropagated errors as the instant specification in para. [0024] cites “The central server 102 can utilize the output to generate corrections (e.g., training feedback) for the initial ANN model 105-2. The corrections can be used to modify the weights, biases, and/or activation functions of the initial ANN model 105-2.”.)
providing commands to the plurality of groups of node devices to update the initial ANN model (Sridharan para. [0215] discloses layer-to-layer communications and “intelligent message scheduling across the defined neural network layers,…”; para. [0241] teaches that a list of communication operations and associated nodes is specified in an advance and repeatedly executed as forward propagation, backpropagation, and gradient distribution are performed.; and para. [0381] cites “the instructions to cause the processor to create one or more groups of the worker nodes, the one or more groups of worker nodes to be created based on a communication pattern for messages to be transmitted between the worker nodes during distributed training of the neural network.”, thus the above operations are performed by way of commands/instructions.) to generate a trained ANN model utilizing the training feedback; and (Sridharan para. [0206] discloses successive forward operations propagating activation data through successive layers. The second portion receivers use the first activation data as input and perform the successive layer operations, thereby generating different, downstream activation data. Para. [0173] teaches that the network output is compared with a desired output using a loss function, this produces error values that are backpropagated to update the weights using Stochastic gradient descent (SGD), this is also taught in para. [0162]. Para. [0206] also teaches backpropagation, using SGD to update weights, followed by Allreduce operations that update each layer’s weights. Then para. [0184] teaches the repeated weight adjustments produce a trained neural network. Examiner’s note: the feedback is the backpropagated errors as the instant specification in para. [0024] cites “The central server 102 can utilize the output to generate corrections (e.g., training feedback) for the initial ANN model 105-2. The corrections can be used to modify the weights, biases, and/or activation functions of the initial ANN model 105-2.”.)
receiving the trained ANN model from the plurality of groups of node devices. (Sridharan para. [0190] teaches combining the results from various nodes and transferring parameters or model updates to a server that maintains global parameters/updates. Also para. [0184] teaches that the repeated updates results in a trained neural network.)
However, Sridharan does not explicitly disclose a plurality of edge devices wherein the edge devices received the neural network, nor does it teach wherein the group processes the second input and the portion of the plurality of groups processes the activation signals concurrently utilizing the initial ANN model.
Xu disclose a plurality of edge device wherein the edge devices receive the neural network. (Xu section 3.2 teaches a server connected to multi-edge devices wherein it cites “This article extends the original MTF framework and proposes a parameter server framework for multi-edge device model parallelism. Unlike the previous hybrid multiprocessor architecture, this framework allows multiple devices to communicate to each other and provides mesh computing between nodes.” It also teaches wherein each node cluster maintains a convolution layer the model wherein it cites “Each convolution layer of the model is maintained by an edge node cluster that implements model parallelism,…”. Page 3.4 page 10 first paragraph each device obtains a convolution filter partition block.)
It would have been obvious to one of ordinary skill in the art before the earliest effective filing date of the claimed invention to modify the teachings of the Sridharan with that of Xu in order to allow for using sever-edge device architecture as both reference deal with distributed machine learning. The benefit of using edge devices in distributed learning is allow for more efficient processing by processing he data on individual devices wherein the data is collected and the sharing the important features, values and update with the server to create global model that is shared with other devices.
However Sridharan in view of Xu does not explicitly disclose wherein the group processes the second input and the portion of the plurality of groups processes the activation signals concurrently utilizing the initial ANN model.
Huang discloses wherein the group processes the second input and the portion of the plurality of groups processes the activation signals concurrently utilizing the initial ANN. (Huang page 2 third paragraph teaches GPIP divides the neural network into cells, places each cell on a separate accelerator and divides a mini-batch into micro-batch, and sends micro batches to the cells wherein it cites “With GPipe, each model can be specified as a sequence of layers, and consecutive groups of layers can be partitioned into cells. Each cell is then placed on a separate accelerator. Based on this partitioned setup, we propose a novel pipeline parallelism algorithm with batch splitting. We first split a mini-batch of training examples into smaller micro-batches, then pipeline the execution of each set of micro-batches over cells.”. Also see figure 2C wherein F1 is working it first micro batch of data (input) while F0,1 is working on activation functions.)
It would have been further obvious to one of ordinary skill in the art before the earliest effective filing date of the claimed invention to modify the teachings of the Sridharan in view of Xu with of the Huang in order to allow for processing two different dataset concurrently as all the references deal with distributed processing. The benefit of doing so is it allows low idle time by scheduling backpropagation as cited in Huang section 2.3 and it also allows for faster or speedup processing with partitioning data across more accelerators as cited on page 5 third paragraph of Huang.
In regards to claim 20, Sridharan in view of Xu in view of Huang discloses the method of claim 19, further comprising generating the training feedback utilizing the different propagation signals and the activation signals. (Sridharan para. [0206] discloses successive forward operations propagating activation data through successive layers. The second portion receivers use the first activation data as input and perform the successive layer operations, thereby generating different, downstream activation data. Para. [0173] teaches that the network output is compared with a desired output using a loss function, this produces error values that are backpropagated to update the weights using Stochastic gradient descent (SGD), this is also taught in para. [0162]. Para. [0206] also teaches backpropagation, using SGD to update weights, followed by Allreduce operations that update each layer’s weights.)
In regards to claim 21, Sridharan in view of Xu in view of Huang discloses the method of claim 19, wherein the second input and the different activation signals are processed by the plurality of groups concurrently. (Huang page 2 third paragraph teaches GPIP divides the neural network into cells, places each cell on a separate accelerator and divides a mini-batch into micro-batch, and sends micro batches to the cells wherein it cites “With GPipe, each model can be specified as a sequence of layers, and consecutive groups of layers can be partitioned into cells. Each cell is then placed on a separate accelerator. Based on this partition setup, we propose a novel pipeline parallelism algorithm with batch splitting. We first split a mini-batch of training examples into smaller micro-batches, then pipeline the execution of each set of micro-batches over cells.”. Also see figure 2C wherein F1 is working it first micro batch of data (input) while F0,1 is working on activation functions.)
In regards to claim 22, Sridharan in view of Xu in view of Huang discloses the method of claim 19, further comprising dividing a plurality of edge devices into the plurality of groups based on computing capabilities of the plurality of edge devices. (Xu section 3.3.1, pages 6–7, teaches determining the floating-point operations per second (FLOPS) of each participating edge device and using those FLOPS measurements to determine an optimal partition of the edge-device mesh. Xu explains that its device-partitioning method obtains the respective FLOPS values of the heterogeneous edge devices and calculates a mesh layout that balances the computing resources among the clusters. Xu Algorithm 1, pages 7–8, receives the FLOPS measurements of the participating devices, executes the cluster partition operation, and broadcasts the resulting cluster assignment to the devices.)
In regards to claim 23, Sridharan in view of Xu in view of Huang discloses the method of claim 19, wherein providing the activation signals further comprises providing matrices representing the activation signals, wherein the matrices are fixed. (Sridharan para. [0204] teaches partitioning input data, weight data, and activation data across multiple nodes “to minimize skewed matrices” and teaches that each node generates a corresponding partial-activation block. Paras. [0203] and [0205] teach reducing those activation blocks and providing the resulting activation data to the nodes operating the next layer. Para. [0241] teaches that the communication pattern and associated nodes are specified in advance and repeatedly reused throughout forward and backward propagation using a persistent communication graph. Xu section 3.1, pages 4-5, represents the neural-network inputs and output activations as (Xi) and (Yi) matrices or tensors and defines the output dimensions using Equations 6-7 based on the input dimensions, convolution-filter dimensions, padding, and stride. Xu section 3.2, page 6, teaches combining the feature map fragments into a complete feature map matrix before the next layer calculation. For a selected ANN architecture and input size, these equations establish predetermined activation-matrix dimensions that do not change during the training operation.)
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to PAULINHO E SMITH whose telephone number is (571)270-1358. The examiner can normally be reached Mon-Fri. 10AM-6PM CST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Abdullah Kawsar can be reached at 571-270-3169. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/PAULINHO E SMITH/Primary Examiner, Art Unit 2127