DETAILED ACTION
This action is in response to claims filed 30 January 2024 for application 18495528 filed 26 October 2023. Currently claims 1-21 are pending.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 1-20 are rejected under 35 U.S.C. 103 as being unpatentable over Chen et al. (US 20210042620 A1) in view of Kale (US 20220044101 A1).
Regarding claims 1, 16 and 19, Chen discloses: A first device for a communication network, comprising at least two second devices, each of the at least two second devices including a neural network (NN) the first device comprising:
A first device NN (Fig 1), and
Processing circuitry configured to:
Obtain outputs of the NNs of the at least two second devices (“The described system allows for effectively distributing the training of giant neural networks, i.e., neural networks with a very large number of parameters, across multiple computing devices. By partitioning the neural network into a plurality of composite layers, the described system can scale arbitrary giant neural network architectures beyond the memory limitations of a single computing device.” [0013]);
Combine the obtained outputs to generate a combined output (“The K composite layers 103 form the sequence defined by the giant neural network 102, starting from a first composite layer P.sub.1 that includes the input layer i=1 for the neural network, and ending with a last composite layer P.sub.K that includes the output layer i=L for the neural network. In this specification, the succeeding and preceding composite layers relative to a particular composite layer in the sequence are sometimes called neighboring composite layers for the particular composite layer.” [0046], see also [0047]);
Provide the combined output as an input to the first device NN [0046], [0047]; and
Run the first device NN based on the input [0046], [0047].
However, Chen does not explicitly disclose: wherein parameters of the first device NN and the NNs of the at least two second devices are independent of each other.
Kale teaches: wherein parameters of the first device NN and the NNs of the at least two second devices are independent of each other (Fig 10, “For example, each integrated circuit device has a Deep Learning Accelerator and random access memory storing a set of instructions and resources to at least implement the computation of an Artificial Neural Network processing the input from a particular sensor. The output of the Artificial Neural Network is independent on the input from other sensors.” [0038], “In one embodiment of sensor fusion, first data representative of parameters of an artificial neural network (201) is stored into random access memory (105) of a device (e.g., 101). For example, the parameters can include kernel and maps matrices (207) of the artificial neural network (201) trained using a machine learning and/or deep learning technique.” [0120], note: each circuit has a NN and associated RAM storing parameters that are independent of other devices).
Chen and Kale are in the same field of endeavor of neural networks and are analogous. Chen discloses combining circuits of neural networks. Kale teaches combining circuits of neural networks with independent neural networks on independent inputs. It would have been obvious to one of ordinary skill in the art before the effective filing date to modify the structure of Chen to have independent NNs as taught by Kale to yield predictable results of independent analysis of different inputs.
Regarding claim 2, Chen discloses: The first device according to claim 1, wherein a size of an input layer of the first device NN is equal to a sum of sizes of output layers of the NNs of the at least two second devices (“Next, the system 100 assigns each of the plurality of composite layers 103 to a respective device of the computing devices 106 for computing the forward computation functions and backpropagation gradients within each composite layer. In some implementations, the system 100 can partition the neural network based on the number of devices available, i.e., so that each device is assigned one composite layer. For example, the computing devices 106 have K computing devices available for processing the composite layers and the system 100 can therefore divide the giant neural network 102 into K composite layers 103.” [0047], “The batch splitting engine 108 of the system 100 takes as input the training examples 101 including a plurality of mini-batches 105, and splits each mini-batch into a plurality of micro-batches 107 of equal size, as shown in FIG. 1.” [0055]).
Regarding claim 3, Chen discloses: The first device according to claim 2, wherein the processing circuitry is configured to combine the obtained outputs, by concatenating activation vectors of the output layers of the NNs of the at least two second devices (“The system can integrate recomputation (or rematerialization) of activation functions during backward propagations with pipeline parallelism for more efficient memory usage and hardware utilization of each individual computing device.” [0025]).
Regarding claim 4, Chen discloses: The first device according to claim 2, the processing circuitry being further configured to:
perform backpropagation to obtain an error vector of the input layer of the first device NN (“A backpropagation function b.sub.i is any function that computes a gradient of the error between an output activation at a network layer i and the expected output of the network at the network layer i. The backpropagation function b.sub.i can use the output gradient computed by a previous layer i+1 to compute the error and obtain the output gradient for the network layer i. Each network layer has a plurality of network parameters that are adjusted as part of neural network training, i.e., weight matrices and bias vectors. The network parameters are utilized within each network layer during training when performing operations such as tensor productions, convolution or attention mechanism.” [0042]);
split the obtained error vector based on the sizes of the output layers of the NNs of the at least two second devices to obtain a set of split error vectors [0042]; and
provide the set of split error vectors to the output layers of the NNs of in the at least two second devices, wherein each NN of the NNs of the at least two second devices is configured to perform backpropagation based on a respective received split error vectors [0042].
Regarding claim 5, Chen discloses: The first device according to claim 1, wherein a size of an input layer of the first device NN is equal to a size of each output layer of the NNs of the at least two second devices (“Next, the system 100 assigns each of the plurality of composite layers 103 to a respective device of the computing devices 106 for computing the forward computation functions and backpropagation gradients within each composite layer. In some implementations, the system 100 can partition the neural network based on the number of devices available, i.e., so that each device is assigned one composite layer. For example, the computing devices 106 have K computing devices available for processing the composite layers and the system 100 can therefore divide the giant neural network 102 into K composite layers 103.” [0047], “The batch splitting engine 108 of the system 100 takes as input the training examples 101 including a plurality of mini-batches 105, and splits each mini-batch into a plurality of micro-batches 107 of equal size, as shown in FIG. 1.” [0055]).
Regarding claim 6, Chen discloses: The first device according to claim 5, wherein the processing circuitry is configured to combine the obtained outputs by performing element-wise summation of activation vectors of the output layers of the NNs of the at least two second devices (“The system can integrate recomputation (or rematerialization) of activation functions during backward propagations with pipeline parallelism for more efficient memory usage and hardware utilization of each individual computing device.” [0025]).
Regarding claim 7, Chen discloses: The first device according to claim 5, the processing circuitry being further configured to:
Perform backpropagation to obtain an error vector of the input layer of the first device NN (“A backpropagation function b.sub.i is any function that computes a gradient of the error between an output activation at a network layer i and the expected output of the network at the network layer i. The backpropagation function b.sub.i can use the output gradient computed by a previous layer i+1 to compute the error and obtain the output gradient for the network layer i. Each network layer has a plurality of network parameters that are adjusted as part of neural network training, i.e., weight matrices and bias vectors. The network parameters are utilized within each network layer during training when performing operations such as tensor productions, convolution or attention mechanism.” [0042]); and
Broadcast the obtained error vector to each of the at least two devices, wherein the NNs of the at least two second devices are configured to perform backpropagation based on the broadcast error vector (“A backpropagation function b.sub.i is any function that computes a gradient of the error between an output activation at a network layer i and the expected output of the network at the network layer i. The backpropagation function b.sub.i can use the output gradient computed by a previous layer i+1 to compute the error and obtain the output gradient for the network layer i. Each network layer has a plurality of network parameters that are adjusted as part of neural network training, i.e., weight matrices and bias vectors. The network parameters are utilized within each network layer during training when performing operations such as tensor productions, convolution or attention mechanism.” [0042]).
Regarding claim 8, Chen discloses: The first device according to claim 1, wherein the first device is a base station or a unit attachable to a base station (Fig 1).
Regarding claims 9, 17 and 20, Chen discloses: A communication network, comprising:
at least two second devices that each comprise a neural network (NN), each respective second device of the at least two second devices being configured to (Fig 1):
Obtain a respective input for its respective NN (“The described system allows for effectively distributing the training of giant neural networks, i.e., neural networks with a very large number of parameters, across multiple computing devices. By partitioning the neural network into a plurality of composite layers, the described system can scale arbitrary giant neural network architectures beyond the memory limitations of a single computing device.” [0013]);
Concurrently run its respective NN based on the respective input to obtain a respective output (Fig 1, “The described system allows for effectively distributing the training of giant neural networks, i.e., neural networks with a very large number of parameters, across multiple computing devices. By partitioning the neural network into a plurality of composite layers, the described system can scale arbitrary giant neural network architectures beyond the memory limitations of a single computing device.” [0013]); and
Provide the respective output to a first device (Fig 1, “The described system allows for effectively distributing the training of giant neural networks, i.e., neural networks with a very large number of parameters, across multiple computing devices. By partitioning the neural network into a plurality of composite layers, the described system can scale arbitrary giant neural network architectures beyond the memory limitations of a single computing device.” [0013]).
Regarding claim 10, Chen discloses: the communication network according to claim 9, wherein the respective inputs of the at least two second devices are distinct (“Each device 106 can be heterogeneous, e.g., have multiple processing units each of a different type. The computing devices 106 can be heterogeneous and include devices with different types of processing units that can vary from device-to-device. Alternatively, each device in the plurality of computing devices 106 can include the same number and types of processing units.” [0049]).
Regarding claim 11, Chen discloses: The communication network according to claim 10, wherein each respective input of the at least two second devices is a subset of a complete dataset obtained from a common data source or from different data sources (“The batch splitting engine 108 of the system 100 takes as input the training examples 101 including a plurality of mini-batches 105, and splits each mini-batch into a plurality of micro-batches 107 of equal size, as shown in FIG. 1.” [0055]).
Regarding claim 12, Chen discloses: The communication network according to claim 9, wherein each respective second device of the at least two second devices is further configured to:
Receive a respective error vector from the first device; and Perform backpropagation based on the received respective error vector (“A backpropagation function b.sub.i is any function that computes a gradient of the error between an output activation at a network layer i and the expected output of the network at the network layer i. The backpropagation function b.sub.i can use the output gradient computed by a previous layer i+1 to compute the error and obtain the output gradient for the network layer i. Each network layer has a plurality of network parameters that are adjusted as part of neural network training, i.e., weight matrices and bias vectors. The network parameters are utilized within each network layer during training when performing operations such as tensor productions, convolution or attention mechanism.” [0042]).
Regarding claim 13, Chen discloses: The communication network according to claim 9, wherein one or more parameters of the respective NNs of the at least two second devices are different, wherein the one or more parameters comprise at least one of the following:
A number of layers;
A number of neurons in each layer;
Weights; and
Biases (“Each network layer has a plurality of network parameters that are adjusted as part of neural network training, i.e., weight matrices and bias vectors. The network parameters are utilized within each network layer during training when performing operations such as tensor productions, convolution or attention mechanism.” [0042]).
Regarding claim 14, Chen discloses: The communication network according to claim 9, wherein each of the at least two second devices is a base station or a unit attachable to a base station (Fig 1).
Regarding claim 15, Chen discloses: A communication system, comprising:
A communication network comprising at least two second devices each respective second device of the at least two second devices: comprising a respective neural network (NN) and
A first device comprising:
A first device NN (Fig 1), and
Processing circuitry configured to:
Obtain outputs of the respective NNs of the at least two second devices (“The described system allows for effectively distributing the training of giant neural networks, i.e., neural networks with a very large number of parameters, across multiple computing devices. By partitioning the neural network into a plurality of composite layers, the described system can scale arbitrary giant neural network architectures beyond the memory limitations of a single computing device.” [0013]);
Combine the obtained outputs to generate a combined output (“The K composite layers 103 form the sequence defined by the giant neural network 102, starting from a first composite layer P.sub.1 that includes the input layer i=1 for the neural network, and ending with a last composite layer P.sub.K that includes the output layer i=L for the neural network. In this specification, the succeeding and preceding composite layers relative to a particular composite layer in the sequence are sometimes called neighboring composite layers for the particular composite layer.” [0046], see also [0047]);
Provide the combined output as an input to the first device NN (Fig 1, [0046-47]); and
Run the first device NN based on the input (Fig 1, [0046-47]),
Wherein the at least two second devices are configured to:
Obtain an input for the respective NNs of the at least two second devices (“The described system allows for effectively distributing the training of giant neural networks, i.e., neural networks with a very large number of parameters, across multiple computing devices. By partitioning the neural network into a plurality of composite layers, the described system can scale arbitrary giant neural network architectures beyond the memory limitations of a single computing device.” [0013]);
Concurrently run the respective NNs of the at least two second devices based on the respective obtained input to obtain outputs (“The described system allows for effectively distributing the training of giant neural networks, i.e., neural networks with a very large number of parameters, across multiple computing devices. By partitioning the neural network into a plurality of composite layers, the described system can scale arbitrary giant neural network architectures beyond the memory limitations of a single computing device.” [0013]); and
Provide output to the first device (Fig 1, [0046-47]).
Regarding claims 18 and 21, Chen discloses: a method for training neural networks (NNs) comprised in a first device and at least two second devices, the method comprising:
Obtaining, by each respective second device of the at least two second devices, an input for a respective NN of the respective second device (Fig 1)
Concurrently running, by the at least two second devices, the respective NNs based on the respective input to obtain respective outputs (“The described system allows for effectively distributing the training of giant neural networks, i.e., neural networks with a very large number of parameters, across multiple computing devices. By partitioning the neural network into a plurality of composite layers, the described system can scale arbitrary giant neural network architectures beyond the memory limitations of a single computing device.” [0013]);
Providing, by each respective second device of the at least two second devices, a respective output to the first device (Fig 1, “The described system allows for effectively distributing the training of giant neural networks, i.e., neural networks with a very large number of parameters, across multiple computing devices. By partitioning the neural network into a plurality of composite layers, the described system can scale arbitrary giant neural network architectures beyond the memory limitations of a single computing device.” [0013]);
Running, by the first device, the NN of the first device based on the input (Fig 1, [0046-47]);
Obtaining, by the first device, an output of the NN of the first device (Fig 1, [0046-47]);
Performing, by the first device, backpropagation to optimize parameters of the NN of the first device based on the obtained output of the NN of the first device, and to obtain an error vector of an input layer of the NN of the first device (“A backpropagation function b.sub.i is any function that computes a gradient of the error between an output activation at a network layer i and the expected output of the network at the network layer i. The backpropagation function b.sub.i can use the output gradient computed by a previous layer i+1 to compute the error and obtain the output gradient for the network layer i. Each network layer has a plurality of network parameters that are adjusted as part of neural network training, i.e., weight matrices and bias vectors. The network parameters are utilized within each network layer during training when performing operations such as tensor productions, convolution or attention mechanism.” [0042]);
Providing, by the first device, a respective error vector to each second device of the at least two second devices (fig 1, [0042]);
Receiving, by each second devices of the at least two second devices, the respective error vector from the first device (fig 1, [0042]); and
Respectively performing, by each respective second device of the at least two second devices, backpropagation based on the received respective error vector to optimize parameters of the respective NN of the respective second device (fig 1, [0013], [0042]).
Response to Arguments
Applicant’s arguments with respect to claim(s) 1-20 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ERIC NILSSON whose telephone number is (571)272-5246. The examiner can normally be reached M-F: 7-3.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, James Trujillo can be reached at (571)-272-3677. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/ERIC NILSSON/Primary Examiner, Art Unit 2151