DETAILED ACTION
This action is responsive to the claims filed on 05/06/2026. Claims 1, 7-10, 15, 17-18, 20, and 22-27 are pending for examination.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Arguments
Applicant’s arguments with respect to the 35 U.S.C. 102/103 rejection of the claims have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 1, 10, 15, 22, and 24-27 are rejected under 35 U.S.C. 103 as being unpatentable by Cui et al, (Cui, G., Xu, J., Zeng, W., Lan, Y., Guo, J., & Cheng, X. (2018, September). MQGrad: Reinforcement learning of gradient quantization in parameter server. In Proceedings of the 2018 ACM SIGIR International Conference on Theory of Information Retrieval (pp. 83-90).), hereafter referred to as Cui, in view of Khronos NNEF (The Khronos NNEF Working Group. (2019). Neural Network Exchange Format. Khronos.Org. registry.khronos.org/NNEF/specs/1.0/nnef-1.0.2.html), hereafter referred to as NNEF, and in further view of Chen et al.,(Z. Chen, K. Fan, S. Wang, L. Duan, W. Lin and A. C. Kot, "Toward Intelligent Sensing: Intermediate Deep Feature Compression," in IEEE Transactions on Image Processing, vol. 29, pp. 2230-2243, 2020, doi: 10.1109/TIP.2019.2941660), hereafter referred to as Chen.
Claim 1: Cui teaches the following limitations:
A method for training a machine learning model, comprising: sending, by a first node, a first quantization strategy to a second node, (Cui, page 86, col. 1, paragraph 2, “The local values at all of the workers are collected by the sever MDP module (step 2 in Fig. 3, line 38 of Alg. 1). After that, the MDP module at server restores the overall global loss, updates its state, calculates the reward, determines the action (the quantization bits), and finally broadcasts the number to all worker nodes (step 3 in Fig. 3, line 39-47 in Alg. 1)”, Cui teaches a server (first node) selecting quantization bits and broadcasting them to worker nodes (second nodes). The broadcasted quantization bits K correspond to the claimed first quantization strategy (i.e., the quantization setting/level sent from the first node to the second node)
receiving, by the first node from the second node, quantized information obtained by the second node performing quantization on to-be-output information according to the first quantization strategy and a first parameter corresponding to the first quantization strategy; (Cui, page 86, col. 1, section 3.1, paragraph 2, “Given the quantization bits, the worker nodes quantize1 their local gradients (step 4 in Fig. 3, line 30 in Alg. 1) and send the quantized local gradients to the parameter server (step 5 in Fig. 3, line 31 in Alg. 1).”, Cui teaches the second node (workers) performing quantization on “to-be-output information” (local gradients) using the selected quantization bits (the corresponding quantization “parameter”) and sending the quantized gradients to the first node (parameter server), which therefore receives the quantized information. )
and restoring the quantized information according to the first quantization strategy and the first parameter, thereby determining a parameter or an output result of the machine learning model. (Cui, page 86, col. 1, section 3.1, paragraph 2, “Then the server broadcasts the quantized global gradient (step 8 in Fig. 3, line 53 in Alg. 1) and the workers receive it, de-quantize the gradient, and update the local model parameters (step 9 in Fig. 3, line 32-33 in Alg. 1).”, Cui expressly teaches de-quantizing (“restoring”) the quantized gradient and then using it to update model parameters, which reads on restoring quantized information to determine/update a parameter (and thus an output result) of the ML model. )
NNEF, in the same field of neural networks, teaches the following which Cui fails to teach:
wherein the first quantization strategy comprises one strategy selected from a uniform quantization strategy and a non-uniform quantization strategy; (NNEF, section 5.2, “The quantization algorithm vendor code 0x0 identifies the Khronos Group. For the algorithm code part, Khronos specifies the following: 0x00 - uncompressed float values in IEEE format. Valid bits per item is 16, 32, 64. 0x01 - uncompressed integer values. 0x10 - quantized values with linear quantization as defined by the linear_quantize operation 0x11 - quantized values with logarithmic quantization as defined by the logarithmic_quantize operation”; section 4.9.5, “fragment linear_quantize( x: tensor<scalar>, min: tensor<scalar>, max: tensor<scalar>, bits: integer ) -> ( y: tensor<scalar> ) { r = scalar(2 ^ bits - 1); z = clamp(x, min, max); q = round((z - min) / (max - min) * r); y = q / r * (max - min) + min; } fragment logarithmic_quantize( x: tensor<scalar>, max: tensor<scalar>, bits: integer ) -> ( y: tensor<scalar> ) { m = ceil(log2(max)); r = scalar(2 ^ bits - 1); q = round(clamp(log2(abs(x)), m - r, m)); y = sign(x) * 2.0 ^ q; }”, NNEF expressly provides two selectable quantization algorithms: linear quantization and logarithmic quantization. The linear quantizer uniformly maps the input range between min and max across the available integer representation, whereas the logarithmic quantizer operates in a logarithmic domain and reconstructs values as powers of two. NNEF therefore teaches the recited selection between a uniform quantization strategy and a non-uniform quantization strategy. )
wherein the first parameter is defined in a protocol; (NNEF, section 5.2, “Each tensor data file consists of two parts, the header and the data. All data is laid out in little-endian byte order. The header size is fixed to 128 bytes, and consists of the following (in this order):… 32 bytes for parameters of the quantization algorithm. The interpretation of these 32 bytes depends on the algorithm code… 0x10 - quantized values with linear quantization as defined by the linear_quantize operation 0x11 - quantized values with logarithmic quantization as defined by the logarithmic_quantize operation… The codes 0x10 and 0x11 have corresponding parameters as follows: min: 32-bit IEEE floating point value indicating the value to which an item containing all-zero bits is mapped max: 32-bit IEEE floating point value indicating the value to which an item containing all-one bits is mapped”, NNEF is itself a standardized exchange-format specification and expressly defines how a quantization algorithm identifier corresponds to the parameters used by that quantization algorithm. Thus, rather than requiring the receiving implementation to independently derive the meaning of a quantization parameter, NNEF prescribes the meaning and interpretation of the parameter as part of the exchange protocol. Accordingly, NNEF teaches a first parameter corresponding to a quantization strategy that is defined in a protocol. NNEF additionally explains that quantization information associated with training is conveyed when transferring a trained network to another implementation. )
wherein the method further comprises: sending, by the first node, to the second node a correspondence between the first quantization strategy and the first parameter, thereby enabling the second node to determine the first parameter upon receiving the first quantization strategy. (NNEF, section 5.2, “Each tensor data file consists of two parts, the header and the data. All data is laid out in little-endian byte order. The header size is fixed to 128 bytes, and consists of the following (in this order):… 32 bytes for parameters of the quantization algorithm. The interpretation of these 32 bytes depends on the algorithm code… 0x10 - quantized values with linear quantization as defined by the linear_quantize operation 0x11 - quantized values with logarithmic quantization as defined by the logarithmic_quantize operation… The codes 0x10 and 0x11 have corresponding parameters as follows: min: 32-bit IEEE floating point value indicating the value to which an item containing all-zero bits is mapped max: 32-bit IEEE floating point value indicating the value to which an item containing all-one bits is mapped”; Section 5.1, “A binary data file (structured according to Tensor File Format) for each variable tensor in the structure description, placed into sub-folders according to the labelling of the variables. The files must have '.dat' extension. For example, a variable with label 'conv1/filter' is placed into the folder 'conv1' under the name 'filter.dat'. Different versions of the same data (for example with different quantization) may be present starting with the same name, having additional (arbitrary) extensions after '.dat'.”, NNEF section 5.2 defines a transmitted tensor header in which the quantization algorithm is represented by a defined algorithm code and the associated quantization parameters are interpreted according to that algorithm code. In particular, algorithm code 0x11 identifies the standardized logarithmic quantizer, and the specification defines the corresponding parameter arrangement and base-2 logarithmic/power reconstruction rule. NNEF also expressly provides a separate quantization-description file, as stated in section 5.1, in which tensor identifiers are associated with the applicable quantization operation and its named parameters, and explains that different algorithms and parameters may be assigned to different tensors. NNEF therefore teaches a protocol-defined correspondence linking an identified quantization strategy to the parameters used by that strategy. Upon receipt of the standardized algorithm identifier and associated quantization information, a receiving implementation determines, from the correspondence prescribed by the NNEF format, the parameters and quantization rule associated with the received quantization strategy. NNEF’s defined algorithm/parameter correspondence would enable a worker node to determine the corresponding quantization parameter upon receipt of the strategy.)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Cui’s distributed machine-learning training system to employ the standardized quantization techniques and quantization-strategy/parameter correspondence taught by NNEF. Cui teaches communicating a selected quantization setting between a parameter server and worker nodes to reduce communication associated with distributed training, while NNEF teaches standardized linear and logarithmic quantization operations and a protocol-defined association between an identified quantization algorithm and its corresponding parameters. NNEF further recognizes that quantization information associated with training is to be conveyed when transferring a trained neural network. A person of ordinary skill therefore would have been motivated to employ NNEF’s standardized quantization definitions and parameter correspondence in Cui so that communicating nodes could consistently identify and apply the selected quantization method and corresponding parameters, thereby promoting interoperability and reliable quantization/dequantization while retaining the communication-efficiency benefits sought by Cui.
Chen, in the same field of neural networks, teaches the following which Cui and NNEF fails to teach:
wherein the first parameter comprises a quantization base number of a non-uniform power quantization; (Chen, page 2239, col. 1, paragraph 2, “In the encoding phase, pre-quantization module is first applied to convert the floating point deep learning feature data to integers. It is necessary since that deep learning features, like the vanilla VGGNets and ResNets features, are in float32 format, which are not compatible with the desired input format of most video codecs. For instance, HEVC and AVC require 8-bit (or higher) integers as the input. In view of that the intermediate deep feature generally has a right-skewed exponential distribution with a wide data span (histograms of intermediate deep features are presented in the supplemental material as Fig. 7), we quantize the features to 8-bit precision with logarithmic sampling in this paper. The quantization and corresponding dequantization (i.e. inverse quantization) are performed as
PNG
media_image1.png
81
333
media_image1.png
Greyscale
where X can be the feature tensor of a certain input from any specific layer of the neural network; Xquant and Xdequant are the corresonding quantized and dequantized feature tensor; round(·) rounds the input float value to the nearest integer; B is the base of logarithm which can be any real number;”, Chen expressly identifies a numerical base B as a parameter of a logarithmic quantization process. Because the quantization operates logarithmically rather than by equal linear spacing and the corresponding inverse process reconstructs values through a power-domain operation, Chen’s B constitutes a quantization base number of a non-uniform power quantization.)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to implement the non-uniform quantization alternative of the modified Cui/NNEF system using the logarithmic quantization technique taught by Chen. Chen teaches logarithmic sampling of deep-learning information and expressly defines a base B for the logarithmic quantization operation, thereby providing a known quantization-base parameter for non-uniform quantization of neural-network information. Chen explains that such quantization is employed to compactly represent deep-learning information having a wide data span for transmission and subsequent reconstruction. Accordingly, use of Chen’s known logarithmic-base quantization in the distributed neural-network communication system of Cui as modified by NNEF would have constituted a predictable application of a known neural-network quantization technique to reduce the representation and communication burden while permitting corresponding dequantization at the receiving node.
Claim 10: Cui teaches:
A method for training a machine learning model, comprising: receiving, by a second node, a first quantization strategy from a first node, (Cui, page 86, section 3.1, paragraph 2, “After that, the MDP module at server restores the overall global loss, updates its state, calculates the reward, determines the action(the quantization bits),and finally broadcasts the number to all worker nodes (step3inFig.3, line39-47inAlg.1).”; Page 86, Algorithm 1, “receive quantize bits K from server”, Cui expressly teaches a worker node receiving the server-selected quantization setting from the parameter server. The parameter server constitutes the first node and the receiving worker constitutes the second node.)
obtaining, by the second node, quantized information by performing quantization on to-be-output information according to the first quantization strategy and a first parameter corresponding to the first quantization strategy; (Cui, page 86, section 3.1, paragraph 2, “Given the quantization bits, the worker nodes quantize1 their local gradients (step 4 in Fig. 3, line 30 in Alg. 1) and send the quantized local gradients to the parameter server (step 5 in Fig. 3, line 31 in Alg. 1).”, Cui teaches the second node (worker) performs quantization on to-be-output information using the received quantization bits K (the corresponding quantization “parameter”))
and sending the quantized information to the first node, wherein the quantized information is used by the first node for restoring the quantized information according to the first quantization strategy and the first parameter, thereby determining a parameter and/or an output result of the machine learning model. (Cui, page 86, section 3.1, paragraph 2, “The server nodes de-quantize and summarize all of the received local gradients to a global gradient for updating the model parameters (step 6 and 7 in Fig. 3, line 51-52 in Alg. 1).”, Cui teaches that workers (second nodes) send the quantized local gradients to the parameter server (first node). Cui further teaches the first node (server) de-quantizes (restores) and summarizes the received quantized local gradients to form a global gradient for updating the model parameters, thereby determining a parameter and/or an output result of the machine learning model.)
NNEF, in the same field of neural networks, teaches the following which Cui fails to teach:
wherein the first quantization strategy comprises one strategy selected from a uniform quantization strategy and a non-uniform quantization strategy; (NNEF, section 4.9.5, “fragment linear_quantize( x: tensor<scalar>, min: tensor<scalar>, max: tensor<scalar>, bits: integer ) -> ( y: tensor<scalar> ) { r = scalar(2 ^ bits - 1); z = clamp(x, min, max); q = round((z - min) / (max - min) * r); y = q / r * (max - min) + min; } fragment logarithmic_quantize( x: tensor<scalar>, max: tensor<scalar>, bits: integer ) -> ( y: tensor<scalar> ) { m = ceil(log2(max)); r = scalar(2 ^ bits - 1); q = round(clamp(log2(abs(x)), m - r, m)); y = sign(x) * 2.0 ^ q; }”, Cui teaches that the server (first node) selects the quantization bits K and broadcasts them to the worker nodes (second nodes), and the worker function explicitly “receive[s] quantize bits K from server.” NNEF teaches expressly defined alternative linear and logarithmic quantization techniques. The linear technique distributes representable values according to a linear scale over the specified range, whereas the logarithmic technique quantizes in the logarithmic domain and reconstructs values as powers of two. Accordingly, NNEF teaches the recited alternative uniform and non-uniform quantization strategies.)
wherein the first parameter is defined in a protocol; (NNEF, section 5.2, “Each tensor data file consists of two parts, the header and the data. All data is laid out in little-endian byte order. The header size is fixed to 128 bytes, and consists of the following (in this order):… 32 bytes for parameters of the quantization algorithm. The interpretation of these 32 bytes depends on the algorithm code… 0x10 - quantized values with linear quantization as defined by the linear_quantize operation 0x11 - quantized values with logarithmic quantization as defined by the logarithmic_quantize operation… The codes 0x10 and 0x11 have corresponding parameters as follows: min: 32-bit IEEE floating point value indicating the value to which an item containing all-zero bits is mapped max: 32-bit IEEE floating point value indicating the value to which an item containing all-one bits is mapped”, NNEF is itself a standardized exchange-format specification and expressly defines how a quantization algorithm identifier corresponds to the parameters used by that quantization algorithm. Thus, rather than requiring the receiving implementation to independently derive the meaning of a quantization parameter, NNEF prescribes the meaning and interpretation of the parameter as part of the exchange protocol. Accordingly, NNEF teaches a first parameter corresponding to a quantization strategy that is defined in a protocol. NNEF additionally explains that quantization information associated with training is conveyed when transferring a trained network to another implementation. )
receiving, by the second node, a correspondence between the first quantization strategy and a first parameter; (NNEF, section 5.3, “The quantization file has a simple line based textual format. Each line contains a tensor identifier … and a quantization algorithm terminated by a semicolon. … "filter1": linear_quantize(min = -2.0, max = 2.5, bits = 8); "bias1": linear_quantize(min = -3.0, max = 0.75, bits = 8);… The quantization algorithm string must not contain an argument for the tensor to be quantized, but all other parameters must be provided as named arguments with compile-time expressions …”; Section 7, paragraph 1, “when transferring trained networks to inference engines, the quantization method used during training needs to be conveyed by NNEF.”, NNEF teaches conveying quantization information that expressly associates a particular quantization algorithm with the named parameters required for that algorithm. Accordingly, the receiving implementation obtains information establishing a correspondence between the applicable quantization strategy and the first parameter associated with that strategy.)
and determining, by the second node, the first parameter corresponding to the first quantization strategy according to the correspondence between the first quantization strategy and the first parameter. (NNEF, section 5.2, “A 4-byte code indicating the quantization (compression) algorithm used to store the data. … 32 bytes for parameters of the quantization algorithm. The interpretation of these 32 bytes depends on the algorithm code… 0x10 - quantized values with linear quantization as defined by the linear_quantize operation 0x11 - quantized values with logarithmic quantization as defined by the logarithmic_quantize operation”; Section 5.3, paragraph 2, “all other parameters must be provided as named arguments with compile-time expressions.”, NNEF expressly defines the interpretation of the parameter information as dependent upon the quantization algorithm code and prescribes the corresponding parameters for the identified algorithm. Thus, after identifying the applicable quantization strategy, the receiving implementation determines the parameter associated with that strategy according to the protocol-defined correspondence between the algorithm and its parameter information.)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the distributed machine-learning training system of Cui to incorporate the standardized quantization techniques and quantization-strategy/parameter correspondence taught by NNEF. Cui teaches reducing communication overhead in distributed training by having a parameter server determine a quantization setting and communicate that setting to worker nodes for use in quantizing training information, while NNEF teaches standardized linear and logarithmic quantization operations and a defined association between an identified quantization algorithm and the parameters corresponding to that algorithm. A person of ordinary skill in the art would have been motivated to apply NNEF’s standardized quantization definitions and corresponding parameter information to Cui’s distributed training architecture in order to provide communicating nodes with a consistent and interoperable mechanism for identifying and applying the selected quantization technique and its corresponding parameters, thereby facilitating proper quantization and dequantization while preserving the communication-efficiency benefits sought by Cui.
Chen, in the same field of neural networks, teaches the following which Cui and NNEF fails to teach:
wherein the first parameter comprises a quantization base number of a non-uniform power quantization; (Chen, page 2239, col. 1, paragraph 2, “In the encoding phase, pre-quantization module is first applied to convert the floating point deep learning feature data to integers. It is necessary since that deep learning features, like the vanilla VGGNets and ResNets features, are in float32 format, which are not compatible with the desired input format of most video codecs. For instance, HEVC and AVC require 8-bit (or higher) integers as the input. In view of that the intermediate deep feature generally has a right-skewed exponential distribution with a wide data span (histograms of intermediate deep features are presented in the supplemental material as Fig. 7), we quantize the features to 8-bit precision with logarithmic sampling in this paper. The quantization and corresponding dequantization (i.e. inverse quantization) are performed as
PNG
media_image1.png
81
333
media_image1.png
Greyscale
where X can be the feature tensor of a certain input from any specific layer of the neural network; Xquant and Xdequant are the corresonding quantized and dequantized feature tensor; round(·) rounds the input float value to the nearest integer; B is the base of logarithm which can be any real number;”, Chen’s expressly defined base B is a numerical parameter governing the disclosed logarithmic/non-uniform quantization process, thereby teaching the claimed quantization base number of a non-uniform power quantization.)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the combined teachings of Cui and NNEF to employ the logarithmic quantization technique taught by Chen. Chen teaches quantizing deep-learning feature information using logarithmic sampling and expressly identifies a logarithmic base B as a parameter of the quantization and corresponding dequantization operation. Because the combined Cui/NNEF system already teaches communicating and applying quantization strategies and their corresponding parameters in a machine-learning environment, a person of ordinary skill in the art would have been motivated to employ Chen’s known logarithmic quantization technique as the non-uniform quantization alternative, including its logarithmic-base parameter, in order to provide an efficient known technique for representing neural-network information having a wide dynamic range with reduced precision while permitting corresponding reconstruction of the quantized information at the receiving node. Such a modification would have amounted to the predictable use of a known neural-network quantization technique according to its established function to further reduce the amount of information required to represent and communicate machine-learning data.
Claim 15: Cui, NNEF, and Chen teaches the limitations of claim 10 and NNEF further teaches the following limitations:
The method as claimed in claim 10, wherein: the correspondence between the quantization strategy and the first parameter is defined in a protocol. (NNEF, section 5.2, “A 4-byte code indicating the quantization (compression) algorithm used to store the data. The code is split to two 2-byte parts. The first part is a vendor code, the second part is an algorithm code. … 32 bytes for parameters of the quantization algorithm. The interpretation of these 32 bytes depends on the algorithm code… 0x10 - quantized values with linear quantization as defined by the linear_quantize operation0x11 - quantized values with logarithmic quantization as defined by the logarithmic_quantize operation… The codes 0x10 and 0x11 have corresponding parameters as follows: min: 32-bit IEEE floating point value indicating the value to which an item containing all-zero bits is mapped; max: 32-bit IEEE floating point value indicating the value to which an item containing all-one bits is mapped.”, The NNEF specification expressly establishes the relationship between the quantization-algorithm identifier and the parameter information applicable to the identified algorithm. The interpretation of the parameter field expressly depends upon the algorithm code, and NNEF prescribes the corresponding parameter meanings for its standardized linear and logarithmic quantization algorithms. Accordingly, NNEF teaches a correspondence between a quantization strategy and the first parameter that is defined in a protocol.)
Claim 22: Cui, NNEF, and Chen teaches the limitations of claim 1 and NNEF further teaches the following limitations:
The method as claimed in claim 1, further comprising: sending to the second node the first parameter corresponding to the first quantization strategy. (NNEF, section 5.1, “An optional quantization file (structured according to Quantization File Format) containing quantization algorithm details for exported tensors. The file must be named ‘graph.quant’.”; section 5.3, “The quantization file has a simple line based textual format. Each line contains a tensor identifier … and a quantization algorithm terminated by a semicolon. … The quantization algorithm string must not contain an argument for the tensor to be quantized, but all other parameters must be provided as named arguments with compile-time expressions… ‘filter1’: linear_quantize(min = -2.0, max = 2.5, bits = 8); ‘bias1’: linear_quantize(min = -3.0, max = 0.75, bits = 8);… when transferring trained networks to inference engines, the quantization method used during training needs to be conveyed by NNEF.”, NNEF expressly teaches conveying quantization information from a producing implementation to a consuming implementation, wherein the conveyed quantization description identifies the applicable quantization algorithm and includes the corresponding named quantization parameters. Thus, NNEF teaches communicating the parameter information corresponding to the quantization strategy to the receiving implementation. In the combined Cui/NNEF system, this known quantization-description mechanism is applied to Cui's first-node/second-node communications such that the first node sends the parameter corresponding to the selected quantization strategy to the second node.)
Claim 24: Cui, NNEF, and Chen teaches the limitations of claim 1 and NNEF further teaches the following limitations:
The method as claimed in claim 1, wherein the uniform quantization strategy is performed by dividing a value range of input signal at equal intervals; (NNEF, section 4.9.5, “fragment linear_quantize( x: tensor<scalar>, min: tensor<scalar>, max: tensor<scalar>, bits: integer ) -> ( y: tensor<scalar> ) { r = scalar(2 ^ bits - 1); z = clamp(x, min, max); q = round((z - min) / (max - min) * r); y = q / r * (max - min) + min; } ”, NNEF’s linear quantization expressly divides the value range extending from min to max according to the number of available quantization indices. Specifically, NNEF defines r = 2^bits - 1 and determines the quantization index according to q = round((z - min)/(max - min) * r). The term (z - min)/(max - min) identifies the relative position of the input value within the range from min to max, and multiplication by r maps that range to successive integer quantization indices from 0 through r. NNEF then reconstructs a value according to y = q/r * (max - min) + min. Therefore, increasing the quantization index from q to q+1 increases the reconstructed value by exactly (max - min)/r, which is constant for every successive pair of quantization levels. In other words, the total value range (max - min) is divided into r equal increments, each having a width of (max - min)/r (for example, an 8-bit quantizer provides r = 255 equal increments between 256 reconstruction levels). Accordingly, NNEF teaches performing uniform quantization by dividing the value range of the input signal at equal intervals.)
and the non-uniform quantization strategy is performed by dividing the value range of input signal at non-equal intervals based on a size of a signal range. (NNEF, section 4.9.5, “fragment logarithmic_quantize( x: tensor<scalar>, max: tensor<scalar>, bits: integer ) -> ( y: tensor<scalar> ) { m = ceil(log2(max)); r = scalar(2 ^ bits - 1); q = round(clamp(log2(abs(x)), m - r, m)); y = sign(x) * 2.0 ^ q; }”; section 5.2, “The codes 0x10 and 0x11 have corresponding parameters as follows: min: 32-bit IEEE floating point value indicating the value to which an item containing all-zero bits is mapped; max: 32-bit IEEE floating point value indicating the value to which an item containing all-one bits is mapped. In case of the logarithmic quantization algorithm (0x11), the min parameter must be set to 0 (unsigned) or -max (signed).”; “In case of logarithmic quantization, let [m = ceil(log2(max))] …”, NNEF’s logarithmic quantization divides the represented input-value range into intervals that are uniform in the logarithmic domain but non-equal in the original linear-value domain. Specifically, NNEF determines m = ceil(log2(max)), where max identifies the maximum represented signal magnitude, and determines the quantization index from log2(abs(x)) over the logarithmic range extending from m-r to m, where r = 2^bits - 1. NNEF then reconstructs the quantized values as powers of two. Thus, successive reconstruction levels have magnitudes corresponding to successive powers of two, e.g., 2^k, 2^(k+1), 2^(k+2), rather than values separated by a constant linear step. The linear interval between 2^k and 2^(k+1) is 2^k, while the next interval between 2^(k+1) and 2^(k+2) is 2^(k+1); therefore, the intervals increase with signal magnitude and are not equal. Further, the specific set of non-equal intervals is based on the size of the represented signal range because NNEF calculates the upper logarithmic boundary m directly from max, and NNEF specifies that the logarithmic quantizer’s lower endpoint is either 0 for an unsigned range or -max for a signed range. Accordingly, max establishes the magnitude/extent of the signal range and, through m = ceil(log2(max)), determines where the non-equally spaced logarithmic quantization intervals occur. NNEF therefore teaches performing non-uniform quantization by dividing the value range of the input signal at non-equal intervals based on the size of the signal range.)
Claim 25: Cui, NNEF, and Chen teaches the limitations of claim 1 and NNEF further teaches the following limitations:
The method as claimed in claim 1, wherein the first parameter further comprises at least one of: a quantization range; a quantization level; and a quantization bit number. (NNEF, section 4.9.5, expressly defines the linear quantization operation with the parameters: “min: tensor<scalar>,max: tensor<scalar>, bits: integer” and uses those parameters in: “r = scalar(2 ^ bits - 1);z = clamp(x, min, max); q = round((z - min) / (max - min) * r);”
NNEF additionally states at section 5.2:
“A 4-byte unsigned integer indicating the number of bits per item, at most 64.” and:
“The codes 0x10 and 0x11 have corresponding parameters as follows: min: 32-bit IEEE floating point value indicating the value to which an item containing all-zero bits is mapped; max: 32-bit IEEE floating point value indicating the value to which an item containing all-one bits is mapped.”, NNEF expressly uses min and max to establish the quantization range and expressly identifies bits as the number of bits used for the quantized representation. Accordingly, NNEF directly teaches both a quantization range and a quantization bit number. Because claim 25 requires only at least one of the recited alternatives, either disclosure independently satisfies the claimed additional first parameter.)
Claims 26-27 are substantially similar to claims 24-25, respectively, as such a similar analysis applies.
Claims 7, 17, and 20 are rejected under 35 U.S.C. 103 as being unpatentable by Cui in view of NNEF, Chen, and Mellempudi et al. (US 2021/0342692 A1), hereafter referred to as Mellempudi.
Claim 7: Cui, NNEF, and Chen teaches the limitations of claim 1 and Mellumpudi, in the same field of neural network quantization, further teaches the following limitations which Cui, NNEF, and Chen fails to teach:
The method as claimed in claim 1, wherein: the first quantization strategy is predefined; or the first quantization strategy is pre-configured. (Mellempudi, paragraph 22, “Prior to sending the message to the receiver computing node ( s ) 102b , the application 202 may request the host fabric interface 124 to reduce the size of the message by quantizing the message based on a quantization level determined by the application”, The quantization strategy may be set in advance by the application or protocol, thus is pre-defined or pre-configured.)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Mellempudi’s quantized-message transmission techniques (e.g., quantizing transmitted training values and including metadata indicative of quantization level) into Cui’s/NNEF/Chen’s parameter-server gradient quantization system, as both address reducing distributed-training communication overhead and both communicate quantization settings/levels along with transmitted training-related information.
Claims 17 are substantially similar to claims 7, as such a similar analysis applies.
Claim 20: Claim recites limitations substantially similar to claim 1, as such a similar analysis applies. Claim 20 also recites the following additional limitation for consideration taught by Mellumpudi:
A node device, comprising: a processor and a memory configured to store a computer program executable on the processor; (Mellempudi, paragraph 16, “The memory 128 may be embodied as any type of volatile or non - volatile memory or data storage capable of performing the functions described herein . In operation , the memory 128 may store various data and software used during operation of the computing node 102 such as operating systems , applications , programs , libraries , and drivers”)
wherein the processor is configured to, when executing the computer program, implement steps of a method for training a machine learning model, the method comprising (Mellempudi, paragraph 21, “In particular , the application 202 may be embodied as a convolutional neural network , recurrent neural network , or other multilayered artificial neural network and / or related training algorithm . The application 202 may be hosted , executed , or otherwise established by one or more of the processor cores 122 of the processor 120.”)
The rationale for the combination of Cui, NNEF, and Chen with Mellumpudi is substantially similar to that applied to claim 7 above.
Claims 8-9 and 18 are rejected under 35 U.S.C. 103 as being unpatentable by Cui in view of NNEF, Chen, and Rahman et al. (US20180167116A1), hereafter referred to as Rahman.
Claim 8: Cui, NNEF, and Chen teaches the limitations of claim 1, Rahman, in the same field of neural networks, teaches the following which Cui, NNEF, and Chen fail to teach:
The method as claimed in claim 1, wherein the first indication information is carried in any one of service layer data, a broadcast message, an RRC message, a MAC CE, a PDCCH and a DCI. (Rahman, paragraph 168, “If both precoder and matrix quantization are supported, then the UE receives configuration information about one of the two quantization types via higher layer signaling such as RRC and MAC CE”, Rahman expressly teaches communicating information identifying an applicable quantization type to a UE using RRC or MAC CE higher-layer signaling, and alternatively through DCI. Because the claim is satisfied by any one of the recited signaling mechanisms, Rahman's express use of RRC, MAC CE, or DCI to communicate quantization configuration information teaches this limitation.)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention, when implementing the communicating nodes of the modified system as wireless terminal and network devices, to employ Rahman’s known control-signaling mechanisms for communicating quantization configuration information into Cui/NNEF/Chen. Rahman expressly teaches communicating a selected quantization type through RRC or MAC CE signaling and alternatively through DCI. Employing such conventional signaling to carry an indication of the quantization configuration between a network device and terminal device would predictably permit the receiving device to determine the quantization configuration using established wireless control channels.
Claim 9: Cui, NNEF, and Chen teaches the limitations of claim 1 and Rahman, in the same field of neural network quantization, further teaches the following limitations which Cui, NNEF, and Chen fails to teach:
The method as claimed in claim 1, wherein the first node comprises a terminal device or a network device; the second node comprises a terminal device or a network device. (Rahman, abstract, “A UE capable of CSI reporting includes a transceiver configured to receive, from a base station (BS), CSI configuration information including a number (L) of beams and a number (T) of CSI reports.”, Rahman further teaches the UE transmitting the generated CSI reports to the base station. Rahman expressly discloses communication between a base station, which is a network device, and a UE, which is a terminal device. Rahman therefore teaches first and second communicating nodes respectively comprising terminal and/or network devices as recited.)
The rationale for the combination of Cui, NNEF, and Chen with Rahman is substantially similar to that applied to claim 8 above.
Claims 18 are substantially similar to claims 8, as such a similar analysis applies.
Claim 23 is rejected under 35 U.S.C. 103 as being unpatentable by Cui in view of NNEF, Chen, and Shih et al., (Wen, C. K., Shih, W. T., & Jin, S. (2018). Deep learning for massive MIMO CSI feedback. IEEE Wireless Communications Letters, 7(5), 748-751.), hereafter referred to as Shih.
Claim 23: Cui, NNEF, and Chen teaches the limitations of claim 1. Shih, in the same field of parameter quantization, teaches the following limitations which Cui, NNEF, and Chen fails to teach:
The method as claimed in claim 1, wherein the to-be-output information is obtained based on measurement information through an encoding neural network model, and the method further comprises: inputting the output result to a decoding neural network model, (Shih, page 749, col. 2, paragraph 2, “In this letter, we are interested in designing the encoder
PNG
media_image2.png
24
191
media_image2.png
Greyscale
… Following the convolutional layer, we reshape the feature maps into a vector and use a fully connected layer to generate the codeword s,”, Shih explicitly defines an encoder mapping from the CSI/channel matrix H (measurement information) to a lower-dimensional codeword s (to-be-output information). The architecture description further explains implementation: feature maps are reshaped and a fully connected layer generates the codeword s, matching “obtained … through an encoding neural network model.” )
and obtaining the measurement information by using the decoding neural network. (Shih, page 749, col. 2, paragraph 2, “In addition, we have to design the inverse transformation (decoder) from the codeword to the original channel, that is,
PNG
media_image3.png
27
191
media_image3.png
Greyscale
… s is returned to the BS, and the BS uses the decoder (4) to obtain H.”. Shih explicitly defines the decoder inverse mapping from codeword s back to channel matrix H. The protocol description states that after the codeword is returned, the base station uses the decoder to obtain H, which reads on “inputting the output result … to a decoding neural network … obtaining the measurement information.”)
It would have been obvious to a POSITA before the effective filing date of the invention to apply Cui/NNEF/Chen with Shih, because Shih identifies that CSI feedback transmissions are limited by feedback overhead, Shih expressly states that CSI feedback “is hindered by excessive feedback overhead,” and describes the solution of learning “representations (or codewords)” for CSI feedback/reconstruction (Shih, Abstract).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure:
Lin, Y., Han, S., Mao, H., Wang, Y., & Dally, W. J. (2017). Deep gradient compression: Reducing the communication bandwidth for distributed training. arXiv preprint arXiv:1712.01887.
Alistarh, D., Grubic, D., Li, J., Tomioka, R., & Vojnovic, M. (2017). QSGD: Communication-efficient SGD via gradient quantization and encoding. Advances in neural information processing systems, 30.
Wen, W., Xu, C., Yan, F., Wu, C., Wang, Y., Chen, Y., & Li, H. (2017). Terngrad: Ternary gradients to reduce communication in distributed deep learning. Advances in neural information processing systems, 30.
Bernstein, J., Wang, Y. X., Azizzadenesheli, K., & Anandkumar, A. (2018, July). signSGD: Compressed optimisation for non-convex problems. In International conference on machine learning (pp. 560-569). PMLR.
Lloyd, S. (1982). Least squares quantization in PCM. IEEE transactions on information theory, 28(2), 129-137.
Horvath, S., Ho, C.-Y., Horvath, L., Narayan, S. A., Canini, M., & Richtarik, P. (2019). Natural Compression for Distributed Deep Learning. arXiv.Org. arxiv.org/abs/1905.10988v3
Li, Y., Dong, X., & Wang, W. (2019). Additive powers-of-two quantization: An efficient non-uniform discretization for neural networks. arXiv preprint arXiv:1909.13144.
Miyashita, D., Lee, E. H., & Murmann, B. (2016). Convolutional neural networks using logarithmic data representation. arXiv preprint arXiv:1603.01025.
Baskin, C., Liss, N., Schwartz, E., Zheltonozhskii, E., Giryes, R., Bronstein, A. M., & Mendelson, A. (2019). UNIQ. ACM Transactions on Computer Systems, 37(1–4), 1–15. doi.org/10.1145/3444943
Any inquiry concerning this communication or earlier communications from the examiner should be directed to HYUNGJUN B YI whose telephone number is (703)756-4799. The examiner can normally be reached M-F 9-5.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Usmaan Saeed can be reached at (571) 272-4046. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/H.B.Y./Examiner, Art Unit 2146
/USMAAN SAEED/Supervisory Patent Examiner, Art Unit 2146