Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Interpretation
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked.
As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph:
(A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function;
(B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and
(C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function.
Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function.
Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function.
Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action.
This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Each of claims 1-19 have the form “Apparatus … configured to” perform two or more limitations recited in functional language. As such, all the limitations in claims 1-19 are being interpreted under 35 U.S.C. 112(f).
Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof.
If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph.
Response to Arguments
Regarding the objection to claims 1–3 and 5–21 in the previous office action, Applicant’s amendments have clarified the language in question and the objection is withdrawn.
Regarding the rejection of claims under 35 U.S.C. 103, Applicant arguments are directed towards amended claims which have not been previously examined, and for which new grounds of rejection are given below, using a new combination of references.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1–3, 5–8, 10, and 20–21 rejected under 35 U.S.C. 103 over Wang et al., US Pre-Grant Publication No. 2022/0083866 (hereafter Wang) in view of Bach et al., US Pre-Grant Publication No. 2018/0018553 (hereafter Bach).
Regarding claim 1 and analogous claims 20-21:
Wang teaches:
“Apparatus for pruning and/or quantizing of a machine learning (ML) predictor”: Wang, paragraph 0005, “According to a first aspect, there is provided an apparatus [Apparatus] comprising means for performing: training a neural network by applying an optimization loss function, wherein the optimization loss function considers empirical errors and model redundancy; pruning a trained neural network [pruning and/or quantizing of a machine learning (ML) predictor] by removing one or more filters that have insignificant contributions from a set of filters; and providing the pruned neural network for transmission.”
“the apparatus being configured to determine relevance scores for portions of the ML predictor on the basis of an activation of the portions of the ML predictor manifesting itself in one or more inferences performed by the ML predictor”: Wang, paragraph 0057, “As another example, scaling factor based pruning may be applied. The filters of the set of filters may be ranked based on importance scaling factors. For example, a BatchNormalization (BN) based scaling factor may be used to quantify the importance of different filters [determine relevance scores for portions of the ML predictor].”
“prune and/or quantize the ML predictor using the relevance scores”: Wang, paragraph 0057, “The filters may be arranged in descending order of the scaling factor, e.g. the BN-based scaling factor. The filters that are below a threshold percentile p % of the ranked filters may be pruned [prune and/or quantize the ML predictor using the relevance scores].”
“wherein the ML predictor comprises nodes and node interconnections”: Wang, paragraph 0030, “A neural network (NN) is a computation graph comprising several layers of computation. Each layer comprises one or more units, where each unit performs an elementary computation. A unit [nodes] is connected to one or more other units [node interconnections], and the connection may have associated a weight.”
“and the apparatus is configured to determine the relevance scores for the nodes and/or the node interconnections of the ML predictor“: Wang, paragraph 0056, “For example, in diversity based pruning, the filters of the set of filters may be ranked based on column-wise summation of the diversity matrix (1). These summations may be used to quantify the diversity of a given filter with regard to other filters in the set of filters [determine the relevance scores for the nodes and/or the node interconnections].”
Wang does not explicitly teach:
“back propagating, along a reverse direction opposite to an activation propagation direction along which activation are propagated through the ML predictor during the one or more inferences, an initial relevance score at an output of the ML predictor by”
“distributing a relevance score R [J] at a predetermined node [J] of the ML predictor onto predecessor nodes of the predetermined node [J] by,”
“for each predecessor node [i], determining a fraction based on a product a[i] ⋅ w[i][J] between an activation a[i] of the respective predecessor node [i] contributing, in the one or more inferences, to an activation a[J] of the predetermined node [J] by weighing the activation a[J] of the respective predecessor node with a weight w[i][J] of the ML predictor between the respective predecessor node [i] and the predetermined node [J], and the weight w[i][J] of the ML predictor between the respective predecessor node [i] and the predetermined node [J], divided by a sum over addends formed by the products a[i] ⋅ w[i][J] for all predecessor nodes [i]”
“distributing the fraction of the relevance score R[J] at the predetermined node [J] to the respective predecessor node [i]”
Bach teaches:
“back propagating, along a reverse direction opposite to an activation propagation direction along which activation are propagated through the ML predictor during the one or more inferences, an initial relevance score at an output of the ML predictor by”: Bach, paragraph 0123, “As an alternative to Taylor-type decomposition, it is possible to compute relevances at each layer in a backward pass [along a reverse direction opposite to an activation propagation direction along which activation are propagated through the ML predictor during the one or more inferences], that is, express relevances Ri(l) as a function of upper-layer relevances |Rj(l+1)| and backpropagating relevances [back propagating … an initial relevance score at an output of the ML predictor] until we reach the input (pixels).”
“distributing a relevance score RJ at a predetermined node j of the ML predictor onto predecessor nodes of the predetermined node j by”: Bach, paragraph 0058, “Thus, as illustrated in FIG. a, the process of reverse propagation may be thought of as distributing the initial relevance value R, starting from the output neuron(s), towards the input side of the network 10 along the reverse propagation direction 32 [distributing a relevance score RJ at a predetermined node j of the ML predictor onto predecessor nodes of the predetermined node j].”
“for each predecessor node i, determining a fraction based on a product ai ⋅ wij between an activation ai of the respective predecessor node i contributing, in the one or more inferences, to an activation aj of the predetermined node j by weighing the activation aj of the respective predecessor node with a weight wij of the ML predictor between the respective predecessor node i and the predetermined node j, and the weight wij of the ML predictor between the respective predecessor node i and the predetermined node j, divided by a sum over addends formed by the products ai ⋅ wij for all predecessor nodes i”: Bach, paragraph 0027, “FIG. 5 shows a neural network-shaped classifier during prediction time. wij are the connection weights. ai is the activation of neuron i”; Bach, paragraphs 0135-0136, “In addition to the redistribution formulas above, we can define alternative formulas as follows:
PNG
media_image1.png
470
854
media_image1.png
Greyscale
where n is the number of upstream neighbor neurons of the respective neuron, Rij is the relevance value redistributed from the respective neuron j to the upstream neighbor neuron i and Rj is the relevance of neuron j which is a downstream neuron of neuron i, xi is the activation of upstream neighbor neuron i during the application of the neural network [hence, the same activations are used for interference and the backpropagation of relevance], wij is the weight connecting the upstream neighbor neuron i to the respective neuron j, wrj is also a weight connecting the upstream neighbor neuron r to the respective neuron j, and bji is a bias term of the respective neuron i, and h( ) is a scalar function [hence, in equation A6, which is determining a fraction, the numerator xi * wij is a product ai ⋅ wij between an activation ai of the respective predecessor node i contributing, in the one or more inferences, to an activation aj of the predetermined node j by weighing the activation aj of the respective predecessor node with a weight wij of the ML predictor between the respective predecessor node i and the predetermined node j, and the weight wij of the ML predictor between the respective predecessor node i and the predetermined node j, and the denominator includes a summation, over variable r, of the product xr * wrj, thus divided by a sum over addends formed by the products ai ⋅ wij for all predecessor nodes i.”
“distributing the fraction of the relevance score Rj at the predetermined node j to the respective predecessor node i”: Bach, paragraph 0058, “Thus, as illustrated in FIG. a, the process of reverse propagation may be thought of as distributing the initial relevance value R, starting from the output neuron(s), towards the input side of the network 10 along the reverse propagation direction 32 [distributing the fraction of the relevance score Rj at the predetermined node j to the respective predecessor node i].”
Bach and Wang are analogous arts as they are both related to the relevance of neural network nodes. It would have been obvious to a person having ordinary skill in the art prior to the effective filing date of the claimed invention to have combined the relevance distribution of Bach with the model pruning of Wang to arrive at the present invention, in order to more efficiently determine the relevance of individual network nodes, as stated in Bach, paragraph 0018, “In particular, this reverse propagation is applicable to a broader set of artificial neural networks and/or at lower computational efforts by performing same in a manner so that for each neuron, preliminarily redistributed relevance scores of a set of downstream neighbor neurons of the respective neuron are distributed on a set of upstream neighbor neurons of the respective neuron according to a distribution function.”
Regarding claim 2:
Wang as modified by Bach teaches the apparatus of claim 1.
Wang further teaches “configured to use a pruned and/or quantized version of the ML predictor which results from the pruning and/or quantizing, to perform one or more further inferences, and recursively repeat the determining the relevance scores and the pruning and/or quantizing on the basis of an activation of the portions of the ML predictor manifesting itself in the one or more further inferences“: Wang, paragraph 0095, “The apparatus may be further caused to, for minibatches of a training stage [one or more further inferences]: rank filters of the set of filters according to scaling factors; select the filters that are below a threshold percentile of the ranked filters; prune the selected filters temporarily during optimization of one of the minibatches; iteratively repeat the ranking, selecting and pruning for the mini-batches [recursively repeat the determining the relevance scores and the pruning and/or quantizing on the basis of an activation of the portions of the ML predictor manifesting itself in the one or more further inferences].”
Regarding claim 3:
Wang as modified by Bach teaches the apparatus of claim 1.
Wang further teaches “configured to subject a non-pruned-away and/or not quantized to zero portion of the ML predictor which results from pruning and/or quantizing, to training using training data“: Wang, paragraph 0095, “The apparatus may be further caused to, for minibatches of a training stage: rank filters of the set of filters according to scaling factors; select the filters that are below a threshold percentile of the ranked filters; prune the selected filters temporarily during optimization of one of the minibatches; iteratively repeat the ranking, selecting and pruning for the mini-batches [showing pruning occurs during training, hence, a non-pruned-away and/or not quantized to zero portion of the ML predictor which results from pruning and/or quantizing is subject to training using training data].”
Regarding claim 5:
Wang as modified by Bach teaches the apparatus of claim 1.
Wang further teaches “configured to determine the relevance score for a predetermined portion of the ML predictor, composed of more than one node and/or node inter connection of the ML predictor by aggregating the relevance scores of the more than one node and/or node interconnection the predetermined portion is composed of”: Wang, paragraph 0056, “The trained neural network may be pruned by removing one or more filters that have insignificant contribution from a set of filters. There are alternative pruning schemes. For example, in diversity based pruning, the filters of the set of filters may be ranked based on column-wise summation of the diversity matrix (1) [the relevance score for a predetermined portion of the ML predictor, composed of more than one node and/or node inter connection of the ML predictor by aggregating the relevance scores of the more than one node and/or node interconnection the predetermined portion is composed of]. These summations may be used to quantify the diversity of a given filter with regard to other filters in the set of filters.”
Regarding claim 6:
Wang as modified by Bach teaches the apparatus of claim 5.
Wang further teaches “configured to determine the predetermined portion by analyzing the distribution of relevance scores over the ML predictor“: Wang, section 0056, “Wang, paragraph 0056, “The trained neural network may be pruned by removing one or more filters that have insignificant contribution from a set of filters. There are alternative pruning schemes. For example, in diversity based pruning, the filters of the set of filters may be ranked based on column-wise summation of the diversity matrix (1). These summations may be used to quantify the diversity of a given filter with regard to other filters in the set of filters [determine the predetermined portion by analyzing the distribution of relevance scores over the ML predictor]. The filters may be arranged in descending order of the column-wise summations of the diversities. The filters that are below a threshold percentile p % of the ranked filters may be pruned.”
Regarding claim 7:
Wang as modified by Bach teaches the apparatus of claim 1.
Wang further teaches “configured to, in pruning and/or quantizing the ML predictor using the relevance scores, prune away predetermined portions of the ML predictor whose relevance according to the relevance score determined for the predetermined portions is lower than a predetermined threshold”: Wang, paragraph 0057, “The filters may be arranged in descending order of the scaling factor, e.g. the BN-based scaling factor. The filters that are below a threshold percentile p % of the ranked filters may be pruned [according to the relevance score determined for the predetermined portions is lower than a predetermined threshold].”
Regarding claim 8:
Wang as modified by Bach teaches the apparatus of claim 1.
Wang further teaches “configured to, in pruning and/or quantizing the ML predictor using the relevance scores, prune away predetermined first portions of the ML predictor whose relevance according to the relevance score determined for the predetermined nodes fulfills a predetermined criterion, and second portions which contribute to an output of the ML predictor via the first portions exclusively”: Wang, paragraph 0077, “Clearly one can observe based on the line 430 that, once the pruning loss is incorporated into the optimization objective function, i.e. minimization objective function, scaling factors associated with pruned filters are significantly suppressed [hence, the scaling factors of second portions which contribute to an output of the ML predictor via the first portions exclusively are minimized] while scaling factors are enhanced for remaining filters”; Wang, paragraph 0095, “The apparatus may be further caused to, for mini-batches of a training stage: rank filters of the set of filters according to scaling factors; select the filters that are below a threshold percentile of the ranked filters; prune the selected filters temporarily during optimization of one of the mini batches [prune away predetermined first portions of the ML predictor whose relevance according to the relevance score determined for the predetermined nodes fulfills a predetermined criterion]; iteratively repeat the ranking, selecting and pruning for the mini-batches [prune away … second portions which contribute to an output of the ML predictor via the first portions exclusively].”
Regarding claim 10:
Wang as modified by Bach teaches the apparatus of claim 1.
Wang further teaches “configured to, in pruning and/or quantizing the ML predictor using the relevance scores, prune away predetermined portions of the ML predictor whose relevance according to the relevance score determined for the predetermined portions is lower than the relevance of more than a predetermined fraction of portions of the ML predictor“: Wang, section 0056, “The trained neural network may be pruned by removing one or more filters that have insignificant contribution from a set of filters. There are alternative pruning schemes. For example, in diversity based pruning, the filters of the set of filters may be ranked based on column-wise summation of the diversity matrix (1). These summations may be used to quantify the diversity of a given filter with regard to other filters in the set of filters. The filters may be arranged in descending order of the column-wise summations of the diversities. The filters that are below a threshold percentile p % of the ranked filters may be pruned [prune away predetermined portions of the ML predictor whose relevance according to the relevance score determined for the predetermined portions is lower than the relevance of more than a predetermined fraction of portions of the ML predictor].”
Claim 9 rejected under 35 U.S.C. 103 over Wang as modified by Bach in view of He et al., “Channel Pruning for Accelerating Very Deep Neural Networks,” 2017, arXiv:1707.06168 (hereafter He).
Wang as modified by Bach teaches the apparatus of claim 7.
Wang further teaches (bold only) “configured to, in pruning and/or quantizing the ML predictor using the relevance scores, perform the pruning away so that the predetermined threshold decreases towards an output of the ML predictor”: Wang, paragraph 0057, “The filters may be arranged in descending order of the scaling factor, e.g. the BN-based scaling factor. The filters that are below a threshold percentile p % of the ranked filters may be pruned [in pruning and/or quantizing the ML predictor using the relevance scores].”
Wang as modified by Bach does not explicitly teach (bold only) “configured to, in pruning and/or quantizing the ML predictor using the relevance scores, perform the pruning away so that the predetermined threshold decreases towards an output of the ML predictor”
He teaches (bold only) “configured to, in pruning and/or quantizing the ML predictor using the relevance scores, perform the pruning away so that the predetermined threshold decreases towards an output of the ML predictor”: He, section 4.1.1, paragraph 3, “Also notice that channel pruning gradually becomes hard, from shallower to deeper layers. It indicates that shallower layers have much more redundancy, which is consistent with [52]. We could prune more aggressively on shallower layers in whole model acceleration [perform the pruning away so that the predetermined threshold decreases towards an output of the ML predictor].”
He and Wang are both related to the same field of endeavor (pruning of neural networks). Wang teaches the use of a relevance threshold in the determination of pruning model components. He teaches that model components may be more aggressively pruned in shallower layers (i.e., layers further from the output). It would have been obvious to a person having ordinary skill in the art prior to the effective filing date of the claimed invention to have combined the aggressive earlier-layer pruning of He to the threshold-based pruning of Wang to arrive at the present invention, in order to take advantage of shallow layer redundancy to prune more of the model, as stated in He, section 4.1.1, paragraph 3, “Also notice that channel pruning gradually becomes hard, from shallower to deeper layers. It indicates that shallower layers have much more redundancy, which is consistent with [52]. We could prune more aggressively on shallower layers in whole model acceleration.”
Claims 11–15 and 17 rejected under 35 U.S.C. 103 over Wang as modified by Bach in view of Choi et al., US Pre-Grant Publication No. 2018/0107926 (hereafter Choi).
Regarding claim 11:
Wang as modified by Bach teaches the apparatus of claim 1.
Wang further teaches (bold only) “configured to prune and/or quantize the ML predictor using an optimization scheme with an objective function which depends on a weighted distance between quantized weights and unquantized weights of the ML predictor, weighted based on the relevance scores”: Wang, paragraph 0048, “In the method disclosed herein, the optimization loss function [an optimization scheme], i.e. the objective function [an objective function] of filter diversity enhanced NN learning may be formulated by:
PNG
media_image2.png
34
220
media_image2.png
Greyscale
wherein λ is the parameter to control relative significance of the original task and the filter diversity enhancement term KΘ, and Θ is the parameter to measure filter diversities used in function K [weighted based on the relevance scores]. W* above represents the first loss function.”
Wang as modified by Bach does not explicitly teach (bold only) “configured to prune and/or quantize the ML predictor using an optimization scheme with an objective function which depends on a weighted distance between quantized weights and unquantized weights of the ML predictor, weighted based on the relevance scores.”
Choi teaches (bold only) “configured to prune and/or quantize the ML predictor using an optimization scheme with an objective function which depends on a weighted distance between quantized weights and unquantized weights of the ML predictor, weighted based on the relevance scores”: Choi, paragraph 0077, “In order to, inter alia, address these issues with conventional k-means clustering, the present disclosure describes (1) utilizing the second-order partial derivatives, i.e., the diagonal of the Hessian matrix, of the loss function with respect to the network parameters as a measure of the significance of different network parameters [weighted based on the relevance scores]; and (2) solving the network quantization problem under a constraint [an optimization scheme with an objective function] of the actual compression ratio resulting from the specific binary encoding scheme that was employed [which depends on a weighted distance between quantized weights and unquantized weights of the ML predictor]. Accordingly, the description below is broken into three sections: I. Network Quantization [configured to prune and/or quantize the ML predictor] using Hessian-Weight; II. Entropy-Constrained Network Quantization; and III. Experimental/Simulation Results.”
Choi and Wang are both related to the same field of endeavor (compression of neural networks). Wang teaches a neural network reduction method using relevancy-based pruning. Choi teaches a neural network quantization method using of a k-means clustering algorithm incorporating relevancy. It would have been obvious to a person having ordinary skill in the art prior to the effective filing date of the claimed invention to have combined the quantization scheme of Choi to the pruning teachings of Wang to arrive at the present invention, in order to further reduce the complexity of a model, as stated in Choi, paragraphs 0006-0008, “As mentioned above, there may be hundreds of millions of network parameters/weights which require a significant amount of memory to be stored. Accordingly, although deep neural networks are extremely powerful, they also require a significant amount of resources to implement, particularly in terms of memory storage. […] This makes it difficult to deploy deep neural networks on devices with limited storage, such as mobile/portable devices. Accordingly, the present disclosure has been made to address at least the problems and/or disadvantages described herein and to provide at least the advantages described below.”
Regarding claim 12:
Wang as modified by Bach and Choi teaches the apparatus of claim 11.
Choi further teaches “wherein the objective function depends on a sum of the weighted distance and a code length of the quantized weights”: Choi, paragraphs 0072-0073: “Assuming k-means clustering is used with variable-length encoding, let Ci be the number of network parameters in cluster i, where 1<=i<=k, and bi is the number of bits of the codeword assigned for the network parameters in cluster i (i.e., the codeword for cluster i). Thus, the binary codewords are only Σi=1k |Ci|bi bits rather than Nb bits. For a lookup table that stores the k binary codewords (bi bits for 1<=i<=k) with their matching quantized network parameter values (b bits each), an additional Σi=1k bi+kb is needed. Thus, in Equation (6)(a):
PNG
media_image3.png
75
408
media_image3.png
Greyscale
. As seen above, in k-means clustering with variable-length encoding, the compression ratio depends not only on the number of clusters, i.e., k, but also on the sizes of the various clusters, i.e., |Ci|, and the assigned number of bits for each cluster codeword, i.e., bi for 1<=i<=k [hence, the distance between codewords, and therefore the quantization error, depends on a code length of the quantized weights]”; Choi, paragraph 0093, “Next, Equation (7)(c) is connected to the problem of network quantization by treating δw`i as the quantization error of network parameter wi at its local optimum wi= w’i, i.e., δw’i=!wi-w`i, , wheres !wi is a quantized value of w’i. Thus, the local impact of quantization on the average loss function at w=w` can be quantified approximately as Equation (7)(d) [the objective function depends on a sum of the weighted distance].”
Choi and Wang are combinable for the rationale given under claim 11.
Regarding claim 13:
Wang as modified by Bach and Choi teaches the apparatus of claim 11.
Choi further teaches “wherein the optimization scheme is a k-means clustering“: Choi, paragraph 0077, “In order to, inter alia, address these issues with conventional k-means clustering [wherein the optimization scheme is a k-means clustering], the present disclosure describes (1) utilizing the second-order partial derivatives, i.e., the diagonal of the Hessian matrix, of the loss function with respect to the network parameters as a measure of the significance of different network parameters; and (2) solving the network quantization problem under a constraint of the actual compression ratio resulting from the specific binary encoding scheme that was employed. Accordingly, the description below is broken into three sections: I. Network Quantization using Hessian-Weight; II. Entropy-Constrained Network Quantization; and III. Experimental/Simulation Results.”
Choi and Wang are combinable for the rationale given under claim 11.
Regarding claim 14:
Wang as modified by Bach and Choi teaches the apparatus of claim 13.
Choi further teaches:
“for each iteration of the k-means clustering, a cluster formation step associating each ML predictor node interconnection with one of a plurality of quantization values so as to reduce the optimization function”: Choi, paragraph 0104, “At 620, each network parameter is assigned to the cluster whose center is closest to the parameter value. This may be done by partitioning the points according to the Voronoi diagram generated by the cluster centers [a cluster formation step associating each ML predictor node interconnection with one of a plurality of quantization values so as to reduce the optimization function]. At 630, after all of the network parameters have been (re-)assigned in 620, a new set of Hessian-weighted cluster center means are calculated. At 640, the number of iterations is updated ("n= n+ 1 ") and, at 645, it is determined whether the iterative process has ended [for each iteration of the k-means clustering].”
“and a quantizer update step updating each quantization value using a weighted sum of the unquantized weights of the predictor node interconnection associated with the respective quantization value, weighted with a relevance score determined for the respective ML predictor node interconnection”: Choi, paragraph 0086, “In this section, the impact of quantization errors on the loss function of a neural network is analyzed and a Hessian-weight that can be used to quantify the significance of different network parameters [a relevance score] in quantization is derived”; Choi, paragraphs 0097-0100, From Equation (7)( d), the optimal clustering that minimizes the Hessian-weighted distortion measure is given by Equation(8)(a):
PNG
media_image4.png
57
372
media_image4.png
Greyscale
where hii is the second-order partial derivative of the loss function with respect to network parameter wi as shown in Equation (8)(b):
PNG
media_image5.png
62
366
media_image5.png
Greyscale
and cj is the Hessian-weighted [weighted with a relevance score] mean of network parameters in cluster CJ as shown in Equation (8)(c) [using a weighted sum of the unquantized weights of the predictor node interconnection associated with the respective quantization value]:
PNG
media_image6.png
69
375
media_image6.png
Greyscale
.
Choi, paragraph 0104, “At 620, each network parameter is assigned to the cluster whose center is closest to the parameter value. This may be done by partitioning the points according to the Voronoi diagram generated by the cluster centers. At 630, after all of the network parameters have been (re-)assigned in 620, a new set of Hessian-weighted cluster center means are calculated [and a quantizer update step updating each quantization value]. At 640, the number of iterations is updated ("n= n+ 1 ") and, at 645, it is determined whether the iterative process has ended.”
Choi and Wang are combinable for the rationale given under claim 11.
Regarding claim 15:
Wang as modified by Bach and Choi teaches the apparatus of claim 13.
Choi further teaches “configured to repeat performing the k-means clustering with, after each performance, accepting the quantized weights for predictor node interconnections for which the acceptance increases the optimization function less than a predetermined threshold or less than a predetermined fraction of remaining unquantized weights“: Choi, paragraph 0104, “At 620, each network parameter is assigned to the cluster whose center is closest to the parameter value. This may be done by partitioning the points according to the Voronoi diagram generated by the cluster centers. At 630, after all of the network parameters have been (re-)assigned in 620, a new set of Hessian-weighted cluster center means are calculated. At 640, the number of iterations is updated ("n= n+1 ") and, at 645, it is determined whether the iterative process has ended. More specifically, it is determined whether the number of iterations has exceeded a limit ("n>N") or there was no change in assignment at 620 (which means the algorithm has effectively converged) [after each performance, accepting the quantized weights for predictor node interconnections for which the acceptance increases the optimization function less than a predetermined threshold or less than a predetermined fraction of remaining unquantized weights]. If it is determined that the process has ended in 645, a codebook for quantized network parameters is generated at 650. If it is determined that the process has not ended in 645, the process repeats to assignment at 620.”
Choi and Wang are combinable for the rationale given under claim 11.
Regarding claim 17:
Wang as modified by Bach and Choi teaches the apparatus of claim 14.
Wang further teaches (bold only) “configured to re-train the ML predictor with respect to unquantized weights for which no quantized weight has yet been accepted”: Wang, paragraph 0066, “The method may comprise estimating accuracy of the network after pruning. For example, the accuracy of the image classification may be estimated using a known dataset. If the accuracy is below a threshold accuracy, the method may comprise retraining the pruned network. Then the accuracy may be estimated again, and the retraining may be repeated until the threshold accuracy is achieved [configured to re-train the ML predictor with respect to unquantized weights for which no quantized weight has yet been accepted].”
Claim 16 rejected under 35 U.S.C. 103 over Wang as modified by Bach and Choi in view of Zeng et al., “Compressing Deep Neural Network for Facial Landmarks Detection,” 2016, Advances in Brain Inspired Cognitive Systems (hereafter Zeng).
Wang as modified by Bach and Choi teaches the apparatus of claim 15.
Wang further teaches (bold only) “configured to re-train the ML predictor with respect to unquantized weights for which no quantized weight has yet been accepted before any repetition of the k-means clustering”: Wang, paragraph 0066, “The method may comprise estimating accuracy of the network after pruning. For example, the accuracy of the image classification may be estimated using a known dataset. If the accuracy is below a threshold accuracy, the method may comprise retraining the pruned network. Then the accuracy may be estimated again, and the retraining may
be repeated until the threshold accuracy is achieved [configured to re-train the ML predictor with respect to unquantized weights for which no quantized weight has yet been accepted].”
Choi further teaches (bold only) “configured to re-train the ML predictor with respect to unquantized weights for which no quantized weight has yet been accepted before any repetition of the k-means clustering”: Choi, paragraph 0104, “At 620, each network parameter is assigned to the cluster whose center is closest to the parameter value. This may be done by partitioning the points according to the Voronoi diagram generated by the cluster centers. At 630, after all of the network parameters have been (re-)assigned in 620, a new set of Hessian-weighted cluster center means are calculated. At 640, the number of iterations is updated ("n= n+1 ") and, at 645, it is determined whether the iterative process has ended. More specifically, it is determined whether the number of iterations has exceeded a limit ("n>N") or there was no change in assignment at 620 (which means the algorithm has effectively converged). If it is determined that the process has ended in 645, a codebook for quantized network parameters is generated at 650. If it is determined that the process has not ended in 645, the process repeats to assignment at 620 [any repetition of the k-means clustering].”
Choi and Wang are combinable for the rationale given under claim 11.
Wang as modified by Bach and Choi does not explicitly teach (bold only) “configured to re-train the ML predictor with respect to unquantized weights for which no quantized weight has yet been accepted before any repetition of the k-means clustering”
Zeng teaches (bold only) “configured to re-train the ML predictor with respect to unquantized weights for which no quantized weight has yet been accepted before any repetition of the k-means clustering”: Zeng, section 3.3, “We compress the network layer by layer from backward to forward. For every iteration, we use the previously trained model to calculate correlated coefficients, and then prune and quantify one additional layer [hence, after the first iteration, the network is a mix of quantized and unquantized weights]. To retrain the network, we use deep learning tools Caffe [23] and simply modify its convolutional and fully connected layers by adding another two blobs to store index and codes [re-train the ML predictor with respect to unquantized weights for which no quantized weight has yet been accepted before any repetition]. Each time before forward-propagation, we use index and codes to reconstruct the weights.”
Zeng and Wang as modified by Choi are both related to the same field of endeavor (compression of neural networks). Wang as modified by Choi teaches a method of reducing model complexity through an iterative k-means clustering-based quantization. Zeng teaches retraining the model during iteration of quantization. It would have been obvious to a person having ordinary skill in the art prior to the effective filing date of the claimed invention to have combined the model retraining of Zeng to the teachings of Wang as modified by Choi to arrive at the present invention, in order to reduce degradation of the network caused by the compression, as stated in Zeng, Abstract, “Network retraining: to reduce training difficulty and performance degradation, we iteratively retrain the network, compressing one layer at a time.”
Claim 18 rejected under 35 U.S.C. 103 over Wang as modified by Bach in view of Sivaraman et al., US Pre-Grant Publication No. 2020/0342324 (hereafter Sivaraman).
Wang as modified by Bach teaches the apparatus of claim 1.
Wang further teaches (bold only) “replace the ML predictor with a pruned and/or quantized version of the ML predictor which results from the pruning and/or quantizing and apply the pruned and/or quantized version of the ML predictor onto further input data such as replenishments of the local input data to subject the further input data to inference”: Wang, paragraph 0066, “The method may comprise estimating accuracy of the network after pruning. For example, the accuracy of the image classification may be estimated using a known data set. If the accuracy is below a threshold accuracy, the method may comprise retraining the pruned network [a pruned and/or quantized version of the ML predictor which results from the pruning and/or quantizing and apply the pruned and/or quantized version of the ML predictor onto further input data]. Then the accuracy may be estimated again, and the retraining may be repeated until the threshold accuracy is achieved.”
Wang as modified by Bach does not explicitly teach:
“configured to retrieve a definition of the ML predictor from a server”
“apply the ML predictor onto local input data so as to make the ML predictor performing the one or more inferences“
(bold only) “replace the ML predictor with a pruned and/or quantized version of the ML predictor which results from the pruning and/or quantizing and apply the pruned and/or quantized version of the ML predictor onto further input data such as replenishments of the local input data to subject the further input data to inference”
Sivaraman teaches:
“configured to retrieve a definition of the ML predictor from a server”: Sivaraman, paragraph 0039, “As a result of the relatively limited computing power of the edge nodes 12 and as discussed in more detail below, the ANNs 24 in the disclosed embodiments are each a pruned version of a parent ANN 32 included by the centralized server 14 [configured to retrieve a definition of the ML predictor from a server]. To this end, the centralized server 14 may include a pruning application that is arranged for commissioning a pruned ANN to each of the edge nodes 12 and the ANNs 24 may be referred to as ‘pruned’. The parent ANN 32 can take any of the forms noted above, such as a CNN, or ANN according to another deep learning model that is capable of detecting and classifying objects in images.”
“apply the ML predictor onto local input data so as to make the ML predictor performing the one or more inferences“: Sivaraman, paragraph 0028, “In view of the foregoing, various embodiments and implementations are directed to systems and methods for enhancing pedestrian safety and/or conducting surveillance via object recognition using pruned neural networks, particularly on resource-constrained edge nodes of communication networks. The pruned neural networks are created by a centralized server of the communication network to be customized for each of the edge nodes on the server. The creation of the pruned neural networks includes identifying a subject of objects (or classes) of a parent neural network and selecting the filters that respond highly to those objects. The subset of objects is determined by analyzing reference images actually captured at each edge node, such that the pruned neural network is customized specifically for each particular edge node [apply the ML predictor onto local input data so as to make the ML predictor performing the one or more inferences]. Furthermore, analysis of the reference images occurs by the parent neural network, which is unpruned, fully trained, and comprehensive. Ultimately, this enables a smaller neural network to be built filter-by-filter from the results of the parent neural network.”
“replace the ML predictor with a pruned and/or quantized version of the ML predictor which results from the pruning and/or quantizing and apply the pruned and/or quantized version of the ML predictor onto further input data such as replenishments of the local input data to subject the further input data to inference”: Sivaraman, paragraph 0028, “In view of the foregoing, various embodiments and implementations are directed to systems and methods for enhancing pedestrian safety and/or conducting surveillance via object recognition using pruned neural networks, particularly on resource-constrained edge nodes of communication networks. The pruned neural networks are created by a centralized server of the communication network to be customized for each of the edge nodes on the server. The creation of the pruned neural networks includes identifying a subject of objects (or classes) of a parent neural network and selecting the filters that respond highly to those objects. The subset of objects is determined by analyzing reference images actually captured at each edge node, such that the pruned neural network is customized specifically for each particular edge node. Furthermore, analysis of the reference images occurs by the parent neural network, which is unpruned, fully trained, and comprehensive. Ultimately, this enables a smaller neural network to be built filter-by-filter from the results of the parent neural network”; Sivaraman, paragraph 0029, “By deploying the pruned neural network to each edge node that includes only the subset of objects which are relevant to each specific edge node, low-capacity or resource-constrained computing devices ( e.g., consumer electronics) can detect certain relevant objects with very high accuracy at fast speeds with minimal resources [replace the ML predictor with a pruned and/or quantized version of the ML predictor which results from the pruning and/or quantizing and apply the pruned and/or quantized version of the ML predictor onto further input data such as replenishments of the local input data to subject the further input data to inference].”
Sivaraman and Wang are both related to the same field of endeavor (neural network pruning). It would have been obvious to a person having ordinary skill in the art prior to the effective filing date of the claimed invention to have combined the server-based pruning techniques of Sivaraman to the pruning teachings of Wang to arrive at the present invention, in order to create pruned networks specialized to local data, as stated in Sivaraman, paragraph 0029, “By deploying the pruned neural network to each edge node that includes only the subset of objects which are relevant to each specific edge node, low-capacity or resource-constrained computing devices ( e.g., consumer electronics) can detect certain relevant objects with very high accuracy at fast speeds with minimal resources.”
Claim 19 rejected under 35 U.S.C. 103 over Wang in view of Bach and Sivaraman.
Wang teaches:
“Apparatus for sub-model extraction”: Wang, paragraph 0005, “According to a first aspect, there is provided an apparatus [Apparatus] comprising means for performing: training a neural network by applying an optimization loss function, wherein the optimization loss function considers empirical errors and model redundancy; pruning a trained neural network [sub-model extraction] by removing one or more filters that have insignificant contributions from a set of filters; and providing the pruned neural network for transmission.”
“remove portions of the general ML predictor exclusively interconnected to one or more predetermined uninterested outputs of the ML predictor to acquire a ML predictor”: Wang, paragraph 0057, “As another example, scaling factor based pruning may be applied. The filters of the set of filters may be ranked based on importance scaling factors. For example, a BatchNormalization (BN) based scaling factor may be used to quantify the importance of different filters. The scaling factor may be obtained from e.g. batch-normalization or additional scaling layer. The filters may be arranged in descending order of the scaling factor, e.g. the BN-based scaling factor. The filters that are below a threshold percentile p % of the ranked filters may be pruned [remove portions of the general ML predictor exclusively interconnected to one or more predetermined uninterested outputs of the ML predictor to acquire a ML predictor].”
“prune and/or quantize the ML predictor by: determining relevance scores for portions of the ML predictor on the basis of an activation of the portions of the ML predictor manifesting itself in one or more inferences performed by the ML predictor”: Wang, paragraph 0057, “As another example, scaling factor based pruning may be applied. The filters of the set of filters may be ranked based on importance scaling factors. For example, a BatchNormalization (BN) based scaling factor may be used to quantify the importance of different filters [determining relevance scores for portions of the ML predictor].”
“pruning and/or quantizing the ML predictor using the relevance scores”: Wang, paragraph 0057, “The filters may be arranged in descending order of the scaling factor, e.g. the BN-based scaling factor. The filters that are below a threshold percentile p % of the ranked filters may be pruned [pruning and/or quantizing the ML predictor using the relevance scores].”
“wherein the ML predictor comprises nodes and node interconnections”: Wang, paragraph 0030, “A neural network (NN) is a computation graph comprising several layers of computation. Each layer comprises one or more units, where each unit performs an elementary computation. A unit [nodes] is connected to one or more other units [node interconnections], and the connection may have associated a weight.”
“and the apparatus is configured to determine the relevance scores for the nodes and/or the node interconnections of the ML predictor“: Wang, paragraph 0056, “For example, in diversity based pruning, the filters of the set of filters may be ranked based on column-wise summation of the diversity matrix (1). These summations may be used to quantify the diversity of a given filter with regard to other filters in the set of filters [determine the relevance scores for the nodes and/or the node interconnections].”
Wang does not explicitly teach:
“configured to retrieve a definition of a general ML predictor from a server”
“back propagating, along a reverse direction opposite to an activation propagation direction along which activation are propagated through the ML predictor during the one or more inferences, an initial relevance score at an output of the ML predictor by”
“distributing a relevance score R [J] at a predetermined node [J] of the ML predictor onto predecessor nodes of the predetermined node [J] by,”
“for each predecessor node [i], determining a fraction based on a product a[i] ⋅ w[i][J] between an activation a[i] of the respective predecessor node [i] contributing, in the one or more inferences, to an activation a[J] of the predetermined node [J] by weighing the activation a[J] of the respective predecessor node with a weight w[i][J] of the ML predictor between the respective predecessor node [i] and the predetermined node [J], and the weight w[i][J] of the ML predictor between the respective predecessor node [i] and the predetermined node [J], divided by a sum over addends formed by the products a[i] ⋅ w[i][J] for all predecessor nodes [i]”
“distributing the fraction of the relevance score R[J] at the predetermined node [J] to the respective predecessor node [i]”
Sivaraman teaches “configured to retrieve a definition of a general machine learning, ML, predictor from a server”: Sivaraman, paragraph 0039, “As a result of the relatively limited computing power of the edge nodes 12 and as discussed in more detail below, the ANNs 24 in the disclosed embodiments are each a pruned version of a parent ANN 32 included by the centralized server 14 [configured to retrieve a definition of a general machine learning, ML, predictor from a server]. To this end, the centralized server 14 may include a pruning application that is arranged for commissioning a pruned ANN to each of the edge nodes 12 and the ANNs 24 may be referred to as ‘pruned’. The parent ANN 32 can take any of the forms noted above, such as a CNN, or ANN according to another deep learning model that is capable of detecting and classifying objects in images.”
Sivaraman and Wang are both related to the same field of endeavor (neural network pruning). It would have been obvious to a person having ordinary skill in the art prior to the effective filing date of the claimed invention to have combined the server-based pruning techniques of Sivaraman to the pruning teachings of Wang to arrive at the present invention, in order to create pruned networks specialized to local data, as stated in Sivaraman, paragraph 0029, “By deploying the pruned neural network to each edge node that includes only the subset of objects which are relevant to each specific edge node, low-capacity or resource-constrained computing devices ( e.g., consumer electronics) can detect certain relevant objects with very high accuracy at fast speeds with minimal resources.”
Bach teaches:
“back propagating, along a reverse direction opposite to an activation propagation direction along which activation are propagated through the ML predictor during the one or more inferences, an initial relevance score at an output of the ML predictor by”: Bach, paragraph 0123, “As an alternative to Taylor-type decomposition, it is possible to compute relevances at each layer in a backward pass [along a reverse direction opposite to an activation propagation direction along which activation are propagated through the ML predictor during the one or more inferences], that is, express relevances Ri(l) as a function of upper-layer relevances |Rj(l+1)| and backpropagating relevances [back propagating … an initial relevance score at an output of the ML predictor] until we reach the input (pixels).”
“distributing a relevance score RJ at a predetermined node j of the ML predictor onto predecessor nodes of the predetermined node j by”: Bach, paragraph 0058, “Thus, as illustrated in FIG. a, the process of reverse propagation may be thought of as distributing the initial relevance value R, starting from the output neuron(s), towards the input side of the network 10 along the reverse propagation direction 32 [distributing a relevance score RJ at a predetermined node j of the ML predictor onto predecessor nodes of the predetermined node j].”
“for each predecessor node i, determining a fraction based on a product ai ⋅ wij between an activation ai of the respective predecessor node i contributing, in the one or more inferences, to an activation aj of the predetermined node j by weighing the activation aj of the respective predecessor node with a weight wij of the ML predictor between the respective predecessor node i and the predetermined node j, and the weight wij of the ML predictor between the respective predecessor node i and the predetermined node j, divided by a sum over addends formed by the products ai ⋅ wij for all predecessor nodes i”: Bach, paragraph 0027, “FIG. 5 shows a neural network-shaped classifier during prediction time. wij are the connection weights. ai is the activation of neuron i”; Bach, paragraphs 0135-0136, “In addition to the redistribution formulas above, we can define alternative formulas as follows:
PNG
media_image1.png
470
854
media_image1.png
Greyscale
where n is the number of upstream neighbor neurons of the respective neuron, Rij is the relevance value redistributed from the respective neuron j to the upstream neighbor neuron i and Rj is the relevance of neuron j which is a downstream neuron of neuron i, xi is the activation of upstream neighbor neuron i during the application of the neural network [hence, the same activations are used for interference and the backpropagation of relevance], wij is the weight connecting the upstream neighbor neuron i to the respective neuron j, wrj is also a weight connecting the upstream neighbor neuron r to the respective neuron j, and bji is a bias term of the respective neuron i, and h( ) is a scalar function [hence, in equation A6, which is determining a fraction, the numerator xi * wij is a product ai ⋅ wij between an activation ai of the respective predecessor node i contributing, in the one or more inferences, to an activation aj of the predetermined node j by weighing the activation aj of the respective predecessor node with a weight wij of the ML predictor between the respective predecessor node i and the predetermined node j, and the weight wij of the ML predictor between the respective predecessor node i and the predetermined node j, and the denominator includes a summation, over variable r, of the product xr * wrj, thus divided by a sum over addends formed by the products ai ⋅ wij for all predecessor nodes i.”
“distributing the fraction of the relevance score Rj at the predetermined node j to the respective predecessor node i”: Bach, paragraph 0058, “Thus, as illustrated in FIG. a, the process of reverse propagation may be thought of as distributing the initial relevance value R, starting from the output neuron(s), towards the input side of the network 10 along the reverse propagation direction 32 [distributing the fraction of the relevance score Rj at the predetermined node j to the respective predecessor node i].”
Bach and Wang are analogous arts as they are both related to the relevance of neural network nodes. It would have been obvious to a person having ordinary skill in the art prior to the effective filing date of the claimed invention to have combined the relevance distribution of Bach with the model pruning of Wang to arrive at the present invention, in order to more efficiently determine the relevance of individual network nodes, as stated in Bach, paragraph 0018, “In particular, this reverse propagation is applicable to a broader set of artificial neural networks and/or at lower computational efforts by performing same in a manner so that for each neuron, preliminarily redistributed relevance scores of a set of downstream neighbor neurons of the respective neuron are distributed on a set of upstream neighbor neurons of the respective neuron according to a distribution function.”
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Yao et al., US Pre-Grant Publication No. 2020/0334537, discloses a method of pruning neural network parameters based on importance measurements.
Evci, Utku, “Detecting Dead Weights and Units in Neural Networks,” 2018, arXiv:1806.06068v1, discloses a method for network pruning that includes (see chapter 4) a measure of importance that is calculated during a backpropagation pass.
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to VINCENT SPRAUL whose telephone number is (703) 756-1511. The examiner can normally be reached M-F 9:00 am - 5:00 pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, MICHAEL HUNTLEY can be reached at (303) 297-4307. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/VAS/
Examiner, Art Unit 2129
/MICHAEL J HUNTLEY/Supervisory Patent Examiner, Art Unit 2129