Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 6/26/2026 has been entered.
Remarks
This Office Action is responsive to Applicants' Amendment filed on June 26, 2026, in which claims 1 and 9 are currently amended. Claims 1, 2, 4-6, 8-10, 12-14, and 16 are currently pending.
Response to Arguments
Applicant’s arguments with respect to rejection of claims 1, 2, 4-6, 8-10, 12-14, and 16 under 35 U.S.C. 103 based on amendment have been considered, however, are not persuasive.
With respect to Applicant’s arguments on pp. 7-8 of the Remarks submitted 6/26/2026 that the combination of Song and Bar does not disclose “both the at least one original feature map and the at least perturbed feature map are generated by the first artificial neural network model that is not quantized, and the respective importance value is determined independently of the second artificial neural network model,” Examiner respectfully disagrees.
Song does derive a combined output loss expression that includes both adversarial input perturbation and quantization induced weight perturbation. In that expression the perturbed/quantized output is compared against the original full-precision output and the resulting difference includes terms involving both dxl and dWl. To that extent, the combined loss reflects both adversarial and quantization effects. But that is not the only relevant disclosure. Song separately identifies the adversarial-loss component generated by the non-quantized full-precision model. Specifically, Song decomposes the output-loss expression and isolates the adversarial term Wldxl. That term is produced using the full-precision weight matrix Wl, not the quantized weight matrix. Song then explains that Wl acts as the amplification factor for adversarial error and introduces the p-norm/Lipschitz constant framework to measure that amplification. The proper mapping does not require using Song’s combined adversarial and quantization loss as the claimed “importance value.” The claimed importance value can instead be mapped to Song’s element-wise adversarial sensitivity component, i.e. the i-th element of the adversarial output difference caused by dxl propagating through Wl. That value is determined from the non-quantized model alone because it depends on Wl and dxl, not on dWl or the quantized model. The quantized model enters later when Song evaluates quantization using the Lipschitz/p-norm metric and compares quantization choices.
With respect to Applicant’s arguments on pp. 9-10 of the Remarks submitted 6/26/2026 that the combination of Song and Bar does not disclose “calculating a distance between a result of applying the respective importance value to each of the at least one original feature map and a result of applying the respective importance value to a quantized feature map, among the at least one quantized feature map, corresponding to each of the at least one original feature map,” Examiner respectfully disagrees. Applicant’s arguments are directly entirely towards an exemplary embodiment of the specification ([¶0055] “For example, a distance […] may be calculated using”) and do not reflect the scope of the instant claims. Song explicitly defines the Lipschitz constant as ||dW||p with distance dW = Wq-W where Wq corresponds to the quantized feature map of the second artificial neural network and W corresponds to the feature map of the original neural network and is calculated based on the element wise importance. Song's distance is the p-norm distance used to bound output error caused by quantization. The comparison between full-precision and quantized feature-map behavior is captured by the quantized weight output deviation from the full-precision output.
For at least these reasons and those further detailed below, Examiner asserts that it is reasonable and appropriate to maintain the rejection under 35 U.S.C. §103 in view of the combination of Song and Bar.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 1, 2, 4-6, 8-10, 12-14, and 16 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Regarding claims 1 and 9, "a result of applying the respective importance value to each of the at least one original feature map" is indefinite. Claim 1 first recites "determining a respective importance value for each element of the at least one original feature map" and then later applies "the respective importance value to each of the at least one original feature map". Respective appears distributive here such that there appears to be a plurality of importance values, and then no antecedent basis for the later recited "the respective importance value". Even if interpreted as singular, the claim would be ambiguous because the claim later refers to "the respective importance value" in a context where multiple feature maps are in play, without identifying which member of that family is "respectively" being used. This affects the scope of the calculation itself, not just the breadth of possible implementations. In other words, the antecedent basis is not broken merely because "the respective importance value" is singular, but the claim is indefinite because the antecedent term is defined at an element level and then later used at a feature-map level without specifying the required correspondence.
The remaining claims are rejected with respect to their dependence on the rejected claims.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 1, 2, 4, 5, 8, 9, 10, 12, 13, and 16 are rejected under U.S.C. §103 as being unpatentable over the combination of Song (“A Layer-wise Adversarial-aware Quantization Optimization for Improving Robustness”, 2021) and Bar (“A Spectral Perspective of DNN Robustness to Label Noise”, 2022).
Regarding claim 1, Song teaches A method for evaluating a quantized artificial neural network, performed by a computing device having one or more processors and a memory for storing one or more programs executed by the one or more processors, the method comprising: ([p. 3 §2.2] "Neural network quantization [6] helps save computation and memory costs and therefore improves power efficiency. Quantization has also been shown to improve accuracy in some cases, and in almost all cases retains the accuracy of the original network. Moreover, quantization enables the deployment of neural networks on limited-precision hardware. Many hardware systems have physical constraints and do not allow full-precision models to directly map to them. For example, GPUs support half precision floating point arithmetic (FP16); ReRAM (a.k.a memristor) [33] only allows limited precision because of process variations [12, 21]")
generating at least one original feature map for input data using a first artificial neural network model; ([p. 4 §3.1] "suppose W is the weight matrix, W + ∆W is the weight after quantization, x is the original input, and x+∆x is the adversarial input. The difference in the output of this layer (δ) can be represented as follow: [See Eqn. 1]" [p. 5 §3.2] "Cisse et al. [4] has already given detailed explanations on how to represent convolution operations as basic matrix multiplications and we have verified the correctness of their derivations" Song explicitly anticipates W representing convolution kernel such that W is interpreted as a convolutional filter producing at least one original feature map for input data (explicitly al(Wlxl)=x^(l+1)) in Eqn. 4).)
generating at least one piece of modified data by applying a predetermined perturbation to said input data; ([p. 2 §2.1] "An adversarial example Xe is generated by injecting adversarial perturbation ε (a.k.a. adversarial strength) to a clean sample X: Xe = X + ε. Usually, adversarial perturbations are so tiny that they are even imperceptible to human eyes. However, carefully designed adversarial perturbations can cause a neural network to misclassify adversarial examples with high confidence levels" See also dxl in Eqn. 3)
generating at least one perturbed feature map for each piece of modified data using the first artificial neural network model;([p. 4 §3.1] "suppose W is the weight matrix, W + ∆W is the weight after quantization, x is the original input, and x+∆x is the adversarial input. The difference in the output of this layer (δ) can be represented as follow: [See Eqn. 1]" [p. 5 §3.2] "Cisse et al. [4] has already given detailed explanations on how to represent convolution operations as basic matrix multiplications and we have verified the correctness of their derivations" [p. 6 §4.1] "the output of layer l is" See Eqn. 3 where dxl is explicitly the layer input perturbation.)
determining a respective importance value for each element of the at least one original feature map by computing a sample variance or p-norm of differences between each element of the at least one original feature map and a corresponding element of the at least one perturbed feature map([p. 6 §4.1] "where a l (·) is the element-wise activation function of layer l [...] The i-th element of ∆x l+1 a,q" [p. 7] "Since each element of ∆xl+1 a,q is proportional to the corresponding element of (Wl∆xl + ∆Wlxl +∆Wl∆xl), we have [Eqn. 7]" [p. 2] "We [...] derive the metrics for error sensitivity" Each element of the original feature map receives a corresponding sensitivity value. Song's importance value can be read as the element-wise adversarial sensitivity caused by Wldxl which is determined before considering the quantized second model because it uses the full-precision weight matrix. In other words, Song calculates element wise loss having both separate adversarial and quantization components, each of these components is explicitly used for sensitivity analysis)
wherein both the at least one original feature map and the at least one perturbed feature map are generated by the first artificial neural network model that is not quantized, ([p. 6] "The full-precision weight matrix and the quantized weight matrix of layer l are Wl and Wl +∆Wl, respectively (both are m×n matrices)" For adversarial loss sensitivity, Song uses the full precision weight matrix Wl not Wq. Thus, the original and perturbed full precision feature-map comparison is supported when the importance values are read from Song's adversarial loss element analysis)
and the respective importance value is determined independently of a second artificial neural network model;([p. 7] "Here, we define the adversarial loss ∆xl+1 a =Wl∆xl […] Wl will be the amplification factor that enlarges the error. An effective and efficient way to measure the adversarial loss amplification effect is the Lipschitz constant of Wl" Song uses only Wl, the full-precision layer, and not the quantized weight. Song then says Wl is "the amplification factor that enlarges the error," and that the Lipschitz constant of Wl is an "effective and efficient way" to measure the adversarial-loss amplification)
generating at least one quantized feature map for the input data using a second artificial neural network model that is a quantized artificial neural network model for the first artificial neural network model; ([p. 4 §3.1] "W + ∆W is the weight after quantization [...] δ = (W +∆W)·(x+∆x)−W x = W∆x+∆W x+∆W∆x" Weight filter (W + ∆W) interpreted as filter of a second artificial neural network generating at least one quantized feature map for the first artificial neural network model having weight filter W.)
and calculating an evaluation value for the second artificial neural network model based on the at least one original feature map, the at least one quantized feature map, and the respective importance value.([p. 7 §4.2] "We use the Lipschitz constant (the matrix norm) as a metric for quantization optimization" Lipschitz constant interpreted as synonymous with evaluation value for the second (quantized) artificial neural network.)
wherein the calculating of the evaluation value comprises a distance between a result of applying the respective importance value to each of the at least one original feature map and a result of applying the respective importance value to a quantized feature map, ([p. 7 §4.2] "If we change the quantization bitwidth or switch the quantization method in layer l, different quantization settings will result in different ∆Wl . Let us assume the two quantization settings are q1 and q2, respectively. Then the quantization error difference between q1 and q2 in the output of layer l is: (∆Wl q1 − ∆Wl q2)x l . We may further define L1 = ∆Wl q1 p , L2 = ∆Wl q2 p , ∆L = ∆Wl q1 −∆Wl q2 p . According to the triangle inequality: ∆L ≥ |L1 −L2|." [p. 7] "the quantization loss ∆xql+1=∆Wlxl […] bounding the quantization loss is equivalent to bounding ∆Wl" Song explicitly defines the Lipschitz constant as ||dW||p with distance dW = Wq-W where Wq corresponds to the quantized feature map of the second artificial neural network and W corresponds to the feature map of the original neural network and is calculated based on the element wise importance. Song's distance is the p-norm distance used to bound output error caused by quantization. The comparison between full-precision and quantized feature-map behavior is captured by the quantized weight output deviation from the full-precision output)
among the at least one quantized feature map, corresponding to each of the at least one original feature map([p. 8] "After we adversarially train a model, we attempt to quantize the model layer by layer" [Algorithm 1] "W[i] = model.layer(i); […] model.layer(i) =Wq,0[i] else model.layer(i) =Wq,1[i]" Song's layer-wise replacement creates a one-to-one correspondence between each original layer and its selected quantization layer).
However, Song does not explicitly teach outputting the evaluation value.
Bar, in the same field of endeavor, teaches outputting the evaluation value([p. 4 §3.1] "Note that the assumption on non-expansive activation functions holds for the currently used functions (e.g., ReLU, sigmoid and tanh). The above proposition provides an upper bound for the regularization term in equation 2, which may suggest to replace the regularization on the network derivative with a regularization on the network weights, which is feasible during training. From equation 4, this can be done through a penalty on the weights’ spectral or Frobenius norm" See Table 4 which shows output evaluation values (Lipschitz constraint)).
Song as well as Bar are directed towards bounded regularization for neural network input perturbation analysis. Therefore, Song as well as Bar are analogous art in the same field of endeavor. It would have been obvious before the effective filing date of the claimed invention to combine the teachings of Song with the teachings of Bar by outputting the Lipschitz constant. Bar provides as additional motivation for combination ([p. 8] “Table 1: Bounding the network weights increases its smoothness.”). This motivation for combination also applies to the remaining claims which depend on this combination.
Regarding claim 2, the combination of Song, and Bar teaches The method of claim 1, wherein the at least one original feature map includes a feature map generated in at least one of a plurality of layers included in the first artificial neural network model, and the at least one quantized feature map includes a quantized feature map generated in layers corresponding to respective layers of the first artificial neural network model, which have generated the at least one original feature map, among a plurality of layers included in the second artificial neural network model.(Song [p. 6 §4.1] "Let us assume that for a neural network which has a total of L layers, the input vector and the input perturbation introduced by the adversarial noise of layer l are x l and ∆x l , respectively (both are column vectors with n elements). The full-precision weight matrix and the quantized weight matrix of layer l are Wl and Wl + ∆Wl , respectively (both are m×n matrices). Here ∆Wl is the error introduced by quantization in layer l. Then the output of layer l is" See also Eqn. 3-7 which show that each filter map is layer (l) specific).
Regarding claim 4, the combination of Song, and Bar teaches The method of claim 1, wherein the determining of the respective importance value comprises determining the respective importance value based on a difference between each element of the at least one original feature map and the corresponding element of the at least one perturbed feature map(Song [p. 6 §4.1] "The loss introduced by adversarial and quantization loss in the output of layer l is: ∆x l+1 a,q = x l+1 a,q −x l+1 = a l [(Wl +∆Wl)(x l +∆x l)]−a l (Wl x l). (4) The i-th element of ∆x l+1 a,q is a l [∑ n j=1 (Wl i j +∆Wl i j)(x l j +∆x l j)]−a l (∑ n j=1Wl i jx l j). To get rid of the activation function, let us consider the overall error before activation: (Wl +∆Wl)(x l +∆x l)−Wl x l = Wl∆x l +∆Wl x l +∆Wl∆x l . (5) Its i-th element can be written as ∑ n j=1 (Wl i j∆x l j +∆Wl i jx l j +∆Wl i j∆x l j)." Song at each layer measures the element-wise difference between the output feature map produced by quantized weights + perturbed input and full-precision weights + clean input and bounds it via p-norm to guide layer-wise quantization. Bounded interaction term ||∆Wl∆xl||p where "The i-th element of ∆x l+1 a,q is a l [∑ n j=1 (Wl i j +∆Wl i j)(x l j +∆x l j)]−a l (∑ n j=1Wl i jx l j)" interpreted as importance value for each element of the original feature map W_ij and each perturbed feature map dW_ij.).
Regarding claim 5, the combination of Song, and Bar teaches The method of claim 1, wherein the determining of the respective importance value comprises determining the respective importance value using a metric learning loss function.(Song [p. 7] "For the mutual loss, we will use the induced p-norm to bound the error" The mutual loss interpreted as a metric learning loss function).
Regarding claim 8, the combination of Song, and Bar teaches The method of claim 1, wherein the calculating of the evaluation value comprises calculating the evaluation value based on a distance between a result of applying the respective importance value of each element of a feature map generated by an i-th layer of the first artificial neural network model among the at least one original feature map to the feature map generated by the i-th layer of the first artificial neural network model and a result of applying the respective importance value of each element of the feature map generated by the i-th layer of the first artificial neural network model to a feature map generated by an i-th layer of the second artificial neural network model among the at least one quantized feature map.(Song [p. 7 §4.2] "If we change the quantization bitwidth or switch the quantization method in layer l, different quantization settings will result in different ∆Wl . Let us assume the two quantization settings are q1 and q2, respectively. Then the quantization error difference between q1 and q2 in the output of layer l is: (∆Wl q1 − ∆Wl q2)x l . We may further define L1 = ∆Wl q1 p , L2 = ∆Wl q2 p , ∆L = ∆Wl q1 −∆Wl q2 p . According to the triangle inequality: ∆L ≥ |L1 −L2|." Song explicitly defines the Lipschitz constant as ||dW||p where distance dW = Wq-W corresponding to (a l [(Wl +∆Wl)(x l +∆x l)]−a l (Wl x l)) where Wq corresponds to the quantized feature map (a l [(Wl +∆Wl)(x l +∆x l)]) of the second artificial neural network and W corresponds to the feature map of the original neural network and is calculated based on the element wise importance x l+1 a,q.).
Regarding claims 9, 10, 12, 13, and 16, claims 9, 10, 12, 13, and 16 are directed towards an apparatus for performing the methods of claims 1, 2, 4, 5, and 8, respectively. Therefore, the rejections applied to claims 1, 2, 4, 5, and 8 also apply to claims 9, 10, 12, 13, and 16.
Claims 6 and 14 are rejected under U.S.C. §103 as being unpatentable over the combination of Song and Bar and in further view of Wang (“Smoothed Geometry for Robust Attribution”, 2020).
Regarding claim 6, the combination of Song, and Bar teaches The method of claim 1.
However, the combination of Song, and Bar doesn't explicitly teach wherein the determining of the respective importance value comprises determining the respective importance value using a gradient of each of the at least one original feature map.
Wang, in the same field of endeavor, teaches the determining of the respective importance value comprises determining the respective importance value using a gradient of each of the at least one original feature map.([p. 2] "Definition 1 (Saliency Map (SM) [38]). Given a model f(x), the Saliency Map for an input x is defined as g(x) = Oxf(x)" [p. 4 §3.2] "Viewing Prop. 2 from an adversarial view, it is possible that an adversary happens to find a certain neighbor whose local geometry is totally different from the input so that the chosen noise level is not large enough to produce semantically similar attribution maps. Similar idea can also be applied to Integrated Gradient")).
The combination of Song, and Bar as well as Wang are directed towards bounded regularization for neural network input perturbation analysis. Therefore, the combination of Song, and Bar as well as Wang are analogous art in the same field of endeavor. Song already bounds layer feature map under input perturbation and weight quantization, it would be obvious to one of ordinary skill in the art before the effective filing date of the claimed invention in view of Wang to also bound the gradient which is explicitly stated in Wang ([p. 4 §3.2] "Viewing Prop. 2 from an adversarial view, it is possible that an adversary happens to find a certain neighbor whose local geometry is totally different from the input so that the chosen noise level is not large enough to produce semantically similar attribution maps. Similar idea can also be applied to Integrated Gradient"). Wang provides as additional motivation for combination ([Abstract] “Our experiments on a range of image models demonstrate that both of these mitigations consistently improve attribution robustness, and confirm the role that smooth geometry plays in these attacks on real, large-scale models”).
Regarding claim 14, claim 14 is directed towards an apparatus for performing the method of claim 6. Therefore, the rejections applied to claim 6 also apply to claim 14.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Elthakeb (“Divide and Conquer: Leveraging Intermediate Feature Representations for Quantized Training of Neural Networks”, 2020).
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SIDNEY VINCENT BOSTWICK whose telephone number is (571)272-4720. The examiner can normally be reached M-F 7:30am-5:00pm EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Miranda Huang can be reached on (571)270-7092. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/SIDNEY VINCENT BOSTWICK/Examiner, Art Unit 2124