DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 2026/03/03 & 2026/06/08. The submission is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1–17 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Regarding independent claims 1, 7, and 13
Step 1 — whether the claim falls within a statutory category. See MPEP 2106.03.
Claim 1 is drawn to a method (process); claim 7 is drawn to a device (machine); claim 13 is drawn to a method (process). Therefore, each of these claims falls under one of the four categories of statutory subject matter (process/method, machine/product/apparatus, manufacture, or composition of matter). (Step 1: YES.)
Step 2A Prong One — whether the claim recites a judicial exception. See MPEP 2106.04, subsection II.
Regarding independent claim 1, the claim is directed to an inference method using a dynamic pruning filter in a convolutional neural network model, the method comprising:
The limitations of "generating an attention weight matrix based on a feature map of at least one channel extracted from an input image," "generating at least one mask matrix by referring to a convolution kernel included in the convolutional neural network model," and "outputting the dynamic pruning filter based on the operation of the attention weight matrix and the at least one mask matrix" are directed to the abstract idea of generating a first matrix from data, generating a second matrix from data, and computing an output by an operation on the first and second matrices.
These limitations are directed towards the abstract idea of a mathematical relationship, specifically organizing information and manipulating information through mathematical correlations (see MPEP § 2106.04(a)(2), subsection I). The recited "attention weight matrix," "mask matrix," and the "operation" producing the "dynamic pruning filter" are mathematical constructs computed by matrix operations, as confirmed by the specification, which expresses the mask matrix by Eq. 1 (M_i = H(σ(|W_ij| − T_ij) − 0.5), spec. p. 13) and the output filter by element-wise multiplication and summation over p patterns (spec. p. 14; FIG. 5). The generation and manipulation of these matrices through mathematical operations falls within the mathematical concepts grouping of abstract ideas. (Step 2A Prong One: YES.)
Independent claim 7 is a device claim reciting limitations corresponding to claim 1, in which a processor is configured to "generate an attention weight matrix," "generate at least one mask matrix," and "output the dynamic pruning filter," and is directed to the same abstract idea for the same reasons.
Independent claim 13 is a method claim reciting limitations corresponding to claim 1, in which the model is trained by "generating an attention weight matrix," "generating at least one mask matrix," and "outputting the dynamic pruning filter," and is directed to the same abstract idea for the same reasons.
Step 2A Prong Two — whether the claim as a whole integrates the recited judicial exception into a practical application. See MPEP 2106.04(d).
Regarding independent claim 1, this claim recites the additional elements of a "convolutional neural network model," "a feature map of at least one channel extracted from an input image," and a "convolution kernel."
Evaluating these additional elements individually and in combination: the "convolutional neural network model" and the "convolution kernel" amount to no more than generally linking the use of the judicial exception to a particular technological environment or field of use, namely neural networks (see MPEP § 2106.05(h)). As set forth in AI SME Example 47 with respect to claim 2, a limitation that confines the use of the abstract idea to the technological environment of neural networks fails to add an inventive concept and does not integrate the exception into a practical application. The recitation of "an input image" from which a feature map is "extracted" is mere data gathering recited at a high level of generality and constitutes insignificant extra-solution activity, because all uses of the recited mathematical operations require such input data (see MPEP § 2106.05(g)).
The claim does not recite any improvement to the functioning of a computer or to another technology or technical field. While the specification states that the dynamic pruning filter reduces computational load and enables inference in IoT/mobile devices (spec. p. 6), the claim itself does not reflect any such improvement; claim 1 recites only the generation and mathematical combination of the matrices and does not recite any specific application of the resulting filter to reduce computation, accelerate inference, or perform a technical task. See MPEP § 2106.04(d)(1) and § 2106.05(a) (the claim itself must reflect the disclosed improvement). The additional elements, individually and in combination, therefore do no more than apply the abstract idea using generic neural-network components as a tool. Accordingly, the claim as a whole does not integrate the recited judicial exception into a practical application (Step 2A Prong Two: NO), and the claim is directed to the abstract idea. (Step 2A: YES.)
Regarding independent claim 7, this claim is drawn to a device reciting limitations corresponding to claim 1 and is rejected under the same rationale. Claim 7 additionally recites "a memory configured to store one or more instructions" and "a processor configured to execute the one or more instructions." These additional elements amount to no more than mere instructions to apply the exception using a generic computer, and recite generic computer components that merely act as a tool on which the abstract idea operates (see MPEP § 2106.05(f)). They therefore fail to integrate the exception into a practical application.
Regarding independent claim 13, this claim is drawn to a method reciting limitations corresponding to claim 1 and is rejected under the same rationale. Claim 13 additionally recites "an electronic device including a memory and a processor," "preparing training data including training input images and training label data," and "inputting the training input images to the convolutional neural network model." The memory and processor recite generic computer components acting as a tool (see MPEP § 2106.05(f)); the preparing of training data and inputting of training images are mere data gathering recited at a high level of generality and constitute insignificant extra-solution activity (see MPEP § 2106.05(g)). These additional elements, individually and in combination, do not integrate the exception into a practical application.
Step 2B — whether the claim as a whole amounts to significantly more than the recited exception. See MPEP 2106.05.
Regarding independent claims 1, 7, and 13, as explained with respect to Step 2A Prong Two, the additional elements are the convolutional neural network model, the convolution kernel, the input image / training images, and (for claims 7 and 13) the generic memory and processor.
The recitation of the convolutional neural network model and convolution kernel amounts to mere instructions to apply the abstract idea in the technological environment of neural networks and does not provide an inventive concept (see MPEP § 2106.05(f) and (h)). The generic memory and processor of claims 7 and 13 amount to no more than mere instructions to apply the exception using a generic computer component (see MPEP § 2106.05(f)).
A conclusion that an additional element is insignificant extra-solution activity in Step 2A Prong Two is re-evaluated in Step 2B to determine whether it is well-understood, routine, and conventional. See MPEP § 2106.05, subsection I.A. The recitation of extracting a feature map from an input image, and of preparing and inputting training images, amounts to receiving or gathering data recited at a high level of generality and is well-understood, routine, and conventional activity in the field of neural-network computing. See MPEP § 2106.05(d).
Even when considered in combination, these additional elements represent mere instructions to implement the abstract idea on generic neural-network and computer components, together with insignificant extra-solution activity, and do not provide an inventive concept. (Step 2B: NO.) Claims 1, 7, and 13 are ineligible.
Regarding dependent claims 2–6, 8–12, and 14–17
Step 2A Prong One — whether the claim recites a judicial exception. See MPEP 2106.04, subsection II.
Claims 2–6, 8–12, and 14–17 merely narrow the abstract idea recited in the independent claims and recite further mathematical operations, as follows:
Claims 2, 8, 14 recite "determining importance of the at least one channel based on an average value of the at least one channel determined through global average pooling (GAP)" and "generating the attention weight matrix based on the importance of the at least one channel." These limitations are directed towards the abstract idea of a mathematical relationship, namely computing an average value and generating a matrix therefrom (see MPEP § 2106.04(a)(2), subsection I).
Claims 3, 9, 15 recite "generating the at least one mask matrix by performing static pruning on the weight matrix of the convolution kernel." This limitation is directed towards the abstract idea of a mathematical relationship, namely deriving a matrix from a weight matrix.
Claims 4, 10, 16 recite "calculating a difference between each element of the weight matrix of the convolution kernel and a pre-determined threshold value," "determining a binarized value for each element by applying a binary step function to the difference," and "generating the at least one mask matrix based on the binarized value for each element." These limitations are directed towards the abstract idea of a mathematical relationship, namely computing a difference, applying a step function, and forming a matrix, as expressed by Eq. 1 of the specification (spec. p. 13).
Claims 5, 11, 17 recite "performing element-wise multiplication on the weight matrix of the convolution kernel, the at least one mask matrix, and the attention weight matrix" and "outputting the dynamic pruning filter based on the element-wise multiplication." These limitations are directed towards the abstract idea of a mathematical relationship, namely element-wise matrix multiplication.
Claims 6, 12 recite "performing inference for image detection or image classification using the dynamic pruning filter." This limitation is addressed under Prong Two below.
(Step 2A Prong One: YES for claims 2–6, 8–12, and 14–17.)
Step 2A Prong Two — whether the claim as a whole integrates the exception into a practical application. See MPEP 2106.04(d).
Claims 2–5, 8–11, and 14–17 recite no additional elements beyond those of the independent claims from which they depend, and for the reasons set forth above with respect to the independent claims, these judicial exceptions are not integrated into a practical application. The claims recite further mathematical operations and do not provide anything more than the mathematical relationships that manipulate matrices through mathematical correlations.
Regarding claims 6 and 12, these claims recite the additional element of "performing inference for image detection or image classification using the dynamic pruning filter." This limitation applies the abstract idea using the dynamic pruning filter without placing any limits on how the inference is performed; it recites only the outcome of "inference for image detection or image classification" and does not recite any details of how the detection or classification is accomplished (see MPEP § 2106.05(f), as applied in AI SME Example 47). The recitation further merely indicates a field of use — image detection or classification — in which the abstract idea is performed (see MPEP § 2106.05(h)). Claims 6 and 12 therefore do not integrate the exception into a practical application.
Step 2B — whether the claims amount to significantly more. See MPEP 2106.05.
For the reasons set forth above with respect to the independent claims, the additional elements of claims 2–6, 8–12, and 14–17, considered individually and in combination, amount to no more than mere instructions to apply the abstract idea using generic neural-network and computer components and insignificant extra-solution activity, and do not provide an inventive concept. The recitation in claims 6 and 12 of "performing inference for image detection or image classification using the dynamic pruning filter" is, at best, mere instructions to apply the abstract idea and does not amount to significantly more. (Step 2B: NO.) Claims 2–6, 8–12, and 14–17 are ineligible.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The text of those sections of Title 35, U.S. Code not included in this action can be found in a prior Office action.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 1-3, 5-9, 11-15, and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Chen et al. (Chen), Non-Patent Literature, "Dynamic Convolution: Attention over Convolution Kernels," arXiv:1912.03458v2 [cs.CV], published 31 Mar 2020, in view of Wang et al. (Wang), U.S. Patent Application Publication No. US 2021/0256385 A1.
Regarding claim 1, (Chen) teaches:
"generating an attention weight matrix based on a feature map of at least one channel extracted from an input image" ((Chen) teaches at p. 3, § 3.2 (Dynamic Convolution): "we apply squeeze-and-excitation to compute kernel attentions {π_k(x)}… The global spatial information is firstly squeezed by global average pooling. Then we use two fully connected layers (with a ReLU between them) and softmax to generate normalized attention weights for K convolution kernels," whereby the normalized attention weights {π_k(x)} constitute the attention weight matrix; (Chen) further teaches at p. 3, § 3.2 that the attention weights are generated from the input feature map: "For an input feature map with dimension H × W × C_in, the attention requires…," and at p. 3, § 3.1 that the attention weights "vary for each input x" and "are functions of input")
(Chen) teaches something related to an inference method using a dynamic pruning filter in a convolutional neural network model and to generating at least one mask matrix by referring to a convolution kernel, in that (Chen) performs inference using a convolutional neural network model in which a filter is dynamically generated from K convolution kernels (p. 3, § 3.2: "dynamic convolution… has K convolution kernels that share the same kernel size and input/output dimensions"). However, (Chen) performs no pruning and generates no mask, and therefore (Chen) does not teach:
"An inference method using a dynamic pruning filter in a convolutional neural network model, the method comprising:"
"generating at least one mask matrix by referring to a convolution kernel included in the convolutional neural network model"
In the same field of endeavor, (Wang) teaches "An inference method using a dynamic pruning filter in a convolutional neural network model, the method comprising:" ((Wang, p. 2, claim 1: "A computer-implemented method… for compressing a deep neural network (DNN) model by DNN weight pruning to accelerate DNN inference on mobile devices"; p. 2, claim 1(a): "performing an intra-convolution kernel pruning of the DNN model wherein a fixed number of weights are pruned in each convolution kernel of the DNN model to generate sparse convolution patterns"; p. 6, ¶ [0075]: the kernel patterns are called "dynamically during DNN execution"), whereby (Wang) performs inference using a dynamically applied pruning filter in a convolutional neural network model.
(Wang) further teaches "generating at least one mask matrix by referring to a convolution kernel included in the convolutional neural network model" ((Wang, p. 3, ¶ [0039]: "we consider it as incorporating an additional convolution kernel P to perform element-wise multiplication with the original kernel. P is termed the Sparse Convolution Pattern (SCP), with dimension H_l × W_l and binary-valued elements (0 and 1)"). The binary-valued (0 and 1) matrix P having the convolution kernel's spatial dimension H_l × W_l reads on the recited "mask matrix," and it is generated "by referring to a convolution kernel" because, per (Wang, p. 3, ¶ [0039]): "the white blocks denote a fixed number of pruned weights in each kernel. The remaining red blocks in each kernel have arbitrary weight values, while their locations form a specific SCP P_i," whereby the 0 and 1 locations of the mask are derived from the convolution kernel.
(Chen) and (Wang) are analogous to the claimed invention as both are from the same field of endeavor of accelerating convolutional neural network inference through model compression. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the input-dependent attention weight matrix of (Chen) with the inference method using a dynamic pruning filter and the mask matrix generated by reference to the convolution kernel of (Wang). The motivation to combine (Chen) and (Wang) is that (Chen) acknowledges that its K parallel convolution kernels increase model size (p. 3, § 3.2: "even though dynamic convolution increases the model size"), while (Wang) teaches applying a limited set of binary sparse convolution patterns to the convolution kernels so as to "accelerate DNN inference on mobile devices" (p. 2, claim 1) and teaches that its pattern pruning "can be integrated in the same algorithm-level solution" (p. 3, ¶ [0040]); one of ordinary skill would apply the mask matrix of (Wang) to the convolution kernels aggregated by the attention weight matrix of (Chen) in order to obtain the input-dependent representation power of (Chen) while reducing computational and memory cost, with a reasonable expectation of success as recited by (Chen) at p. 1, Abstract ("Assembling multiple kernels is not only computationally efficient due to the small kernel size, but also has more representation power since these kernels are aggregated in a non-linear way via attention").
The combination of (Chen) and (Wang) teaches something related to outputting the dynamic pruning filter based on the operation of the attention weight matrix and the at least one mask matrix. (Chen) teaches the operation of the attention weight matrix at p. 3, § 3.1, Eq. 1: "W̃(x) = Σ_{k=1}^{K} π_k(x) W̃_k," and at p. 3, § 3.2: "They are aggregated by using the attention weights {π_k}," whereby an aggregated kernel W̃(x) is output as the result of the attention operation. (Wang) teaches the operation of the mask matrix at p. 3, ¶ [0039]: the sparse convolution pattern performs "element-wise multiplication with the original kernel," whereby a pruned kernel is output as the result of the mask operation. However, the combination of (Chen) and (Wang) does not expressly teach:
"outputting the dynamic pruning filter based on the operation of the attention weight matrix and the at least one mask matrix"
as a single operation in which the attention weight matrix and the at least one mask matrix jointly operate to produce one output filter.
One of ordinary skill in the art, however, would have readily inferred this limitation from the combined teachings of (Chen) and (Wang). Because (Chen) teaches that the attention weight matrix operates upon the convolution kernels to output the aggregated kernel (p. 3, § 3.1, Eq. 1) and (Wang) teaches that the mask matrix operates upon those same convolution kernels by "element-wise multiplication with the original kernel" (p. 3, ¶ [0039]), one of ordinary skill would have readily inferred that applying the attention aggregation of (Chen) to the masked kernels of (Wang) outputs a single dynamic pruning filter that is both input-adaptively aggregated and sparse. Such a combination amounts to the combination of familiar elements according to known methods yielding predictable results. It therefore would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to output the dynamic pruning filter based on the operation of the attention weight matrix of (Chen) and the at least one mask matrix of (Wang).
As to Claim 2, which depends from claim 1, the limitations of claim 1 are rejected under the same rationale set forth above with respect to claim 1 over (Chen) in view of (Wang) and further in view of (Azarian Yazdi). Claim 2 further recites:
"determining importance of the at least one channel based on an average value of the at least one channel determined through global average pooling (GAP)"
"generating the attention weight matrix based on the importance of the at least one channel"
Regarding the limitation "determining importance of the at least one channel based on an average value of the at least one channel determined through global average pooling (GAP)," (Chen) teaches this element at p. 3, § 3.2 (Dynamic Convolution): "we apply squeeze-and-excitation to compute kernel attentions {π_k(x)}… The global spatial information is firstly squeezed by global average pooling." (Chen) thereby teaches determining the importance of the at least one channel from an average value obtained by global average pooling of the input feature map, where the global average pooling operation of (Chen) computes the average value of the at least one channel of the H × W × C_in feature map ((Chen, p. 3, § 3.2: "For an input feature map with dimension H × W × C_in").
Regarding the limitation "generating the attention weight matrix based on the importance of the at least one channel," (Chen) teaches this element at p. 3, § 3.2 (Dynamic Convolution): "Then we use two fully connected layers (with a ReLU between them) and softmax to generate normalized attention weights for K convolution kernels." (Chen) thereby teaches generating the attention weight matrix {π_k(x)} from the squeezed channel information produced by the global average pooling, such that the attention weight matrix is generated based on the importance of the at least one channel.
As to Claim 3, which depends from claim 1, the limitations of claim 1 are rejected under the same rationale set forth above with respect to claim 1 over (Chen) in view of (Wang). Claim 3 further recites:
"generating the at least one mask matrix by performing static pruning on the weight matrix of the convolution kernel"
Regarding the limitation "generating the at least one mask matrix by performing static pruning on the weight matrix of the convolution kernel," (Chen) teaches something related in that it operates upon convolution kernels (p. 3, § 3.2: "dynamic convolution… has K convolution kernels that share the same kernel size and input/output dimensions"). However, (Chen) does not teach:
"generating the at least one mask matrix by performing static pruning on the weight matrix of the convolution kernel"
In the same field of endeavor, (Wang) teaches "generating the at least one mask matrix by performing static pruning on the weight matrix of the convolution kernel" (Wang, p. 3, ¶ [0039]: "we consider it as incorporating an additional convolution kernel P to perform element-wise multiplication with the original kernel. P is termed the Sparse Convolution Pattern (SCP), with dimension H_l × W_l and binary-valued elements (0 and 1)"; "the white blocks denote a fixed number of pruned weights in each kernel. The remaining red blocks in each kernel have arbitrary weight values, while their locations form a specific SCP P_i"). (Wang) teaches that the sparse convolution pattern P is a fixed, predetermined binary mask applied to the weight values of the convolution kernel, which reads on "performing static pruning on the weight matrix of the convolution kernel" because the pattern prunes a fixed number of weights in each kernel independent of the input. (Wang) further teaches that the number of such patterns is limited (p. 3, ¶ [0039]: "the total number of SCP types is limited"; p. 10, claim 4: "assigning a pattern from a limited set of sparse convolution patterns to each kernel of the DNN model"), corresponding to the recited generation of the at least one mask matrix by static pruning.
(Chen) and (Wang) are analogous to the claimed invention as both are from the same field of endeavor of accelerating convolutional neural network inference through model compression. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the attention-based kernel aggregation of (Chen) with the static pruning of the convolution kernel weight matrix by a limited set of sparse convolution patterns of (Wang). The motivation to combine (Chen) and (Wang) is that (Wang) teaches that applying a limited set of static sparse convolution patterns to the convolution kernels produces the same sparsity ratio in each filter and "a limited number of pattern shapes" (p. 10, claim 3), so as to accelerate DNN inference on mobile devices (p. 2, claim 1); one of ordinary skill would generate the at least one mask matrix by performing static pruning on the weight matrix of the convolution kernel as taught by (Wang) in order to reduce the computational and memory cost of the combined network in a hardware-friendly manner, with a reasonable expectation of success.
As to Claim 5, which depends from claim 1, the limitations of claim 1 are rejected under the same rationale set forth above with respect to claim 1 over (Chen) in view of (Wang). Claim 5 further recites:
"performing element-wise multiplication on the weight matrix of the convolution kernel, the at least one mark matrix, and the attention weight matrix"
"outputting the dynamic pruning filter based on the element-wise multiplication"
Claim 5 recites "the at least one mark matrix." The specification does not describe a "mark matrix"; the specification uniformly describes a "mask matrix" (see spec., pp. 3, 12–13). Under the broadest reasonable interpretation consistent with the specification, "the at least one mark matrix" is interpreted as "the at least one mask matrix." See MPEP § 2111. The limitation is mapped accordingly below.
Regarding claim 5, (Chen) teaches:
the multiplication of the attention weight matrix with the weight matrix of the convolution kernel ((Chen) teaches at p. 3, § 3.1, Eq. 1: "W̃(x) = Σ_{k=1}^{K} π_k(x) W̃_k," and at p. 3, § 3.2: "They are aggregated by using the attention weights {π_k}," whereby each convolution kernel W̃_k is multiplied by its corresponding attention weight π_k(x) and the products are summed to output the aggregated kernel W̃(x))
(Chen) teaches something related to "performing element-wise multiplication on the weight matrix of the convolution kernel, the at least one mark matrix, and the attention weight matrix," in that (Chen) multiplies the weight matrix of the convolution kernel by the attention weight matrix. However, (Chen) generates no mask matrix, and therefore (Chen) does not teach:
"performing element-wise multiplication on the weight matrix of the convolution kernel, the at least one mark matrix, and the attention weight matrix"
In the same field of endeavor, (Wang) teaches "performing element-wise multiplication on the weight matrix of the convolution kernel, the at least one mark matrix, and the attention weight matrix" to the extent of the element-wise multiplication of the weight matrix of the convolution kernel and the mask matrix ((Wang, p. 3, ¶ [0039]: "we consider it as incorporating an additional convolution kernel P to perform element-wise multiplication with the original kernel. P is termed the Sparse Convolution Pattern (SCP), with dimension H_l × W_l and binary-valued elements (0 and 1)"), where the sparse convolution pattern P is the mask matrix and the original kernel is the weight matrix of the convolution kernel.
(Chen) and (Wang) are analogous to the claimed invention as both are from the same field of endeavor of accelerating convolutional neural network inference through model compression. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the multiplication of the attention weight matrix with the convolution kernel of (Chen) with the element-wise multiplication of the mask matrix and the convolution kernel of (Wang), such that the weight matrix of the convolution kernel, the mask matrix, and the attention weight matrix are element-wise multiplied. The motivation to combine (Chen) and (Wang) is that (Chen) teaches that aggregating the kernels by attention provides greater representation power (p. 1, Abstract: "these kernels are aggregated in a non-linear way via attention") while acknowledging that the K parallel kernels increase model size (p. 3, § 3.2: "even though dynamic convolution increases the model size"), and (Wang) teaches that element-wise multiplication of the kernel by a binary sparse convolution pattern accelerates DNN inference on mobile devices (p. 3, ¶ [0039]; p. 2, claim 1); one of ordinary skill would element-wise multiply the weight matrix by both the mask matrix of (Wang) and the attention weight matrix of (Chen) in order to obtain an input-adaptive and sparse result, with a reasonable expectation of success.
The combination of (Chen) and (Wang) teaches something related to "outputting the dynamic pruning filter based on the element-wise multiplication." (Chen) teaches outputting the aggregated kernel W̃(x) as the result of the attention multiplication (p. 3, § 3.1, Eq. 1), and (Wang) teaches outputting the pruned kernel as the result of the element-wise mask multiplication (p. 3, ¶ [0039]). However, the combination of (Chen) and (Wang) does not expressly teach:
"outputting the dynamic pruning filter based on the element-wise multiplication"
as the output of a single element-wise multiplication involving all three of the weight matrix, the mask matrix, and the attention weight matrix.
One of ordinary skill in the art, however, would have readily inferred this limitation from the combined teachings of (Chen) and (Wang). (Wang) teaches that "the total number of SCP types is limited" and that a pattern is assigned "from a limited set of sparse convolution patterns to each kernel" (p. 3, ¶ [0039]; p. 10, claim 4), such that a plurality of mask matrices corresponding to the limited set of patterns is present, and (Chen) teaches summing the attention-weighted kernels to output a single filter (p. 3, § 3.1, Eq. 1: "W̃(x) = Σ_{k=1}^{K} π_k(x) W̃_k"). Because (Wang) already provides plural mask matrices applied to the convolution kernel by element-wise multiplication and (Chen) already outputs a single filter from the attention-weighted sum of plural kernels, one of ordinary skill would have readily inferred that element-wise multiplying the weight matrix, the mask matrix, and the attention weight matrix outputs the dynamic pruning filter. Such a combination amounts to the combination of familiar elements according to known methods yielding predictable results. It therefore would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to output the dynamic pruning filter based on the element-wise multiplication.
As to Claim 6, which depends from claim 1, the limitations of claim 1 are rejected under the same rationale set forth above with respect to claim 1 over (Chen) in view of (Wang) and further in view of (Azarian Yazdi). Claim 6 further recites:
"performing inference for image detection or image classification using the dynamic pruning filter"
Regarding "performing inference for image detection or image classification using the dynamic pruning filter," (Chen) teaches this element at p. 1, Abstract: "the top-1 accuracy of ImageNet classification is boosted by 2.9% with only 4% additional FLOPs and 2.9 AP gain is achieved on COCO keypoint detection," and at p. 2, § 1: "on both image classification (ImageNet) and keypoint detection (COCO)." (Chen) thereby teaches performing inference for image classification and image detection using the dynamically aggregated convolution kernel, which in the combination is the dynamic pruning filter.
Claim 7 is a device claim reciting the corresponding limitations of claim 1 in the form of a memory and a processor configured to perform the recited functions, and is rejected under the same rationale set forth above with respect to claim 1.
Claim 8 is rejected under the same rationale as claim 2.
Claim 9 is rejected under the same rationale as claim 3.
Claim 11 is rejected under the same rationale as claim 5.
Claim 12 is rejected under the same rationale as claim 6.
Regarding independent claim 13, (Chen) teaches:
"inputting the training input images to the convolutional neural network model" ((Chen) teaches at p. 5, § 5 (Experiments: ImageNet Classification): the dynamic convolution networks are trained on ImageNet, whereby training input images are input to the convolutional neural network model)
"training the convolutional neural network model by generating an attention weight matrix based on a feature map of at least one channel extracted from the training input images" ((Chen) teaches at p. 3, § 3.2 (Dynamic Convolution): "we apply squeeze-and-excitation to compute kernel attentions {π_k(x)}… The global spatial information is firstly squeezed by global average pooling. Then we use two fully connected layers (with a ReLU between them) and softmax to generate normalized attention weights for K convolution kernels," whereby the normalized attention weights {π_k(x)} constitute the attention weight matrix; (Chen) further teaches at p. 3, § 3.2 that the attention weights are generated from the input feature map: "For an input feature map with dimension H × W × C_in, the attention requires…," and at p. 3, § 3.1 that the attention weights "vary for each input x" and "are functions of input"; (Chen) teaches training the model so as to generate the attention weight matrix at p. 5, § 5)
(Chen) teaches something related to a method for training a convolutional neural network model for use in an electronic device, to preparing training data including training label data including a dynamic pruning filter, and to generating at least one mask matrix by referring to a convolution kernel, in that (Chen) trains a convolutional neural network model end-to-end to output a dynamically aggregated convolution kernel (p. 3, § 3.2: "dynamic convolution… has K convolution kernels that share the same kernel size and input/output dimensions"; p. 5, § 5). However, (Chen) performs no pruning, generates no mask, and does not recite an electronic device including a memory and a processor, and therefore (Chen) does not teach:
"A method for training a convolutional neural network model for a use in electronic device including a memory and a processor, the method comprising:"
"preparing training data including training input images and training label data including a dynamic pruning filter"
"generating at least one mask matrix by referring to a convolution kernel included in the convolutional neural network model"
In the same field of endeavor, (Wang) teaches "A method for training a convolutional neural network model for a use in electronic device including a memory and a processor, the method comprising:" ((Wang, p. 2, claim 1: "A computer-implemented method… for compressing a deep neural network (DNN) model by DNN weight pruning to accelerate DNN inference on mobile devices," and p. 2, claim 1(c): "training the DNN model compressed in steps (a) and (b)"; (Wang, p. 10, claim 17: "a non-transitory computer readable medium having stored thereon a computer program… that, when executed by a computer processor, cause that computer processor to" perform the recited steps, whereby the method is for use in an electronic device including a memory and a processor).
(Wang) further teaches "preparing training data including training input images and training label data including a dynamic pruning filter" ((Wang, p. 3, ¶ [0038]: "An L-layer DNN can be expressed as a feature extractor F_L(F_{L−1}(…F_1(X)…))," where X is an input image; (Wang, p. 2, claim 1(c): "training the DNN model compressed in steps (a) and (b)," whereby the model to which the sparse convolution patterns have been applied is trained; and (Wang, p. 6, ¶ [0075]: the kernel patterns are called "dynamically during DNN execution," whereby the pattern applied to the kernel constitutes a dynamic pruning filter).
(Wang) further teaches "generating at least one mask matrix by referring to a convolution kernel included in the convolutional neural network model" ((Wang, p. 3, ¶ [0039]: "we consider it as incorporating an additional convolution kernel P to perform element-wise multiplication with the original kernel. P is termed the Sparse Convolution Pattern (SCP), with dimension H_l × W_l and binary-valued elements (0 and 1)"). The binary-valued (0 and 1) matrix P having the convolution kernel's spatial dimension H_l × W_l reads on the recited "mask matrix," and it is generated "by referring to a convolution kernel" because, per (Wang, p. 3, ¶ [0039]): "the white blocks denote a fixed number of pruned weights in each kernel. The remaining red blocks in each kernel have arbitrary weight values, while their locations form a specific SCP P_i," whereby the 0 and 1 locations of the mask are derived from the convolution kernel.
(Chen) and (Wang) are analogous to the claimed invention as both are from the same field of endeavor of accelerating convolutional neural network inference through model compression. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the training of the convolutional neural network model to generate the input-dependent attention weight matrix of (Chen) with the training of the pruned model, the preparation of training data, and the mask matrix generated by reference to the convolution kernel of (Wang). The motivation to combine (Chen) and (Wang) is that (Chen) acknowledges that its K parallel convolution kernels increase model size (p. 3, § 3.2: "even though dynamic convolution increases the model size"), while (Wang) teaches applying a limited set of binary sparse convolution patterns to the convolution kernels and training the resulting compressed model so as to "accelerate DNN inference on mobile devices" (p. 2, claims 1 and 1(c)) and teaches that its pattern pruning "can be integrated in the same algorithm-level solution" (p. 3, ¶ [0040]); one of ordinary skill would train the convolutional neural network model of (Chen) using the mask matrix and compressed-model training of (Wang) in order to obtain the input-dependent representation power of (Chen) while reducing computational and memory cost, with a reasonable expectation of success as recited by (Chen) at p. 1, Abstract ("Assembling multiple kernels is not only computationally efficient due to the small kernel size, but also has more representation power since these kernels are aggregated in a non-linear way via attention").
The combination of (Chen) and (Wang) does not clearly teach "outputting the dynamic pruning filter, which is the training label data, based on the operation of the attention weight matrix and the at least one mask matrix." (Chen) trains the network end-to-end to output the dynamically aggregated convolution kernel based on the operation of the attention weight matrix (p. 3, § 3.1, Eq. 1: "W̃(x) = Σ_{k=1}^{K} π_k(x) W̃_k"; p. 5, § 5), and (Wang) trains the compressed model in which the mask matrix operates upon the convolution kernel by "element-wise multiplication with the original kernel" (p. 3, ¶ [0039]; p. 2, claim 1(c)); however, neither (Chen) nor (Wang) expressly teaches that the output dynamic pruning filter is itself the training label data, as the limitation literally recites.
The limitation is addressed in the specification of the instant application. Per the specification, p. 5, the training method comprises "preparing training data including training input images and training label data including a dynamic pruning filter" and "outputting the dynamic pruning filter, which is the training label data, based on the operation of the attention weight matrix and the at least one mask matrix." Per the specification, p. 15: "the convolutional neural network may receive an image as input and output a dynamic pruning filter," and "training is carried out in the direction that the threshold value becomes relatively large or small, and a pattern corresponding to at least one mask matrix may be determined." Read in light of this disclosure, the limitation "outputting the dynamic pruning filter, which is the training label data" is given its broadest reasonable interpretation as a convolutional neural network model that is trained to output a dynamic pruning filter.
Under this broadest reasonable interpretation, the limitation is met by the combination of (Chen) in view of (Wang), as set forth above: (Chen) trains the network end-to-end to output the aggregated kernel based on the operation of the attention weight matrix (p. 3, § 3.2; p. 5, § 5), and (Wang) applies the mask matrix to the convolution kernel by element-wise multiplication and trains the resulting compressed model (p. 3, ¶ [0039]; p. 2, claim 1(c)). It therefore would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to train the convolutional neural network model to output the dynamic pruning filter based on the operation of the attention weight matrix and the at least one mask matrix, for the motivation set forth above.
Claim 14 is rejected under the same rationale as claims 2 and 8.
Claim 15 is rejected under the same rationale as claims 3 and 9.
Claim 17 is rejected under the same rationale as claims 5 and 11.
Claims 4, 10, and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Chen et al. (Chen), Non-Patent Literature, "Dynamic Convolution: Attention over Convolution Kernels," arXiv:1912.03458v2 [cs.CV], published 31 Mar 2020, in view of Wang et al. (Wang), U.S. Patent Application Publication No. US 2021/0256385 A1, and further in view of Azarian Yazdi et al. (Azarian Yazdi), U.S. Patent Application Publication No. US 2021/0158166 A1.
As to Claim 4, which depends from claim 3, the limitations of claims 1 and 3 are rejected under the same rationale set forth above with respect to claims 1 and 3 over (Chen) in view of (Wang) and further in view of (Azarian Yazdi). Claim 4 further recites:
"calculating a difference between each element of the weight matrix of the convolution kernel and a pre-determined threshold value"
"determining a binarized value for each element by applying a binary step function to the difference"
"generating the at least one mask matrix based on the binarized value for each element"
The combination of (Chen) and (Wang) does not teach the above limitations. In the same field of endeavor, (Azarian Yazdi) teaches these limitations.
Regarding "calculating a difference between each element of the weight matrix of the convolution kernel and a pre-determined threshold value," (Azarian Yazdi) teaches this element at p. 5, ¶ [0053], Eq. 1: "V_kl = W_kl × sigm((W_kl − τ_l)/T)," where the term (W_kl − τ_l) is the difference between each element W_kl of the convolution kernel weight matrix and the threshold τ_l. (Azarian Yazdi) teaches that τ_l is a threshold value at p. 5, ¶ [0052]: "the LTP method learns a threshold for each layer of the neural network. The learned threshold may be referred to as a layer threshold," and at p. 6, ¶ [0066]: "the kernel norms may be pruned based on a comparison with the learned threshold τ_l."
Regarding "determining a binarized value for each element by applying a binary step function to the difference," (Azarian Yazdi) teaches this element at p. 6, ¶ [0054]: "the sigmoid function outputs zero if a value of an input to the sigmoid function… is less than 0.5 and outputs a one if the value of the input is equal to or greater than 0.5," and at p. 6, ¶ [0064]: "During inference, the sigmoid function may be replaced with a hard-limiter, such that all weights below the corresponding threshold are pruned," and at p. 6, ¶ [0065]: the differentiable functions "may converge to a hard-limiter or step function through annealing the temperature parameter." (Azarian Yazdi) thereby teaches applying a binary step function (hard-limiter) to the difference (W_kl − τ_l) to determine a binarized value of zero or one for each element.
Regarding "generating the at least one mask matrix based on the binarized value for each element," (Azarian Yazdi) teaches this element at p. 6, ¶ [0054]: "An output of one represents an un-pruned weight," and at p. 5, ¶ [0053]: "V_kl = W_kl × sigm(...)," whereby the binarized value (zero or one) for each element forms the mask that is multiplied with the weight matrix, such that the mask matrix is generated based on the binarized value for each element.
(Chen), (Wang) and (Azarian Yazdi) are analogous to the claimed invention as all are from the same field of endeavor of accelerating convolutional neural network inference through model compression. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the attention-based kernel aggregation of (Chen) and the static pruning of the convolution kernel weight matrix of (Wang) with the threshold-comparison and binary-step mask generation of (Azarian Yazdi). The motivation to combine (Chen), (Wang) and (Azarian Yazdi) is that (Azarian Yazdi) teaches learning the pruning threshold during training (p. 5, ¶ [0052]) and considering the regularization loss "in conjunction with the classification loss L to balance classification performance and a number of pruned weights" (p. 6, ¶ [0059]); one of ordinary skill would generate the mask matrix of the combined network by calculating the difference between each kernel weight element and a threshold and applying a binary step function as taught by (Azarian Yazdi) in order to determine which kernel elements to prune in a principled, trainable manner, thereby improving the accuracy-versus-sparsity trade-off with a reasonable expectation of success.
Claim 10 is rejected under the same rationale as claim 4.
Claim 16 is rejected under the same rationale as claims 4 and 10.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to HUNG VAN LE whose telephone number is (571)270-0164. The examiner can normally be reached 8 a.m. - 5 p.m..
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Cesar Paula can be reached at (571) 272-4128. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/HUNG VAN LE/Examiner, Art Unit 2145
/CESAR B PAULA/Supervisory Patent Examiner, Art Unit 2145