yesDETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
Claims 1-9 have been interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because they use a generic placeholder “a pruning unit”, “an optimizing unit”, “a learning” coupled with functional language "configured to” without reciting sufficient structure to achieve the function. Furthermore, the generic placeholder is not preceded by a structural modifier.
Since the claim limitation(s) invokes 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, claims 1-9 have been interpreted to cover the corresponding structure described in the specification that achieves the claimed function, and equivalents thereof.
A review of the specification shows that the following appears to be the corresponding structure described in the specification for the 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph limitation: Fig. 1 and [0035] “The terms ‘unit,’ ‘part,’ ‘module,’ ‘member,’ ‘block,’ and the like as used may be implemented in software and hardware. Further, a plurality of ‘unit,’ ‘part,’ ‘module,’ ‘member,’ ‘block,’ and the like as used may be implemented as one component. It is also possible that ‘unit,’ ‘part,’ ‘module,’ ‘member,’ ‘block,’ and the like includes a plurality of components.”; [0042] “When a component, device, element, or the like of the present disclosure is described as having a purpose or performing an operation, function, or the like, the component, device, or element should be considered herein as being ‘configured to’ meet that purpose or to perform that operation or function.”; examples and functions of pruning unit, optimizing unit and learning unit are disclosed in [0045]-[0054]. For claims directed to computer-implemented functions, the corresponding structure is the particular algorithm disclosed in the specification. Harris Corp. v. Ericsson Inc., 417 F.3d 1241, 1249 (Fed. Cir. 2005). As such, Fig. 1 [0035], [0042] and [0045]-[0054] of the specification cited above will be interpreted to cover the corresponding structure disclosed in the specification and equivalents thereof.
If applicant wishes to provide further explanation or dispute the examiner’s interpretation of the corresponding structure, applicant must identify the corresponding structure with reference to the specification by page and line number, and to the drawing, if any, by reference characters in response to this Office action.
If applicant does not intend to have the claim limitation(s) treated under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112 , sixth paragraph, applicant may amend the claim(s) so that it/they will clearly not invoke 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, or present a sufficient showing that the claim recites/recite sufficient structure, material, or acts for performing the claimed function to preclude application of 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph.
For more information, see MPEP § 2173 et seq. and Supplementary Examination Guidelines for Determining Compliance With 35 U.S.C. 112 and for Treatment of Related Issues in Patent Applications, 76 FR 7162, 7167 (Feb. 9, 2011).
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefore, subject to the conditions and requirements of this title.
Claims 1 and 10 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Claim 1 recites the same limitations as claim 10.
Step 1: Is the claim to a process, machine, manufacture, or composition of matter?
Yes, claim 10 is directed to a method. Claim 1 is directed to a system.
Step 2A Prong 1: Does claim 10 recite an abstract idea, law of nature, or natural phenomenon?
The limitations of:
Obtaining a pruned neural network by performing pruning on a base neural network; (Mental process; A human can obtain a pruned network by performing pruning on a base network.)
Obtaining an optimized hyperparameter set by performing HPO a predetermined number of times on the pruned neural network; (Mental process; A human can determine an optimized hyperparameter by performing HPO a certain number of times.)
Step 2A Prong 2: Does claim 10 recite elements that integrate the judicial exception into a practical application?
Training the base neural network using the optimized hyperparameter set (The limitation amounts to merely applying the abstract idea to a generic component, in this case applying the optimized hyperparameter set to the neural network (generic machine learning model). This does not amount to significantly more than the exception itself (MPEP 2106.05(f))).
Step 2B: Does claim 10 recite elements that amount to significantly more than the judicial exception?
Training the base neural network using the optimized hyperparameter set (The limitation amounts to merely applying the abstract idea to a generic component, in this case applying the optimized hyperparameter set to the neural network (generic machine learning model). This does not amount to significantly more than the exception itself (MPEP 2106.05(f))) and does not provide an inventive concept.
Claims 2 and 11 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Claim 2 recites the same limitations as claim 11.
Step 1: Is claim 11 to a process, machine, manufacture, or composition of matter?
Yes, claim 11 is directed to a method. Claim 2 is directed to a system.
Step 2A Prong 1: Does claim 11 recite an abstract idea, law of nature, or natural phenomenon?
The limitations of:
Performing pruning on the base neural network comprises performing structured single shot pruning (Mental process; A human can perform structured single shot pruning on a simple neural network)
Step 2A Prong 2: Does claim 11 recite elements that integrate the judicial exception into a practical application?
The base neural network has not been pre-trained; and
Performing pruning on the base neural network comprises performing structured single shot pruning (The limitations amount to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself (MPEP 2106.05(h))).
Step 2B: Does claim 11 recite elements that amount to significantly more than the judicial exception?
The base neural network has not been pre-trained; and
Performing pruning on the base neural network comprises performing structured single shot pruning (The limitations amount to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself (MPEP 2106.05(h)) and does not provide an inventive concept.
Claims 3 and 12 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Claim 3 recites the same limitations as claim 12.
Step 1: Is claim 12 to a process, machine, manufacture, or composition of matter?
Yes, claim 12 is directed to a method. Claim 3 is directed to a system.
Step 2A Prong 1: Does claim 12 recite an abstract idea, law of nature, or natural phenomenon?
The limitations of:
obtaining the pruned neural network includes, when pruning the base network, pruning in a channel unit for each layer; (Mental process; A human can obtain a pruned network by performing pruning in a channel unit.)
Step 2A Prong 2: Does claim 12 recite elements that integrate the judicial exception into a practical application?
obtaining the pruned neural network includes, when pruning the base network, pruning in a channel unit for each layer (The limitation amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself (MPEP 2106.05(h))).
Step 2B: Does claim 12 recite elements that amount to significantly more than the judicial exception?
obtaining the pruned neural network includes, when pruning the base network, pruning in a channel unit for each layer (The limitation amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself (MPEP 2106.05(h)) and cannot provide an inventive concept.
Claims 4 and 13 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Claim 4 recites the same limitations as claim 13.
Step 1: Is claim 13 to a process, machine, manufacture, or composition of matter?
Yes, claim 13 is directed to a method. Claim 4 is directed to a system.
Step 2A Prong 1: Does claim 13 recite an abstract idea, law of nature, or natural phenomenon?
The limitations of:
obtaining the pruned network includes, when pruning in the channel unit, maintaining a minimum channel remaining ratio for each layer (Mental process; A human can obtain a pruned network by performing pruning on a base network.)
Step 2A Prong 2: Does claim 13 recite elements that integrate the judicial exception into a practical application?
obtaining the pruned network includes, when pruning in the channel unit, maintaining a minimum channel remaining ratio for each layer (The limitation amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself (MPEP 2106.05(h))).
Step 2B: Does claim 13 recite elements that amount to significantly more than the judicial exception?
obtaining the pruned network includes, when pruning in the channel unit, maintaining a minimum channel remaining ratio for each layer (The limitation amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself (MPEP 2106.05(h)) and cannot provide an inventive concept.
Claims 5 and 14 are eligible under 35 U.S.C. 101 because the claimed invention is directed to statutory subject matter. Claim 5 recites the same limitations as claim 14.
Step 1: Is claim 14 to a process, machine, manufacture, or composition of matter?
Yes, claim 14 is directed to a method. Claim 5 is directed to a system.
Step 2A Prong 1: Does claim 14 recite an abstract idea, law of nature, or natural phenomenon?
The limitations of:
Setting a pruning ratio and the minimum channel remaining ratio,
inputting a mini-batch only once,
calculating gradients according to input values,
calculating a score for each weight using the calculated gradients,
calculating a score for each channel using the calculated score for each weight, and
pruning a channel having a score lower than a threshold set based on the pruning ratio
(These steps individually recite mental processes or mathematical calculations, but in a full sequence they are likely too complicated for a mental process. Therefore claim 14 does not recite an abstract idea)
Claims 6 and 15 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Claim 6 recites the same limitations as claim 15.
Step 1: Is claim 15 to a process, machine, manufacture, or composition of matter?
Yes, claim 15 is directed to a method. Claim 6 is directed to a system.
Step 2A Prong 1: Does claim 15 recite an abstract idea, law of nature, or natural phenomenon?
The limitations of:
obtaining the pruned network includes including a first layer as a layer of the pruned neural network when the pruning is performed such that a number of channels remains equal to or greater than the minimum channel remaining ratio with respect to the first layer (Mental process; A human can obtain a pruned network such that the number of channels is greater than or equal to the channel remaining ratio.)
Step 2A Prong 2: Does claim 15 recite elements that integrate the judicial exception into a practical application?
obtaining the pruned network includes including a first layer as a layer of the pruned neural network when the pruning is performed such that a number of channels remains equal to or greater than the minimum channel remaining ratio with respect to the first layer (The limitation amounts to the claim reciting only the idea of a solution or an outcome. We often see this when a generic machine learning model is used to perform the abstract idea. i.e., using a neural network to calculate a value 2106.05(f). This does not integrate the abstract idea into a practical application.)
Step 2B: Does claim 15 recite elements that amount to significantly more than the judicial exception?
obtaining the pruned network includes including a first layer as a layer of the pruned neural network when the pruning is performed such that a number of channels remains equal to or greater than the minimum channel remaining ratio with respect to the first layer (The limitation amounts to the claim reciting only the idea of a solution or an outcome. We often see this when a generic machine learning model is used to perform the abstract idea. i.e., using a neural network to calculate a value 2106.05(f). This does not amount to significantly more than the exception itself and cannot provide an inventive concept.)
Claims 7 and 16 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Claim 7 recites the same limitations as claim 16.
Step 1: Is claim 16 to a process, machine, manufacture, or composition of matter?
Yes, claim 16 is directed to a method. Claim 7 is directed to a system.
Step 2A Prong 1: Does claim 16 recite an abstract idea, law of nature, or natural phenomenon?
The limitations of:
obtaining the pruned network includes including a second layer as a layer of the pruned network after additionally assigning a number of channels to the second layer so as to be equal to or greater than the minimum channel remaining ratio when the pruning is performed such that the number of channels remains less than the minimum channel remaining ratio with respect to the second layer (Mental process; A human can obtain a pruned network that includes assigning a number of channels to the second layer that satisfy the minimum channel remaining ratio when pruning is performed such that the remaining channels fall below the channel remaining ratio.)
Step 2A Prong 2: Does claim 16 recite elements that integrate the judicial exception into a practical application?
obtaining the pruned network includes including a second layer as a layer of the pruned network after additionally assigning a number of channels to the second layer so as to be equal to or greater than the minimum channel remaining ratio when the pruning is performed such that the number of channels remains less than the minimum channel remaining ratio with respect to the second layer (The limitation amounts to the claim reciting only the idea of a solution or an outcome. We often see this when a generic machine learning model is used to perform the abstract idea. i.e., using a neural network to calculate a value 2106.05(f). This does not integrate an abstract idea into a practical application.)
Step 2B: Does claim 16 recite elements that amount to significantly more than the judicial exception?
The neural network learning method of claim 10, wherein obtaining the pruned network includes including a second layer as a layer of the pruned network after additionally assigning a number of channels to the second layer so as to be equal to or greater than the minimum channel remaining ratio when the pruning is performed such that the number of channels remains less than the minimum channel remaining ratio with respect to the second layer (The limitation amounts to the claim reciting only the idea of a solution or an outcome. We often see this when a generic machine learning model is used to perform the abstract idea. i.e., using a neural network to calculate a value 2106.05(f). This does not amount to significantly more than the exception itself and cannot provide an inventive concept.)
Claims 8 and 17 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Claim 8 recites the same limitations as claim 17.
Step 1: Is claim 17 to a process, machine, manufacture, or composition of matter?
Yes, claim 17 is directed to a method. Claim 8 is directed to a system.
Step 2A Prong 1: Does claim 17 recite an abstract idea, law of nature, or natural phenomenon?
The limitations of:
additionally assigning the number of channels to the second layer includes assigning, to the second layer, a channel having a higher score among channels pruned in the second layer (Mental process; A human can assign channels with higher scores among channels pruned in the second layer.)
Step 2A Prong 2: Does claim 17 recite elements that integrate the judicial exception into a practical application?
additionally assigning the number of channels to the second layer includes assigning, to the second layer, a channel having a higher score among channels pruned in the second layer (The limitation amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself (MPEP 2106.05(h))).
Step 2B: Does claim 17 recite elements that amount to significantly more than the judicial exception?
additionally assigning the number of channels to the second layer includes assigning, to the second layer, a channel having a higher score among channels pruned in the second layer (The limitation amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself (MPEP 2106.05(h)) and cannot provide an inventive concept.
Claims 9 and 18 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Claim 9 recites the same limitations as claim 18.
Step 1: Is the claim to a process, machine, manufacture, or composition of matter?
Yes, claim 18 is directed to a method. Claim 9 is directed to a system.
Step 2A Prong 1: Does claim 18 recite an abstract idea, law of nature, or natural phenomenon?
The limitations of:
the HPO uses one of random search, evolutionary optimization, or Bayesian optimization (Mathematical concepts).
Step 2A Prong 2: Does claim 18 recite elements that integrate the judicial exception into a practical application?
the HPO uses one of random search, evolutionary optimization, or Bayesian optimization (The limitation amounts to the claim reciting only the idea of a solution or an outcome. We often see this when a generic machine learning model is used to perform the abstract idea. i.e., using a neural network to calculate a value 2106.05(f). This does not integrate an abstract idea into a practical application)
Step 2B: Does claim 18 recite elements that amount to significantly more than the judicial exception?
the HPO uses one of random search, evolutionary optimization, or Bayesian optimization (The limitation amounts to the claim reciting only the idea of a solution or an outcome. We often see this when a generic machine learning model is used to perform the abstract idea. i.e., using a neural network to calculate a value 2106.05(f). This does not amount to significantly more than the exception itself and cannot provide an inventive concept.)
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or non-obviousness.
Claims 1-3 and 10-12 are rejected under 35 U.S.C. 103 as being unpatentable over Yang “Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer” in view of “Single Shot Structured Pruning Before Training” to Amersfoort et al. (hereby referred to as Amersfoort).
Regarding independent claim 1:
Yang teaches “A neural network learning apparatus comprising:” (Abstract, “A Pytorch implementation of our technique can be found at github.com/microsoft/mup and installable via pip install mup.” Yang premises that PyTorch, an open-source deep learning library, is used for implementing the Maximal Update Parameterization (MUP) package. This can be understood as a neural network apparatus).
“an optimizing unit configured to obtain an optimized hyperparameter set by performing hyperparameter optimization (HPO) a predetermined number of times…” (On pg 1, “We show that, in the recently discovered Maximal Update Parametrization (µP), many optimal HPs remain stable even as model size changes. This leads to a new HP tuning paradigm we call µTransfer: parametrize the target model in µP, tune the HP indirectly on a smaller model, and zero-shot transfer them to the full-sized model, i.e., without directly tuning the latter at all.” On pg 8, “To improve the reproducibility of our result: 1) we repeat the entire HP search process (a trial) 25 times for each setup.” Yang mentions tuning the HP on a smaller proxy model, and this can be understood as performing HPO. Yang also mentions repeating the entire HP search process 25 times to improve the reproducibility of results, and this can be understood as performing HPO a predetermined number of times).
“…a learning unit configured to train the base neural network using the optimized hyperparameter set” (On pg 1, “tune the HP indirectly on a smaller model, and zero-shot transfer them to the full-sized model, i.e., without directly tuning the latter at all.” Yang mentions tuning HP on a smaller proxy model before zero-transferring them to the full-sized model. Because the HP tuning is repeated 25 times for improving reproducibility of results, this can be understood as performing HPO and training a base neural network using the optimized hyperparameter set).
Yang does not expressly disclose, but Amersfoort does teach, “a pruning unit configured to obtain a pruned neural network by performing pruning on a base neural network;” (On pg 1, “In this work, we propose a structured and compute-aware Pruning Before Training (PBT) method. Structuring pruning methods remove entire channels in convolutional layers and hidden units in linear layers, leading to speed ups and reduced memory consumption on standard compute devices. Our method, Single Shot Structured Pruning (3SP), is easy to implement and has few hyper-parameters to tune.” Amersfoort classifies 3SP as a computationally cheaper pruning approach compared to the iterative and unstructured pruning approaches).
“…using the pruned neural network; and…” (Introduction pg 2, “we show how 3SP can be used to identify valuable data to acquire faster than a full model, allowing us to achieve better accuracy within a time-budget than an un-pruned model could.” Amersfoort lays out the motivation for using a pruned network, which is to achieve high accuracy under a limited budget).
Review of the specification [0035] shows that a unit may be constructed as software or hardware. Regarding the optimizing unit and learning unit, the abstract of Yang teaches “A Pytorch implementation of our technique can be found at github.com/microsoft/mup and installable via pip install mup.” PyTorch is an open-source deep learning library. The introduction of Yang further teaches “While large model training needs to be distributed across many GPUs, the small model tuning can happen on individual GPUs, greatly increasing the level of parallelism for tuning (and in the context of organizational compute clusters, better scheduling and utilization ratio).” Regarding the pruning unit, the introduction of Amersfoort teaches “Our method, Single Shot Structured Pruning (3SP), is easy to implement and has few hyper parameters to tune. The pruned model trains 2x faster (on a GPU) and performs inference 3x faster (on a CPU), with only a 0.5% loss in accuracy on CIFAR-10.”
Yang section 7 teaches “We perform HP tuning only on a smaller proxy model, test the obtained HPs on the large target model directly, and compare against baselines tuned using the target model.” Section 10 of Yang further teaches “Our approach is distinct from all of the above in that it does not work on the HP optimization process itself. Instead, it decouples the size of the target model from the tuning cost, which was not feasible prior to this work. This means that no matter how large the target model is, we can always use a fixed-sized proxy model to probe its HP landscape. Nevertheless, our method is complementary, as the above approaches can naturally be applied to the tuning of the proxy model.” Amersfoort section 3.1 teaches “SNIP defines M with the same shape as the weights, allowing it to turn off individual weights. We instead define M~ to remove entire operations. In particular, for convolutional layers, each output channel gets one binary mask variable governing its entire spatial extent. Linear layers have a binary mask per hidden unit; we visualize this in Figure 1. Masking, therefore, can be implemented by changing the shape of the weight tensors and is equivalent to using a smaller model, unlike unstructured methods.”
Yang and Amersfoort are analogous art because they are from the same field of endeavor, specifically reducing the computational cost of training and tuning deep neural networks. They are further reasonably pertinent to the particular problem involved, namely the prohibitive time and compute cost of hyperparameter optimization carried out on a full-size deep neural network.
It would therefore have been obvious to one having ordinary skill in the art at the time that the claimed invention was effectively filed to combine the teachings of Yang and Amersfoort. The motivation for doing so is given by Yang addressing the issue in the abstract, “Hyperparameter (HP) tuning in deep learning is an expensive process, prohibitively so for neural networks (NNs) with billions of parameters,” and Amersfoort providing a benefit in the introduction, “Our method, Single Shot Structured Pruning (3SP), is easy to implement and has few hyper parameters to tune. The pruned model trains 2x faster (on a GPU) and performs inference 3x faster (on a CPU), with only a 0.5% loss in accuracy on CIFAR-10.”
Note that independent claim 10 recites the same substantial subject matter as claim 1, only differing in embodiment. The difference in embodiment, a method as opposed to an apparatus, is an obvious variation of the other.
Regarding dependent claim 2:
Yang combined with Amersfoort discloses the limitations of claim 1. Amersfoort further teaches “the base neural network is a neural network that has not been pre-trained” (On pg 1, “We introduce a method to speed up training by 2x and inference by 3x in deep neural networks using structured pruning applied before training. Unlike previous works on pruning before training which prune individual weights, our work develops a methodology to remove entire channels and hidden units with the explicit aim of speeding up training and inference.” Amersfoort mentions using 3SP on a neural network before training, which can be understood as the neural network not being pre-trained).
Amersfoort further teaches “the pruning unit is configured to perform pruning on the base neural network by structured single-shot pruning” (In the abstract, “Unlike previous works on pruning before training which prune individual weights, our work develops a methodology to remove entire channels and hidden units with the explicit aim of speeding up training and inference.” On pg 2, “In structured pruning, only entire channels in convolutional layers and columns of linear layers can be removed (Figure 1). This is a significant restriction compared to unstructured pruning where any individual weight can be removed. However, structured pruning reduces the computational cost of training and evaluating a model, because when entire channels are removed the size of the activations is also reduced, leading to a smaller model.” Amersfoort explains what 3SP is and how it differs from iterative pruning).
Note that dependent claim 11 recites the same substantial subject matter as claim 2, only differing in embodiment. The difference in embodiment, a method as opposed to an apparatus, is an obvious variation of the other.
Regarding dependent claim 3:
Yang combined with Amersfoort discloses the limitations of claim 1. Amersfoort further teaches “the pruning unit is configured to, when performing pruning on the base neural network, perform the pruning in a channel unit for each layer” (Section 3.1, “SNIP defines M with the same shape as the weights, allowing it to turn off individual weights. We instead define M~ to remove entire operations. In particular, for convolutional layers, each output channel gets one binary mask variable governing its entire spatial extent. Linear layers have a binary mask per hidden unit; we visualize this in Figure 1. Masking, therefore, can be implemented by changing the shape of the weight tensors and is equivalent to using a smaller model, unlike unstructured methods.” Amersfoort specifically mentions defining M~ to remove entire operations, as well as each output channel getting a binary mask. This can be understood as pruning being performed in a channel unit for each layer).
Note that dependent claim 12 recites the same substantial subject matter as claim 3, only differing in embodiment. The difference in embodiment, a method as opposed to an apparatus, is an obvious variation of the other.
Claims 4 and 13 are rejected under 35 U.S.C. 103 as being unpatentable over Yang and Amersfoort as applied to claims 3 and 12 above, and further in view of “Revisiting Random Channel Pruning for Neural Network Compression” to Li et al. (hereby referred to as Li).
Regarding dependent claim 4:
Yang combined with Amersfoort discloses the limitations of claim 3. Li further teaches “the pruning unit is configured to, when performing the pruning in the channel unit, maintain a minimum channel remaining ratio for each layer” (In section 4, “The importance score is the indicator of which channels should be pruned in the next step. Then we select a number of sub-architectures and prune the channels with the lowest score. A sub-architecture is formed by sampling pruning ratios for each layer separately, and then pruning the number of channels given by the ratio.” In section 5.1, “The minimum number of channels is reset and rounded to multiples of 8. This again avoids the very narrow bottleneck in the network. For example, when pruning 30% of the FLOPs of ResNet-18, we empirically require 40% of the channels must be kept.” Li mentions the minimum channel number is reset and rounded to multiples of 8, and then gives an example where 40% of channels must be kept. This can be understood as maintaining a minimum channel remaining ratio for each layer, since the goal is avoiding a narrow bottleneck in the network).
Yang, Amersfoort and Li are analogous art because they are from the same field of endeavor, specifically reducing the computational cost of training and accelerating the inference of deep neural networks through pruning or smaller proxy networks. They are further reasonably pertinent to the particular problem involved, namely the prohibitive time and compute cost of hyperparameter optimization carried out on a full-size deep neural network.
It would therefore have been obvious to one having ordinary skill in the art at the time that the claimed invention was effectively filed to combine the teachings of Yang, Amersfoort, and Li. The motivation for doing so is given by Li in section 4.1, “A bottleneck in the network could harm the performance of the pruned network,” and in section 5.1, “The minimum number of channels is reset and rounded to multiples of 8. This again avoids the very narrow bottleneck in the network. For example, when pruning 30% of the FLOPs of ResNet-18, we empirically require 40% of the channels must be kept.”
Note that dependent claim 13 recites the same substantial subject matter as claim 4, only differing in embodiment. The difference in embodiment, a method as opposed to an apparatus, is an obvious variation of the other.
Claims 5 and 14 are rejected under 35 U.S.C. 103 as being unpatentable over Yang, Amersfoort and Li as applied to claims 4 and 13 above, in view of “SNIP: Single-Shot Network Pruning Based on Connection Sensitivity” to Lee et al. (hereby referred to as Lee), “Optimizing the Deep Neural Networks by Layer-Wise Refined Pruning and the Acceleration on FPGA” to Hengyi at al. (hereby referred to as Hengyi), and further in view of “Importance Estimation for Neural Network Pruning” to Molchanov et al. (hereby referred to as Molchanov).
Regarding dependent claim 5:
Yang combined with Amersfoort and Li discloses the limitations of claim 4. Li further teaches “wherein the pruning unit is configured to: set a pruning ratio and the minimum channel remaining ratio,” (In section 4, “A sub architecture is formed by sampling pruning ratios for each layer separately, and then pruning the number of channels given by the ratio. A minimum ratio of remaining channels is set. That is, the range for sampling the pruning ratio is [η,1].” In section 4.1, “Specifically, let C_prune and C_orig denote the floating point operations (FLOPs) of the pruned network and the original network, respectively. Then the samples that meet the following criteria are kept, i.e. | C_prune / C_orig | −γ <=T, where γ is the overall pruning ratio of the network and T is the threshold that confines the difference between the actual and target pruning ratio.” Li explicitly mentions the minimum ratio of remaining channels is set. Li gives a range from while a pruning ratio is sampled. This can be understood as the pruning unit configured to set a pruning ratio and the minimum channel remaining ratio).
Yang combined with Amersfoort and Li do not expressly disclose, but Hengyi does teach, “prune a channel having a score lower than a threshold set based on the pruning ratio” (Section 3.3, “Second, the layer-wise pruning ratios are calculated,” “Third, the redundant channels that may be pruned are identified,” and “the thresholds, T[ T0,T1,T2 ,..., TN ] where Ti is defined as the pruning threshold with regard to the parameters γ of the ith BN layer, are obtained by multiplying the pruning ratio with the number of channels N. Based on the set of scalars within range of the pruning ratio that is smaller than the threshold, the correlated channels of the BN layer are identified as redundant.” Hengyi describes a process of setting pruned ratios and thresholds, then identifying redundant channels based on the set of scalars within range of the pruning ratio that is smaller than the threshold. This can be understood as pruning a channel having a score lower than a threshold set based on the pruning ratio).
Yang combined with Amersfoort, Li and Hengyi do not expressly disclose, but Lee does teach, “input a mini-batch only once, calculate gradients according to input values, calculate a score for each weight using the calculated gradients,” (Abstract pg 1, “We present a new approach that prunes a given network once at initialization prior to training.” Algorithm 1 step 2, “Sample a mini-batch of training data.” Section 4.1, “We take the magnitude of the derivatives g_j as the saliency criterion. Note if the magnitude of the derivative is high (regardless of the sign), it essentially means that the connection c_j has a considerable effect on the loss (either positive or negative), and it has to be preserved to allow learning on w_j. Based on this hypothesis, we define connection sensitivity as the normalized magnitude of the derivatives: [Equation 6].” Algorithm 1 step 3, “Connection sensitivity.” Lee mentions calculating the magnitude of the derivatives as well as how high magnitudes impact learning for w_j, and the connection sensitivity is defined as the normalized magnitude. This can be understood as calculating gradients according to input values and calculating a score for each weight using the gradients).
Yang combined with Amersfoort, Li, Hengyi and Lee do not expressly disclose, but Molchanov does teach, the limitation “calculate a score for each channel using the calculated score for each weight, and” (On pg 3, “To approximate the joint importance of a structural set of parameters W_s e.g. a convolutional filter, we have two alternatives. We define it as a group contribution: [Equation 7], or alternatively, sum the importance of the individual parameters in the set, [Equation 8].” Molchanov defines two methods for approximating the joint importance of a convolutional filter using its set of weights. From equations 7 and 8, this can be understood as calculating a score for each channel using the calculated score for each weight).
Yang combined with Amersfoort, Li, Hengyi, Lee and Molchanov are analogous art because they are from the same field of endeavor, specifically reducing the computational cost of training and tuning deep neural networks through techniques such as using smaller proxy networks or pruning larger networks. They are further reasonably pertinent to the particular problem involved, namely the prohibitive time and compute cost of hyperparameter optimization carried out on a full-size deep neural network.
It would therefore have been obvious to one having ordinary skill in the art at the time that the claimed invention was effectively filed to combine the teachings of Yang, Amersfoort, Li, Hengyi, Lee, and Molchanov. The motivation for doing so is given by Molchanov on pg 3, which is “To approximate the joint importance of a structural set of parameters W_s e.g. a convolutional filter, we have two alternatives. We define it as a group contribution: [Equation 7], or alternatively, sum the importance of the individual parameters in the set, [Equation 8],” and in the abstract, “We propose a novel method that estimates the contribution of a neuron (filter) to the final loss.”
Note that dependent claim 14 recites the same substantial subject matter as claim 5, only differing in embodiment. The difference in embodiment, a method as opposed to an apparatus, is an obvious variation of the other.
Claims 9 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Yang and Amersfoort as applied to claims 1 and 10 above, and further in view of “Practical Bayesian Optimization of Learning Algorithms” to Snoek et al. (hereby referred to as Snoek).
Regarding dependent claim 9:
Yang combined with Amersfoort discloses the limitations of claim 1. Snoek further teaches “wherein the optimizing unit is configured to perform HPO using one of random search, evolutionary optimization, or Bayesian optimization” (Abstract, “Machine learning algorithms frequently require careful tuning of model hyperparameters, regularization terms, and optimization parameters. Unfortunately, this tuning is often a “black art” that re quires expert experience, unwritten rules of thumb, or sometimes brute-force search.” Introduction pg 1, “In this setting where function evaluations are expensive, it is desirable to spend computational time making better choices about where to seek the best parameters. Bayesian optimization (Mockus et al., 1978) provides an elegant approach and has been shown to outperform other state of the art global optimization algorithms on a number of challenging optimization benchmark functions (Jones, 2001).” Snoek explicitly mentions HP using Bayesian optimization, which can be understood as HPO using one of random search, evolutionary optimization, or Bayesian optimization).
Yang, Amersfoort and Snoek are analogous art because they are from the same field of endeavor, specifically reducing the computational cost of training and tuning deep neural networks. They are further reasonably pertinent to the particular problem involved, namely the prohibitive time and compute cost of hyperparameter optimization carried out on a deep neural network.
It would therefore have been obvious to one having ordinary skill in the art at the time that the claimed invention was effectively filed to combine the teachings of Yang, Amersfoort, and Snoek. The motivation for doing so is given by Snoek in the introduction, “In this setting where function evaluations are expensive, it is desirable to spend computational time making better choices about where to seek the best parameters. Bayesian optimization (Mockus et al., 1978) provides an elegant approach and has been shown to outperform other state of the art global optimization algorithms on a number of challenging optimization benchmark functions (Jones, 2001).”
Note that dependent claim 18 recites the same substantial subject matter as claim 9, only differing in embodiment. The difference in embodiment, a method as opposed to an apparatus, is an obvious variation of the other.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to TALHA IBN MUHIB whose telephone number is (571)270-7331. The examiner can normally be reached Mon-Fri (8AM - 5PM) ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Miranda Huang can be reached at (571) 270-7092. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/TALHA IBN MUHIB/Examiner, Art Unit 2124
/ALAN CHEN/Primary Examiner, Art Unit 2125