Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Remarks
This Office Action is responsive to Applicants' Amendment filed on July 7, 2026, in which claims 1, 8, 11, 12, and 18 are currently amended. Claims 1-4,6-14 and 16-19 are currently pending.
Response to Arguments
Applicant’s arguments with respect to rejection of claims 1-4,6-14 and 16-19 under 35 U.S.C. 101 based on amendment have been considered, however, are not persuasive.
With respect to Applicant's arguments on p. 8 of the Remarks submitted 7/7/2026 that the amendments "remove elements that may be performed in the mind to some degree", Examiner respectfully disagrees. Claim 1, for example, is still entirely directed towards determining "a neural architecture search space subject to one or more design constraints" which is a mental process which can be performed entirely in the mind with or without the assistance of tools such as pen and paper or a generic computer "executing computer-readable instructions by one or more processors of a computing device". With respect to Applicant's arguments that the human mind is not equipped to perform calculations regarding the performance, Examiner notes that said calculations in claim 1 are "a plurality of gradient functions" which is a mathematical calculation which is itself a judicial exception that one of ordinary skill in the art could perform in the mind with the assistance of tools such as pen and paper as a mental process. There is nothing recited in the claims that would limit the scope of the judicial exception in such a way that would make it impractical to perform entirely in the mind.
With respect to Applicant's arguments on p. 9 of the Remarks submitted 7/7/2026 that "the pending claims now more clearly reflect an improvement to the functioning of a computer or to another technology or technical field", Examiner respectfully disagrees. As noted above, the claims are directed towards a judicial exception (2106.05(a) "It is important to note, the judicial exception alone cannot provide the improvement") performed on a generic computer system which does not integrate the judicial exception into a practical application (MPEP 2106.07(a)(II) "employing well-known computer functions to execute an abstract idea, even when limiting the use of the idea to one particular environment, does not integrate the exception into a practical application"). For these reasons Examiner asserts that the rejection under 35 USC 101 is appropriate and should be maintained.
Applicant’s arguments with respect to rejection of claims 1-4,6-14 and 16-19 under 35 U.S.C. 102/103 based on amendment have been considered and are persuasive. The argument is moot in view of a new ground of rejection set forth below.
Claim Rejections - 35 USC § 101
101 Rejection
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-4, 6-14, 16-19 are rejected under 35 USC § 101 because the claimed invention is directed to non-statutory subject matter.
Regarding Claim 1: Claim 1 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1 Analysis: Claim 1 is directed to a method, which is directed to a process, one of the statutory categories.
Step 2A Prong One Analysis: Claim 1 under its broadest reasonable interpretation is a series of mental processes. For example, but for the generic computer components language, the above limitations in the context of this claim encompass neural network processing, including the following:
determine a neural architecture search space subject to one or more design constraints based, at least in part, on hardware computing resources for execution of a neural network, (observation, evaluation, and judgement. One of ordinary skill in the art could readily determine a neural architecture search space subject to design constraints. For a non-exhaustive example, if given an image processing task as motivation, one of ordinary skill in the art would readily constrain the search space to convolutional neural network models entirely in the mind without the assistance of tools such as pen and paper)
to construct a supernet based on the neural architecture search space, (observation, evaluation, and judgement. Given the previous non-exhaustive example of image processing task, one of ordinary skill in the art could readily construct the supernet architecture entirely in the mind with or without the assistance of tools such as pen and paper. As a non-limiting example: a supernet comprising three convolutional layers followed by a fully connected layer. This example was generated entirely in the mind without the assistance of any tools)
to perform a neural architecture search employing a plurality of gradient functions, each of the plurality of gradient functions associated with one of a plurality of design parameters over the supernet for selection of a neural network architecture subsequent to construction of the supernet, the plurality of design parameters comprising convolutional neural network channel pruning, quantization, and at least one additional design parameter (observation, evaluation, and judgement. As one non-limiting example, this could amount to simply selecting a training routine hyperparameter i.e. SGD, ADAM, Momentum, etc. all which employ a plurality of gradient functions. This search/hyperparameter selection can readily be performed entirely in the mind)
Therefore, claim 1 recites an abstract idea which is a judicial exception.
Step 2A Prong Two Analysis: Claim 1 recites additional elements “executing computer-readable instructions by one or more processors of a computing device to”. However, these additional features are computer components recited at a high-level of generality, such that they amount to no more than mere instructions to apply the judicial exception using a generic computer component. An additional element that merely recites the words “apply it” (or an equivalent) with the judicial exception, or merely includes instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea, does not integrate the judicial exception into a practical application (See MPEP 2106.05(f)). Therefore, claim 1 is directed to a judicial exception.
Step 2B Analysis: Claim 1 does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to the lack of integration of the abstract idea into a practical application, the additional elements recited in claim 1 amount to no more than mere instructions to apply the judicial exception using a generic computer component.
For the reasons above, claim 1 is rejected as being directed to non-patentable subject matter under §101. This rejection applies equally to dependent claims 2-10. The additional limitations of the dependent claims are addressed briefly below:
Dependent claim 2 recites additional observation, evaluation, and judgement “determine the neural architecture search space based, at least in part, on combinations of activation bit width and weight bit width to be implemented in one or more layers of the neural network.” And “quantify costs associated with the combinations of activation bit width and weight bit width”
Dependent claim 3 recites additional observation, evaluation, and judgement “determine the neural architecture search space further based, at least in part, on candidate channel sizes for at least one of the one or more layers of the neural network based, at least in part, on the combinations of activation bit width and weight bit width” and “quantify costs associated with the candidate channel sizes based, at least in part, on the quantified costs associated with the combinations of activation bit width and weight bit width”
Dependent claim 4 recites additional instructions to apply the judicial exception using generic computer components “at least one of the one or more layers of the neural network comprises a convolution layer to be implemented at least in part by application of a kernel” as well as additional observation, evaluation, and judgement “determine the neural architecture search space further based, at least in part, on available kernel sizes for the kernel”.
Dependent claim 6 recites additional observation, evaluation, and judgement “determine the neural architecture search space further based, at least in part, on candidate operator types for the at least one of the one or more layers of the neural network based, at least in part, on the combinations of activation bit width and weight bit width; and quantify costs associated with the candidate operator types based, at least in part, on the quantified costs associated with the candidate channel sizes”
Dependent claim 7 recites additional observation, evaluation, and judgement “determine a union of candidate design options over the candidate operator types for the at least one of the one or more layers of the neural network; and determine a union of the candidate design options over the combinations of activation bit width and weight bit width based, at least in part, on the determined union of the candidate design options over the candidate operator types for the at least one of the one or more layers of the neural network”
Dependent claim 8 recites additional observation, evaluation, and judgement “determine candidate design options for implementation of the neural network based, at least in part, on the hardware computing resources and the one or more design constraints” as well as additional insignificant extra-solution activity of gathering and outputting data (See MPEP 2106.05(g)) “express and/or structure the candidate design options as the neural architecture search space in a non-transitory storage medium” (expressing the candidate design options in a non-transitory storage medium interpreted as storing in memory) which is well-understood, routine, and conventional in the art (See MPEP 2106.05(d)(II)(i))
Dependent claim 9 recites additional observation, evaluation, and judgement “execute the NAS process to select a design option from the neural architecture search space to implement the neural network”
Dependent claim 10 recites additional observation, evaluation, and judgement “wherein at least one of the one or more design constraints are defined by execution latency, operation count, model size, power consumption or memory usage, or a combination thereof”
Regarding Claim 11: Claim 11 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1 Analysis: Claim 11 is directed to a method, which is directed to a process, one of the statutory categories.
Step 2A Prong One Analysis: Claim 11 under its broadest reasonable interpretation is a series of mental processes. For example, but for the generic computer components language, the above limitations in the context of this claim encompass neural network processing, including the following:
determine a neural architecture search space subject to one or more design constraints based, at least in part, on hardware computing resources for execution of a neural network, (observation, evaluation, and judgement. One of ordinary skill in the art could readily determine a neural architecture search space subject to design constraints. For a non-exhaustive example, if given an image processing task as motivation, one of ordinary skill in the art would readily constrain the search space to convolutional neural network models entirely in the mind without the assistance of tools such as pen and paper)
to construct a supernet based on the neural network architecture search space, (observation, evaluation, and judgement. Given the previous non-exhaustive example of image processing task, one of ordinary skill in the art could readily construct the supernet architecture entirely in the mind with or without the assistance of tools such as pen and paper. As a non-limiting example: a supernet comprising three convolutional layers followed by a fully connected layer. This example was generated entirely in the mind without the assistance of any tools)
to perform a neural architecture search employing a plurality of gradient functions, each of the plurality of gradient functions associated with one of a plurality of design parameters, over the supernet for selection of a neural network architecture subsequent to construction of the supernet, the plurality of design parameters comprising convolutional neural network channel pruning, quantization, and at least one additional design parameter (observation, evaluation, and judgement. As one non-limiting example, this could amount to simply selecting a training routine hyperparameter i.e. SGD, ADAM, Momentum, etc. all which employ a plurality of gradient functions. This search/hyperparameter selection can readily be performed entirely in the mind)
Therefore, Claim 11 recites an abstract idea which is a judicial exception.
Step 2A Prong Two Analysis: Claim 11 recites additional elements “A computing device, comprising: one or more processors to”. However, these additional features are computer components recited at a high-level of generality, such that they amount to no more than mere instructions to apply the judicial exception using a generic computer component. An additional element that merely recites the words “apply it” (or an equivalent) with the judicial exception, or merely includes instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea, does not integrate the judicial exception into a practical application (See MPEP 2106.05(f)). Claim 11 also recites additional elements “express and/or structure the neural network architecture search space in a non-transitory storage medium” (expressing the candidate design options in a non-transitory storage medium interpreted as storing in memory) which amounts to gathering and outputting data (See MPEP 2106.05(g)). Therefore, Claim 11 is directed to a judicial exception.
Step 2B Analysis: Claim 11 does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to the lack of integration of the abstract idea into a practical application, the additional elements recited in Claim 11 amount to no more than mere instructions to apply the judicial exception using a generic computer component.
For the reasons above, Claim 11 is rejected as being directed to non-patentable subject matter under §101. This rejection applies equally to dependent claims 12-17. The additional limitations of the dependent claims are addressed briefly below:
Dependent claim 12 recites additional observation, evaluation, and judgement “determine the neural network architecture search space based, at least in part, on combinations activation bit width and weight bit width to be implemented in one or more layers of the neural network-based inference engine; and quantify costs associated with the combinations of activation bit width and weight bit width”
Dependent claim 13 recites additional observation, evaluation, and judgement “determine the neural network architecture search space further based, at least in part, on candidate channel sizes for at least one of the one or more layers of the neural network-based inference engine based, at least in part, on the combinations of activation bit width and weight bit width; and quantify costs associated with the candidate channel sizes based, at least in part, on the quantified costs associated with the combinations of activation bit width and weight bit width”
Dependent claim 4 recites additional instructions to apply the judicial exception using generic computer components “wherein at least one of the one or more layers of the neural network-based inference engine comprises a convolution layer to be implemented at least in part by application of a kernel, and wherein the one or more processors are further to” as well as additional observation, evaluation, and judgement “determine the neural network architecture search space further based, at least in part, on available kernel sizes for the kernel”
Dependent claim 16 recites additional observation, evaluation, and judgement “determine the neural network architecture search space further based, at least in part, on candidate operator types for the at least one of the one or more layers of the neural network-based inference engine based, at least in part, on the combinations of activation bit width and weight bit width; and quantify costs associated with the candidate operator types based, at least in part, on the quantified costs associated with the candidate channel sizes”
Dependent claim 7 recites additional observation, evaluation, and judgement “wherein at least one of the one or more design constraints are defined by execution latency, operation count, model size, power consumption or memory usage, or a combination thereof”
Regarding Claim 18: Claim 18 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1 Analysis: Claim 18 is directed to a method, which is directed to a process, one of the statutory categories.
Step 2A Prong One Analysis: Claim 18 under its broadest reasonable interpretation is a series of mental processes. For example, but for the generic computer components language, the above limitations in the context of this claim encompass neural network processing, including the following:
to execute a neural network architecture search process to select a design option for a neural network-based inference engine from a plurality of candidate design options expressed and/or structured as a neural network architecture search space in a non-transitory storage medium, the plurality of candidate design options defining a neural network architecture search space having been determined based, at least in part, on (observation, evaluation, and judgement),
determination of a neural network architecture search space subject to one or more design constraints by application of the one or more design constraints to an identification of hardware computing resources based, at least in part, on hardware computing resources for execution of a neural network, and to construct a supernet based on the neural network architecture search space, and to perform a neural architecture search employing a plurality of gradient functions, each of the plurality of gradient functions associated with one of a plurality of design parameters, over the supernet for selection of the neural network architecture subsequent to construction of the supernet, the plurality of design parameters comprising convolutional neural network channel pruning, quantization, and at least one additional design parameter (observation, evaluation, and judgement)
Therefore, Claim 18 recites an abstract idea which is a judicial exception.
Step 2A Prong Two Analysis: Claim 18 recites additional elements “executing computer-readable instructions by one or more processors of a computing device”. However, these additional features are computer components recited at a high-level of generality, such that they amount to no more than mere instructions to apply the judicial exception using a generic computer component. An additional element that merely recites the words “apply it” (or an equivalent) with the judicial exception, or merely includes instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea, does not integrate the judicial exception into a practical application (See MPEP 2106.05(f)). Therefore, Claim 18 is directed to a judicial exception.
Step 2B Analysis: Claim 18 does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to the lack of integration of the abstract idea into a practical application, the additional elements recited in Claim 18 amount to no more than mere instructions to apply the judicial exception using a generic computer component.
For the reasons above, Claim 18 is rejected as being directed to non-patentable subject matter under §101. This rejection applies equally to dependent claim 19. The additional limitations of the dependent claims are addressed briefly below:
Dependent claim 19 recites additional observation, evaluation, and judgement “wherein at least one of the one or more design constraints are defined by execution latency, operation count, model size, power consumption or memory usage, or a combination thereof”
Therefore, when considering the elements separately and in combination, they do not add significantly more to the inventive concept. Accordingly, claims 1-19 are rejected under 35 U.S.C. § 101.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 1, 2, 8-12, and 17-19 are rejected under U.S.C. §103 as being unpatentable over the combination of Hu (“TF-NAS: Rethinking Three Search Freedoms of Latency-Constrained Differentiable Neural Architecture Search”, 2020) and Benmeziane (“A Comprehensive Survey on Hardware-Aware Neural Architecture Search”, 2021).
PNG
media_image1.png
468
1068
media_image1.png
Greyscale
FIG. 1 of Hu
Regarding claim 1, Hu teaches A method of generating a neural architecture search space for a neural architecture search (NAS) process, comprising:([p. 1] "Abstract. With the flourish of differentiable neural architecture search (NAS), automatically searching latency-constrained architectures gives a new perspective to reduce human labor and expertise. However, the searched architectures are usually suboptimal in accuracy and may have large jitters around the target latency. In this paper, we rethink three freedoms of differentiable NAS, i.e. operation-level, depth-level and width level, and propose a novel method, named Three-Freedom NAS (TF NAS), to achieve both good classification accuracy and precise latency constraint")
executing computer-readable instructions by one or more processors of a computing device([p. 10] "All the experiments are conducted on ImageNet [10] under the mobile set ting. Similar with [3], the latency is measured with a batch size of 32 on a Titan RTX GPU. We set the number of threads for OpenMP to 1 and use Py torch1.1+cuDNN7.6.0 […] This procedure takes about 1.8 days on 1 Titan RTX GPU")
to determine a neural architecture search space ([p. 4 §3.2] "we build a layer-wise search space, which is depicted in Fig. 1 and Tab. 1.")
subject to one or more design constraints based, at least in part, on hardware computing resources for execution of a neural network, ([p. 2 §1] "Our TF-NAScan search architectures with precise latency on target devices" [p. 14 §5] "we have proposed Three-Freedom NAS (TF-NAS) to seek an architecture with good accuracy as well as precise latency on the target devices.")
and to construct a supernet based on the neural architecture search space, ([p. 4] "ω and α are the supernet weights and the architecture distribution parameters, respectively. Given a supernet A, we aim to search a subnet α∗ ∈ A that minimizes the validation loss Lval (ω∗,α) and the latency constraint C (LAT(α)), where the weights ω∗ of supernet are obtained by minimizing the training loss Ltrain (ω,α) and λ is a trade-off hyperparameter" See also Table 1 "Macro architecture of the supernet")
and to perform a neural architecture search employing a plurality of gradient functions, each of the plurality of gradient functions associated with one of a plurality of design parameters over the supernet for selection of a neural network architecture ([p. 4] "Sampling a subnet from supernet A is a non-differentiable process w.r.t. the architecture distribution parameters α. Therefore, a continuous relaxation is needed to allow back-propagation. Assuming there are N operations to be searched in each layer, we define opl i and αl i as the i-th operation in layer l and its architecture distribution parameter, respectively [...] the gradients of all the αli can be back-propagated through Eq. (3).")
subsequent to construction of the supernet,([p. 9] "Algorithm 1 […] from the supernet A" Hu explicitly discloses that TF-NAS finds latency-constrained architectures from the supernet. Algorithm 1 also derives a seed network for the supernet and puts it back into the supernet, and the experiments section states the supernet is trained while architecture-distribution parameters are updated, directly placing the search over the supernet.).
However, Hu does not explicitly teach the plurality of design parameters comprising convolutional neural network channel pruning, quantization, and at least one additional design parameter..
Benmeziane, in the same field of endeavor, teaches the plurality of design parameters comprising convolutional neural network channel pruning, quantization, and at least one additional design parameter.([p. 1] "This kind of NAS, called hardware-aware NAS (HW-NAS), makes searching the most efficient architecture more complicated and opens several questions" [p. 3] "we introduce other considerations that either tackle the hardware efficiency objective, model compressions such as automatic quantization and pruning or add other objectives to the NAS like the robustness against adversarial attacks" [p. 14] "Gumbel Softmax: [156] One way to relax the discrete variables is to use the Gumbel softmax function. It helps insert some random noise following the Gumbel distribution so that the gradient computation is possible" [p. 17] "mixed-precision quantization that applies different bitwidth values for different layers in the same network is more commonly used [...] X. Dong and Y. Yang in [176] propose to prune the over parameterized network without performance damage. They directly search within their NAS process for a network with a flexible channel and layer sizes" See also p. 4 "Other Considerations for HW-NAS" explicitly including Quantization, pruning, and security and reliability with implementation details. Benmeziane explicitly describes automatic mixed-precision quantization as NAS that searches layer bitwidths and explicitly recites gradient based convolutional channel pruning methods by Dong and Yang.).
Hu as well as Benmeziane are directed towards hardware aware neural architecture search. Therefore, Hu as well as Benmeziane are analogous art in the same field of endeavor. It would have been obvious before the effective filing date of the claimed invention to combine the teachings of Hu with the teachings of Benmeziane by including channel pruning and quantization among other design parameters searched using Hu's gradient-based supernet approach. Benmeziane provides as additional motivation for combination ([p. 23] “HW-NAS offers a paradigm that opens up the design space and pushes forward the Pareto frontier between hardware efficiency and model accuracy for efficient and improved hardware/software co-design, hence pushing AI to its next frontiers”). This motivation for combination also applies to the remaining claims which depend on this combination.
Regarding claim 2, the combination of Hu and Benmeziane teaches The method of claim 1, and further comprising executing computer-readable instructions by one or more processors of the computing device to: determine the neural architecture search space based, at least in part, on combinations of activation bit width and weight bit width to be implemented in one or more layers of the neural network; and(Benmeziane [p. 17] "Recently, mixed-precision quantization that applies different bitwidth values for different layers in the same network is more commonly used. HAQ [62] implemented a dedicated reinforcement learning agent that learns to assign the right bitwidth to each layer. Their goal is to specialize the architecture to a specific hardware platform by incorporating hardware constraints and accuracy into the reward function. For each layer, the agent takes two decisions: one for the weights and another one for the activation")
quantify costs associated with the combinations of activation bit width and weight bit width.(Benmeziane [p. 17] "Recently, mixed-precision quantization that applies different bitwidth values for different layers in the same network is more commonly used. HAQ [62] implemented a dedicated reinforcement learning agent that learns to assign the right bitwidth to each layer. Their goal is to specialize the architecture to a specific hardware platform by incorporating hardware constraints and accuracy into the reward function. For each layer, the agent takes two decisions: one for the weights and another one for the activation" Benmeziane is explicit that HAQ quantifies reward/cost associated with the combinations of activation and weight bit widths).
Regarding claim 8, the combination of Hu and Benmeziane teaches The method of claim 1, and further comprising executing computer-readable instructions by the one or more processors of the computing device to: determine candidate design options for implementation of the neural network based, at least in part, on the identified hardware computing resources and the one or more design constraints;(Hu [p. 4] "we focus on latency-constrained macro search" [p. 8] "Due to the coarse-grained search space for latency, some target latency cannot be precisely satisfied, leading to instability during architecture searching. For example, setting the target latency to be 15ms, we search two architectures: one is 14.32ms and the other is 15.76ms. Both of them have around 0.7ms gaps for the target latency." [p. 23] "Each candidate operation has a kernel size k=3 or k=5 and a continuous expansion ratio eE[2,4] or eE[4,8]")
express and/or structure the candidate design options as the neural architecture search space in a non-transitory storage medium.(Hu [p. 4] "we build a layer-wise search space, which is depicted in Fig. 1 and Tab. 1" [p. 10] "Overall Algorithm. Our Three-Freedom NAS (TF-NAS) contains all above components: the bi-sampling search algorithm, the sink-connecting search space and the elasticity-scaling strategy. It finds latency-constrained architectures from the supernet").
Regarding claim 9, the combination of Hu and Benmeziane teaches The method of claim 1, and further comprising executing computer-readable instructions by the one or more processors of the computing device to: execute the NAS process to select a design option from the neural architecture search space to implement the neural network.(Hu [p. 10] "Overall Algorithm. Our Three-Freedom NAS (TF-NAS) contains all above components: the bi-sampling search algorithm, the sink-connecting search space and the elasticity-scaling strategy. It finds latency-constrained architectures from the supernet (Tab. 1) by solving the following bi-level problem [...] the best architecture is derived from the supernet based on α and β, where the strongest operation in each layer and the strongest depth in each stage are chosen.").
Regarding claim 10, the combination of Hu and Benmeziane teaches The method of claim 1, wherein at least one of the one or more design constraints are defined by execution latency, operation count, model size, power consumption or memory usage, or a combination thereof.(Hu [p. 1 §1] "specific resource constraints (e.g. FLOPs, latency, energy) is more important in practice" [p. 3] "the number of parameters, FLOPs and latency for neural architecture search" [p. 8] "Due to the coarse-grained search space for latency, some target latency cannot be precisely satisfied, leading to instability during architecture searching. For example, setting the target latency to be 15ms, we search two architectures: one is 14.32ms and the other is 15.76ms. Both of them have around 0.7ms gaps for the target latency.").
Regarding claim 11, Hu teaches A computing device, comprising: one or more processors to:([p. 1] "Abstract. With the flourish of differentiable neural architecture search (NAS), automatically searching latency-constrained architectures gives a new perspective to reduce human labor and expertise. However, the searched architectures are usually suboptimal in accuracy and may have large jitters around the target latency. In this paper, we rethink three freedoms of differentiable NAS, i.e. operation-level, depth-level and width level, and propose a novel method, named Three-Freedom NAS (TF NAS), to achieve both good classification accuracy and precise latency constraint")
determine a neural network architecture search space subject to one or more design constraints based, at least in part, on hardware computing resources for execution of a neural network, ([p. 2 §1] "Our TF-NAScan search architectures with precise latency on target devices" [p. 14 §5] "we have proposed Three-Freedom NAS (TF-NAS) to seek an architecture with good accuracy as well as precise latency on the target devices.")
and to construct a supernet based on the neural architecture search space, ([p. 4] "ω and α are the supernet weights and the architecture distribution parameters, respectively. Given a supernet A, we aim to search a subnet α∗ ∈ A that minimizes the validation loss Lval (ω∗,α) and the latency constraint C (LAT(α)), where the weights ω∗ of supernet are obtained by minimizing the training loss Ltrain (ω,α) and λ is a trade-off hyperparameter" See also Table 1 "Macro architecture of the supernet")
and to perform a neural architecture search employing a plurality of gradient functions, each of the plurality of gradient functions associated with one of a plurality of design parameters, ([p. 4] "Sampling a subnet from supernet A is a non-differentiable process w.r.t. the architecture distribution parameters α. Therefore, a continuous relaxation is needed to allow back-propagation. Assuming there are N operations to be searched in each layer, we define opl i and αl i as the i-th operation in layer l and its architecture distribution parameter, respectively [...] the gradients of all the αli can be back-propagated through Eq. (3).")
over the supernet for selection of a neural network architecture subsequent to construction of the supernet; ([p. 9] "Algorithm 1 […] from the supernet A" Hu explicitly discloses that TF-NAS finds latency-constrained architectures from the supernet. Algorithm 1 also derives a seed network for the supernet and puts it back into the supernet, and the experiments section states the supernet is trained while architecture-distribution parameters are updated, directly placing the search over the supernet.)
and express and/or structure the neural network architecture search space in a non-transitory storage medium. ([p. 4] "we build a layer-wise search space, which is depicted in Fig. 1 and Tab. 1" [p. 10] "Overall Algorithm. Our Three-Freedom NAS (TF-NAS) contains all above components: the bi-sampling search algorithm, the sink-connecting search space and the elasticity-scaling strategy. It finds latency-constrained architectures from the supernet" [p. 10] "All the experiments are conducted on ImageNet [10] under the mobile set ting. Similar with [3], the latency is measured with a batch size of 32 on a Titan RTX GPU. We set the number of threads for OpenMP to 1 and use Py torch1.1+cuDNN7.6.0 […] This procedure takes about 1.8 days on 1 Titan RTX GPU").
However, Hu does not explicitly teach the plurality of design parameters comprising convolutional neural network channel pruning, quantization, and at least one additional design parameter.
Benmeziane, in the same field of endeavor, teaches the plurality of design parameters comprising convolutional neural network channel pruning, quantization, and at least one additional design parameter.([p. 1] "This kind of NAS, called hardware-aware NAS (HW-NAS), makes searching the most efficient architecture more complicated and opens several questions" [p. 3] "we introduce other considerations that either tackle the hardware efficiency objective, model compressions such as automatic quantization and pruning or add other objectives to the NAS like the robustness against adversarial attacks" [p. 14] "Gumbel Softmax: [156] One way to relax the discrete variables is to use the Gumbel softmax function. It helps insert some random noise following the Gumbel distribution so that the gradient computation is possible" [p. 17] "mixed-precision quantization that applies different bitwidth values for different layers in the same network is more commonly used [...] X. Dong and Y. Yang in [176] propose to prune the over parameterized network without performance damage. They directly search within their NAS process for a network with a flexible channel and layer sizes" See also p. 4 "Other Considerations for HW-NAS" explicitly including Quantization, pruning, and security and reliability with implementation details. Benmeziane explicitly describes automatic mixed-precision quantization as NAS that searches layer bitwidths and explicitly recites gradient based convolutional channel pruning methods by Dong and Yang.).
Hu as well as Benmeziane are directed towards hardware aware neural architecture search. Therefore, Hu as well as Benmeziane are analogous art in the same field of endeavor. It would have been obvious before the effective filing date of the claimed invention to combine the teachings of Hu with the teachings of Benmeziane by including channel pruning and quantization among other design parameters searched using Hu's gradient-based supernet approach. Benmeziane provides as additional motivation for combination ([p. 23] “HW-NAS offers a paradigm that opens up the design space and pushes forward the Pareto frontier between hardware efficiency and model accuracy for efficient and improved hardware/software co-design, hence pushing AI to its next frontiers”). This motivation for combination also applies to the remaining claims which depend on this combination.
Regarding claim 12, the combination of Hu and Benmeziane teaches The computing device of claim 11, wherein the one or more processors are further to: determine the neural network architecture search space based, at least in part, on combinations activation bit width and weight bit width to be implemented in one or more layers of a neural network-based inference engine; and(Benmeziane [p. 17] "Recently, mixed-precision quantization that applies different bitwidth values for different layers in the same network is more commonly used. HAQ [62] implemented a dedicated reinforcement learning agent that learns to assign the right bitwidth to each layer. Their goal is to specialize the architecture to a specific hardware platform by incorporating hardware constraints and accuracy into the reward function. For each layer, the agent takes two decisions: one for the weights and another one for the activation")
quantify costs associated with the combinations of activation bit width and weight bit width.(Benmeziane [p. 17] "Recently, mixed-precision quantization that applies different bitwidth values for different layers in the same network is more commonly used. HAQ [62] implemented a dedicated reinforcement learning agent that learns to assign the right bitwidth to each layer. Their goal is to specialize the architecture to a specific hardware platform by incorporating hardware constraints and accuracy into the reward function. For each layer, the agent takes two decisions: one for the weights and another one for the activation" Benmeziane is explicit that HAQ quantifies reward/cost associated with the combinations of activation and weight bit widths).
Regarding claim 17, the combination of Hu and Benmeziane teaches The computing device of claim 11, wherein at least one of the one or more design constraints are defined by execution latency, operation count, model size, power consumption or memory usage, or a combination thereof.(Hu [p. 1 §1] "specific resource constraints (e.g. FLOPs, latency, energy) is more important in practice" [p. 3] "the number of parameters, FLOPs and latency for neural architecture search" [p. 8] "Due to the coarse-grained search space for latency, some target latency cannot be precisely satisfied, leading to instability during architecture searching. For example, setting the target latency to be 15ms, we search two architectures: one is 14.32ms and the other is 15.76ms. Both of them have around 0.7ms gaps for the target latency.").
Regarding claim 18, Hu teaches A method comprising: executing computer-readable instructions by one or more processors of a computing device ([p. 10] "All the experiments are conducted on ImageNet [10] under the mobile set ting. Similar with [3], the latency is measured with a batch size of 32 on a Titan RTX GPU. We set the number of threads for OpenMP to 1 and use Py torch1.1+cuDNN7.6.0 […] This procedure takes about 1.8 days on 1 Titan RTX GPU")
to execute a neural network architecture search process ([p. 1] "Abstract. With the flourish of differentiable neural architecture search (NAS), automatically searching latency-constrained architectures gives a new perspective to reduce human labor and expertise. However, the searched architectures are usually suboptimal in accuracy and may have large jitters around the target latency. In this paper, we rethink three freedoms of differentiable NAS, i.e. operation-level, depth-level and width level, and propose a novel method, named Three-Freedom NAS (TF NAS), to achieve both good classification accuracy and precise latency constraint")
to select a design option for a neural network-based inference engine from a plurality of candidate design options expressed ([p. 4] "we focus on latency-constrained macro search" [p. 8] "Due to the coarse-grained search space for latency, some target latency cannot be precisely satisfied, leading to instability during architecture searching. For example, setting the target latency to be 15ms, we search two architectures: one is 14.32ms and the other is 15.76ms. Both of them have around 0.7ms gaps for the target latency." [p. 23] "Each candidate operation has a kernel size k=3 or k=5 and a continuous expansion ratio eE[2,4] or eE[4,8]" [p. 10] "After searching, the best architecture is derived from the supernet based on α and β, where the strongest operation in each layer and the strongest depth in each stage are chosen")
and/or structured as a neural network architecture search space in a non-transitory storage medium, ([p. 4] "we build a layer-wise search space, which is depicted in Fig. 1 and Tab. 1" [p. 10] "Overall Algorithm. Our Three-Freedom NAS (TF-NAS) contains all above components: the bi-sampling search algorithm, the sink-connecting search space and the elasticity-scaling strategy. It finds latency-constrained architectures from the supernet")
the plurality of candidate design options defining a neural network architecture search space having been determined based, at least in part, on: determination of a neural network architecture search space subject to one or more design constraints ([p. 2 §1] "Our TF-NAScan search architectures with precise latency on target devices" [p. 14 §5] "we have proposed Three-Freedom NAS (TF-NAS) to seek an architecture with good accuracy as well as precise latency on the target devices.")
by application of the one or more design constraints to the identification of the hardware computing resources based, at least in part, on hardware computing resources for execution of a neural network, ([p. 4] "ω and α are the supernet weights and the architecture distribution parameters, respectively. Given a supernet A, we aim to search a subnet α∗ ∈ A that minimizes the validation loss Lval (ω∗,α) and the latency constraint C (LAT(α)), where the weights ω∗ of supernet are obtained by minimizing the training loss Ltrain (ω,α) and λ is a trade-off hyperparameter" See also Table 1 "Macro architecture of the supernet")
and to construct a supernet based on the neural architecture search space, and to perform a neural architecture search employing a plurality of gradient functions, each of the plurality of gradient functions associated with one of a plurality of design parameters, ([p. 4] "Sampling a subnet from supernet A is a non-differentiable process w.r.t. the architecture distribution parameters α. Therefore, a continuous relaxation is needed to allow back-propagation. Assuming there are N operations to be searched in each layer, we define opl i and αl i as the i-th operation in layer l and its architecture distribution parameter, respectively [...] the gradients of all the αli can be back-propagated through Eq. (3).")
over the supernet for selection of the neural network architecture subsequent to construction of the supernet,(]p. 9] "Algorithm 1 […] from the supernet A" Hu explicitly discloses that TF-NAS finds latency-constrained architectures from the supernet. Algorithm 1 also derives a seed network for the supernet and puts it back into the supernet, and the experiments section states the supernet is trained while architecture-distribution parameters are updated, directly placing the search over the supernet.).
However, Hu does not explicitly teach the plurality of design parameters comprising convolutional neural network channel pruning, quantization, and at least one additional design parameter..
Benmeziane, in the same field of endeavor, teaches the plurality of design parameters comprising convolutional neural network channel pruning, quantization, and at least one additional design parameter.([p. 1] "This kind of NAS, called hardware-aware NAS (HW-NAS), makes searching the most efficient architecture more complicated and opens several questions" [p. 3] "we introduce other considerations that either tackle the hardware efficiency objective, model compressions such as automatic quantization and pruning or add other objectives to the NAS like the robustness against adversarial attacks" [p. 14] "Gumbel Softmax: [156] One way to relax the discrete variables is to use the Gumbel softmax function. It helps insert some random noise following the Gumbel distribution so that the gradient computation is possible" [p. 17] "mixed-precision quantization that applies different bitwidth values for different layers in the same network is more commonly used [...] X. Dong and Y. Yang in [176] propose to prune the over parameterized network without performance damage. They directly search within their NAS process for a network with a flexible channel and layer sizes" See also p. 4 "Other Considerations for HW-NAS" explicitly including Quantization, pruning, and security and reliability with implementation details. Benmeziane explicitly describes automatic mixed-precision quantization as NAS that searches layer bitwidths and explicitly recites gradient based convolutional channel pruning methods by Dong and Yang.).
Hu as well as Benmeziane are directed towards hardware aware neural architecture search. Therefore, Hu as well as Benmeziane are analogous art in the same field of endeavor. It would have been obvious before the effective filing date of the claimed invention to combine the teachings of Hu with the teachings of Benmeziane by including channel pruning and quantization among other design parameters searched using Hu's gradient-based supernet approach. Benmeziane provides as additional motivation for combination ([p. 23] “HW-NAS offers a paradigm that opens up the design space and pushes forward the Pareto frontier between hardware efficiency and model accuracy for efficient and improved hardware/software co-design, hence pushing AI to its next frontiers”). This motivation for combination also applies to the remaining claims which depend on this combination.
Regarding claim 19, the combination of Hu and Benmeziane teaches The method of claim 18, wherein at least one of the one or more design constraints are defined by execution latency, operation count, model size, power consumption or memory usage, or a combination thereof.(Hu [p. 1 §1] "specific resource constraints (e.g. FLOPs, latency, energy) is more important in practice" [p. 3] "the number of parameters, FLOPs and latency for neural architecture search" [p. 8] "Due to the coarse-grained search space for latency, some target latency cannot be precisely satisfied, leading to instability during architecture searching. For example, setting the target latency to be 15ms, we search two architectures: one is 14.32ms and the other is 15.76ms. Both of them have around 0.7ms gaps for the target latency.").
Claims 3, 4, 6, 7, 13, 14, and 16 are rejected under U.S.C. §103 as being unpatentable over the combination of Hu and Benmeziane and Dong (“HAO: Hardware-aware Neural Architecture Optimization for Efficient Inference”, 2021).
Regarding claim 3, the combination of Hu and Benmeziane teaches The method of claim 2.
However, the combination of Hu and Benmeziane doesn't explicitly teach, and further comprising executing computer-readable instructions by one or more processors of the computing device to: determine the neural architecture search space further based, at least in part, on candidate channel sizes for at least one of the one or more layers of the neural network based, at least in part, on the combinations of activation bit width and weight bit width; and
quantify costs associated with the candidate channel sizes based, at least in part, on the quantified costs associated with the combinations of activation bit width and weight bit width.
Dong, in the same field of endeavor, teaches computer-readable instructions by one or more processors of the computing device to: determine the neural architecture search space further based, at least in part, on candidate channel sizes for at least one of the one or more layers of the neural network based, at least in part, on the combinations of activation bit width and weight bit width; and([p. 1] "As an example, the quantization algorithm may select a mixture of every bitwidth from 1 bit to 8 bit, and the NAS algorithm may choose to jointly use convolution with different kernel and group sizes" [p. 4 §III] "Given a layer with input channel size IC, output channel size" [p. 5] "We set no limit on the total number of subgraphs and choose the channel size for different layers")
quantify costs associated with the candidate channel sizes based, at least in part, on the quantified costs associated with the combinations of activation bit width and weight bit width.(See Table II which quantifies costs (framerates) associated with models having a channel size of 192x192, 256x256, etc. on mixed bit-width quantized models.).
The combination of Hu and Benmeziane as well as Dong are directed towards neural architecture search. Therefore, the combination of Hu and Benmeziane as well as Dong are analogous art in the same field of endeavor. It would have been obvious before the effective filing date of the claimed invention to combine the teachings of the combination of Hu and Benmeziane with the teachings of Dong by implementing the bit-width selection and quantization as part of the per layer search space in Hu. the combination of Hu and Benmeziane already optimizes towards a target hardware latency Dong provides as additional motivation for combination ([p. 2 §2] “the authors integrate several model compression techniques in the search framework and use quantization to reduce the latency of weight loading. [25] proposes a uniformed differentiable search algorithm using gumbel-softmax to sample discrete implementation hyperparameters including quantization bitwidth”). This motivation for combination also applies to the remaining claims which depend on this combination.
Regarding claim 4, the combination of Hu, Benmeziane, and Dong teaches The method of claim 3, wherein at least one of the one or more layers of the neural network comprises a convolution layer to be implemented at least in part by application of a kernel, the method further comprising executing computer-readable instructions by one or more processors of the computing device to:(Dong [p. 8] "Hyperparameter Optimization for a fixed architecture: In this approach a neural architecture is given including its operator choices. The objective is limited to optimizing the architecture hyperparameters (e.g., number of channels, stride, kernel size)")
determine the neural architecture search space further based, at least in part, on available kernel sizes for the kernel.(Dong [p. 1] "As an example, the quantization algorithm may select a mixture of every bitwidth from 1 bit to 8 bit, and the NAS algorithm may choose to jointly use convolution with different kernel and group sizes").
Regarding claim 6, the combination of Hu, Benmeziane, and Dong teaches The method of claim 3, and further comprising executing computer-readable instructions by one or more processors of the computing device to: determine the neural architecture search space further based, at least in part, on candidate operator types for the at least one of the one or more layers of the neural network based, at least in part, on the combinations of activation bit width and weight bit width; and(Dong [p. 3 §IIIA] "1) Hardware Subgraph Template: As shown in Fig. 1, in HAO, we adopt a subgraph-based hardware design. A subgraph consists of several convolution kernels that are spatially mapped on hardware, which also corresponds to the major building block of neural architecture [...] We implement a parameterizable accelerator template in high-level synthesis (HLS). The generated dataflow accelerator can contain M convolution kernels chained through FIFOs to exploit pipeline-level parallelism. Each convolution kernel can be chosen from one of the three convolutions from the kernel pool: Conv k×k, Depthwise Conv k×k [8], and Conv 1 × 1. The hardware implementation of each kernel typically comprises a weight buffer, a line buffer, a MAC engine, and a quantization unit to rescale outputs.All the computational units are implemented using integer-only arithmetics" Conv kxk, DW Conv kxk, and Conv 1x1 are interpreted as candidate operator types the neural architecture search space is based on. See also FIG. 1.)
quantify costs associated with the candidate operator types based, at least in part, on the quantified costs associated with the candidate channel sizes.(Dong [p. 8] "In Fig. 7, we show one of the searched results by HAO. A subgraph {1x1 convolution, 3x3 depthwise convolution, 1x1 convolution} is used in this solution. As can be seen, HAO finds that a 6-bit/7-bit mixed-precision quantization setting is better than 8-bit uniform quantization for weights. In general, lower bit-width means more computation units under the same resource constraints, but it can lead to larger quantization perturbation [...] the results of HAO show that, for our implementation on Zynq ZU3EG, solutions with solely 3 × 3 depthwise convolution perform better than those with a mixture of 3 × 3 and 5 × 5 depthwise convolution. This is due to the fact that when using a mixture of 3 × 3 and 5 × 5 depthwise convolution, either 3×3 or 5×5 kernel will be idle when invoking the accelerator").
Regarding claim 7, the combination of Hu, Benmeziane, and Dong teaches The method of claim 6, and further comprising executing computer-readable instructions by the one or more processors of the computing device to: determine a union of candidate design options over the candidate operator types for the at least one of the one or more layers of the neural network; and(Dong [p. 8] "In Fig. 7, we show one of the searched results by HAO. A subgraph {1x1 convolution, 3x3 depthwise convolution, 1x1 convolution} is used in this solution. As can be seen, HAO finds that a 6-bit/7-bit mixed-precision quantization setting is better than 8-bit uniform quantization for weights. In general, lower bit-width means more computation units under the same resource constraints, but it can lead to larger quantization perturbation [...] the results of HAO show that, for our implementation on Zynq ZU3EG, solutions with solely 3 × 3 depthwise convolution perform better than those with a mixture of 3 × 3 and 5 × 5 depthwise convolution. This is due to the fact that when using a mixture of 3 × 3 and 5 × 5 depthwise convolution, either 3×3 or 5×5 kernel will be idle when invoking the accelerator")
determine a union of the candidate design options over the combinations of activation bit width and weight bit width based, at least in part, on the determined union of the candidate design options over the candidate operator types for the at least one of the one or more layers of the neural network. (Dong [p. 8] "In Fig. 7, we show one of the searched results by HAO. A subgraph {1x1 convolution, 3x3 depthwise convolution, 1x1 convolution} is used in this solution. As can be seen, HAO finds that a 6-bit/7-bit mixed-precision quantization setting is better than 8-bit uniform quantization for weights. In general, lower bit-width means more computation units under the same resource constraints, but it can lead to larger quantization perturbation [...] the results of HAO show that, for our implementation on Zynq ZU3EG, solutions with solely 3 × 3 depthwise convolution perform better than those with a mixture of 3 × 3 and 5 × 5 depthwise convolution. This is due to the fact that when using a mixture of 3 × 3 and 5 × 5 depthwise convolution, either 3×3 or 5×5 kernel will be idle when invoking the accelerator").
Regarding claim 13, the combination of Hu and Benmeziane teaches The computing device of claim 12.
However, the combination of Hu and Benmeziane doesn't explicitly teach wherein the one or more processors are further to: determine the neural network architecture search space further based, at least in part, on candidate channel sizes for at least one of the one or more layers of the neural network-based inference engine based, at least in part, on the combinations of activation bit width and weight bit width; and
quantify costs associated with the candidate channel sizes based, at least in part, on the quantified costs associated with the combinations of activation bit width and weight bit width.
Dong, in the same field of endeavor, teaches the one or more processors are further to: determine the neural network architecture search space further based, at least in part, on candidate channel sizes for at least one of the one or more layers of the neural network-based inference engine based, at least in part, on the combinations of activation bit width and weight bit width; and([p. 1] "As an example, the quantization algorithm may select a mixture of every bitwidth from 1 bit to 8 bit, and the NAS algorithm may choose to jointly use convolution with different kernel and group sizes" [p. 4 §III] "Given a layer with input channel size IC, output channel size" [p. 5] "We set no limit on the total number of subgraphs and choose the channel size for different layers")
quantify costs associated with the candidate channel sizes based, at least in part, on the quantified costs associated with the combinations of activation bit width and weight bit width.(See Table II which quantifies costs (framerates) associated with models having a channel size of 192x192, 256x256, etc. on mixed bit-width quantized models.).
The combination of Hu and Benmeziane as well as Dong are directed towards neural architecture search. Therefore, the combination of Hu and Benmeziane as well as Dong are analogous art in the same field of endeavor. It would have been obvious before the effective filing date of the claimed invention to combine the teachings of the combination of Hu and Benmeziane with the teachings of Dong by implementing the bit-width selection and quantization as part of the per layer search space in Hu. the combination of Hu and Benmeziane already optimizes towards a target hardware latency Dong provides as additional motivation for combination ([p. 2 §2] “the authors integrate several model compression techniques in the search framework and use quantization to reduce the latency of weight loading. [25] proposes a uniformed differentiable search algorithm using gumbel-softmax to sample discrete implementation hyperparameters including quantization bitwidth”). This motivation for combination also applies to the remaining claims which depend on this combination.
Regarding claim 14, the combination of Hu, Benmeziane, and Dong teaches The computing device of claim 13, wherein at least one of the one or more layers of the neural network-based inference engine comprises a convolution layer to be implemented at least in part by application of a kernel, and wherein the one or more processors are further to: (Dong [p. 3 §IIIA] "We implement a parameterizable accelerator template in high-level synthesis (HLS). The generated dataflow accelerator can contain M convolution kernels chained through FIFOs to exploit pipeline-level parallelism. Each convolution kernel can be chosen from one of the three convolutions from the kernel pool")
determine the neural network architecture search space further based, at least in part, on available kernel sizes for the kernel.(Dong [p. 1] "As an example, the quantization algorithm may select a mixture of every bitwidth from 1 bit to 8 bit, and the NAS algorithm may choose to jointly use convolution with different kernel and group sizes").
Regarding claim 16, the combination of Hu, Benmeziane, and Dong teaches The computing device of claim 13, wherein the one or more processors are further to: determine the neural network architecture search space further based, at least in part, on candidate operator types for the at least one of the one or more layers of the neural network-based inference engine based, at least in part, on the combinations of activation bit width and weight bit width; and(Dong [p. 3 §IIIA] "1) Hardware Subgraph Template: As shown in Fig. 1, in HAO, we adopt a subgraph-based hardware design. A subgraph consists of several convolution kernels that are spatially mapped on hardware, which also corresponds to the major building block of neural architecture [...] We implement a parameterizable accelerator template in high-level synthesis (HLS). The generated dataflow accelerator can contain M convolution kernels chained through FIFOs to exploit pipeline-level parallelism. Each convolution kernel can be chosen from one of the three convolutions from the kernel pool: Conv k×k, Depthwise Conv k×k [8], and Conv 1 × 1. The hardware implementation of each kernel typically comprises a weight buffer, a line buffer, a MAC engine, and a quantization unit to rescale outputs.All the computational units are implemented using integer-only arithmetics" Conv kxk, DW Conv kxk, and Conv 1x1 are interpreted as candidate operator types the neural architecture search space is based on. See also FIG. 1.)
quantify costs associated with the candidate operator types based, at least in part, on the quantified costs associated with the candidate channel sizes.(Dong [p. 8] "In Fig. 7, we show one of the searched results by HAO. A subgraph {1x1 convolution, 3x3 depthwise convolution, 1x1 convolution} is used in this solution. As can be seen, HAO finds that a 6-bit/7-bit mixed-precision quantization setting is better than 8-bit uniform quantization for weights. In general, lower bit-width means more computation units under the same resource constraints, but it can lead to larger quantization perturbation [...] the results of HAO show that, for our implementation on Zynq ZU3EG, solutions with solely 3 × 3 depthwise convolution perform better than those with a mixture of 3 × 3 and 5 × 5 depthwise convolution. This is due to the fact that when using a mixture of 3 × 3 and 5 × 5 depthwise convolution, either 3×3 or 5×5 kernel will be idle when invoking the accelerator").
Conclusion
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any extension fee pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SIDNEY VINCENT BOSTWICK whose telephone number is (571)272-4720. The examiner can normally be reached M-F 7:30am-5:00pm EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Miranda Huang can be reached on (571)270-7092. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/SIDNEY VINCENT BOSTWICK/Examiner, Art Unit 2124