Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 04-09-2025 is in compliance
with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Objections
Claim 3 is objected to because of the following informalities: Line 5 of claim 3 contains two consecutive commas. Correction is required by deleting one of the duplicate commas.
Appropriate correction is required.
Claim Rejections - 35 USC § 112(b)
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claim 2 is rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claim 2 recites the limitation " wherein the cost associated with selection of the candidate layer comprises one or more constraints" in lines 1-2. There is insufficient antecedent basis for this limitation in the claim, because neither claim 2 nor independent claim 1, from which claim 2 depends, previously introduces or recites selecting a candidate layer. Accordingly, it is unclear what selection the recited cost is associated with.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-18 are rejected under U.S.C 101 for containing an abstract idea without significantly more.
Regarding claim 1:
Step 1 – Is the claim to a process, machine, manufacture or composition of matter?
Yes, the claim is a process.
Step 2A – Prong 1 – Does the claim recite an abstract idea, law of nature, or natural phenomenon?
Yes, the claim recites an abstract idea.
determining a cost metric for each of the candidate model layers, wherein a cost metric is indicative of a cost associated with inclusion of the candidate model layer in an optimized machine-learned model; This limitation is directed to the abstract idea of a mental process (including an observation, evaluation, judgement, opinion) which can be performed in the human mind, or by a human using pen and paper (see MPEP 2106.04(a)(2) Ill. C.)
for each model layer, grouping the one or more respective candidate model layers into one or more candidate layer clusters, wherein each candidate layer cluster is associated with a range of cost metrics; This limitation is directed to the abstract idea of a mental process (including an observation, evaluation, judgement, opinion) which can be performed in the human mind, or by a human using pen and paper (see MPEP 2106.04(a)(2) Ill. C.)
filtering at least one candidate model layer based the cost metric associated with the candidate model layer being greater than a threshold cost; and This limitation is directed to the abstract idea of a mental process (including an observation, evaluation, judgement, opinion) which can be performed in the human mind, or by a human using pen and paper (see MPEP 2106.04(a)(2) Ill. C.)
wherein the cost function maximizes a performance metric of the optimized machine-learned model subject to a sum of the cost metrics associated with each candidate model layer included in the optimized machine-learned model being less than a maximum cost. This limitation is directed to mathematical calculation (see MPEP 2106.04(a)(2) l. C.)
Step 2A – Prong 2 – Does the claim recite additional elements that integrate the judicial exception into a practical application?
No, there are no additional elements that integrate the judicial exception into a practical application. The additional elements:
A computing system for layer-wise neural architecture search with polynomial
complexity to combinatorically construct an optimized machine-learned model, comprising: This limitation is directed to a computer merely used as a tool to perform an existing process (see MPEP 2106.05(f) (2)).
one or more processors; This limitation is directed to a computer merely used as a tool to perform an existing process (see MPEP 2106.05(f) (2)).
one or more non-transitory computer-readable media that store instructions that, when executed by the one or more processors, cause the participant computing device to perform operations, the operations comprising: This limitation is directed to a computer merely used as a tool to perform an existing process (see MPEP 2106.05(f) (2)).
iteratively constructing, for each model layer of a plurality of model layers, one or more candidate model layers; Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea [see MPEP 2106.05(f)] and therefore fails to integrate the exception into a practical application.
constructing an optimized machine-learned model comprising a candidate model layer for each of the plurality of layers based on a cost function, Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea [see MPEP 2106.05(f)] and therefore fails to integrate the exception into a practical application.
Step 2B – Does the claim recite additional elements that amount to significantly more than the judicial exception?
No, there are no additional elements that amount to significantly more than the judicial exception. The additional elements are:
A computing system for layer-wise neural architecture search with polynomial
complexity to combinatorically construct an optimized machine-learned model, comprising: This limitation is directed to a computer merely used as a tool to perform an existing process (see MPEP 2106.05(f) (2)).
one or more processors; This limitation is directed to a computer merely used as a tool to perform an existing process (see MPEP 2106.05(f) (2)).
one or more non-transitory computer-readable media that store instructions that, when executed by the one or more processors, cause the participant computing device to perform operations, the operations comprising: This limitation is directed to a computer merely used as a tool to perform an existing process (see MPEP 2106.05(f) (2)).
iteratively constructing, for each model layer of a plurality of model layers, one or more candidate model layers; Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea [see MPEP 2106.05(f)] and therefore fails to integrate the exception into a practical application.
constructing an optimized machine-learned model comprising a candidate model layer for each of the plurality of layers based on a cost function, Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea [see MPEP 2106.05(f)] and therefore fails to integrate the exception into a practical application.
Regarding claim 2,
Claim 2 is rejected under 35 U.S.C 101 because the claimed invention is directed to an abstract idea without significantly more. The claim is dependent on claim 1 which includes an abstract idea (see rejection for claim 1). The additional limitations:
wherein the cost associated with selection of the candidate layer comprises one or more constraints, comprising: a size of the candidate layer; a degree of energy consumption associated with the candidate layer; or an inference latency associated with the candidate layer. This limitation is directed to the abstract idea of a mental process (including an observation, evaluation, judgement, opinion) which can be performed in the human mind, or by a human using pen and paper (see MPEP 2106.04(a)(2) Ill. C.)
Regarding claim 3,
Claim 3 is rejected under 35 U.S.C 101 because the claimed invention is directed to an abstract idea without significantly more. The claim is dependent on claim 2 which includes an abstract idea (see rejection for claim 2). The additional limitations:
determining, by the computing system, that the performance metric for the
optimized machine-learned model is less than a threshold degree of performance; and This limitation is directed to the abstract idea of a mental process (including an observation, evaluation, judgement, opinion) which can be performed in the human mind, or by a human using pen and paper (see MPEP 2106.04(a)(2) Ill. C.)
constructing, by the computing system, a second optimized machine-learned model comprising a candidate model layer for each of the plurality of layers based on the cost function,, Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea [see MPEP 2106.05(f)] and therefore fails to integrate the exception into a practical application.
wherein the second cost function maximizes the performance metric of the optimized machine-learned model subject to a sum of the cost metrics associated with each candidate model layer included in the optimized machine-learned model being less than a second maximum cost greater than the maximum cost. This limitation is directed to mathematical calculation (see MPEP 2106.04(a)(2) l. C.)
Regarding claim 4,
Claim 4 is rejected under 35 U.S.C 101 because the claimed invention is directed to an abstract idea without significantly more. The claim is dependent on claim 1 which includes an abstract idea (see rejection for claim 1). The additional limitations:
for each model layer N of a plurality of model layers M: selecting, [by a computing system comprising one or more computing devices], one or more layer search options from a plurality of layer search options; This limitation is directed to the abstract idea of a mental process (including an observation, evaluation, judgement, opinion) which can be performed in the human mind, or by a human using pen and paper (see MPEP 2106.04(a)(2) Ill. C.)
by a computing system comprising one or more computing devices This limitation is directed to a computer merely used as a tool to perform an existing process (see MPEP 2106.05(f) (2)).
based on the model layer N, using, [by the computing system], the one or more search options to construct one or more candidate model layers for a model layer N+1 of the plurality of model layers, Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea [see MPEP 2106.05(f)] and therefore fails to integrate the exception into a practical application.
by the computing system This limitation is directed to a computer merely used as a tool to perform an existing process (see MPEP 2106.05(f) (2)).
wherein the one or more candidate model layers are respectively associated with one or more cost metrics, This limitation is directed to the abstract idea of a mental process (including an observation, evaluation, judgement, opinion) which can be performed in the human mind, or by a human using pen and paper (see MPEP 2106.04(a)(2) Ill. C.)
wherein a cost metric is indicative of a cost associated with inclusion of the candidate model layer in an optimized machine-learned model; and This limitation is directed to mathematical calculation (see MPEP 2106.04(a)(2) l. C.)
constructing, by the computing system, an optimized machine-learned model comprising M model layers based on a cost function, Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea [see MPEP 2106.05(f)] and therefore fails to integrate the exception into a practical application.
wherein the cost function maximizes a performance metric of the optimized machine-learned model subject to a sum of the cost metrics associated with each candidate model layer included in the optimized machine-learned model being less than a maximum cost. This limitation is directed to mathematical calculation (see MPEP 2106.04(a)(2) l. C.)
Regarding claim 5, this claim is rejected under the same rationale with claim 2 (as shown in the rejections above) because they are analogous claims.
Regarding claim 6,
Claim 6 is rejected under 35 U.S.C 101 because the claimed invention is directed to an abstract idea without significantly more. The claim is dependent on claim 4 which includes an abstract idea (see rejection for claim 4). The additional limitations:
wherein the one or more candidate layers comprises a plurality of candidate layers respectively associated with a plurality of cost metrics, wherein each of the plurality of cost metrics is different, and This limitation is directed to the abstract idea of a mental process (including an observation, evaluation, judgement, opinion) which can be performed in the human mind, or by a human using pen and paper (see MPEP 2106.04(a)(2) Ill. C.)
wherein each candidate layer represents an optimal layer for a respectively associated cost metric; and This limitation is directed to the abstract idea of a mental process (including an observation, evaluation, judgement, opinion) which can be performed in the human mind, or by a human using pen and paper (see MPEP 2106.04(a)(2) Ill. C.)
wherein using the one or more search options to identify one or more candidate layers further comprises: This limitation is directed to the abstract idea of a mental process (including an observation, evaluation, judgement, opinion) which can be performed in the human mind, or by a human using pen and paper (see MPEP 2106.04(a)(2) Ill. C.)
grouping, [by the computing system], the plurality of candidate layers into a plurality of candidate layer clusters, This limitation is directed to the abstract idea of a mental process (including an observation, evaluation, judgement, opinion) which can be performed in the human mind, or by a human using pen and paper (see MPEP 2106.04(a)(2) Ill. C.)
by the computing system This limitation is directed to a computer merely used as a tool to perform an existing process (see MPEP 2106.05(f) (2)).
wherein each candidate layer cluster is associated with a range of cost metrics. This limitation is directed to the abstract idea of a mental process (including an observation, evaluation, judgement, opinion) which can be performed in the human mind, or by a human using pen and paper (see MPEP 2106.04(a)(2) Ill. C.)
Regarding claim 7,
Claim 7 is rejected under 35 U.S.C 101 because the claimed invention is directed to an abstract idea without significantly more. The claim is dependent on claim 4 which includes an abstract idea (see rejection for claim 4). The additional limitations:
wherein constructing the optimized machine-learned model comprises, for each layer of the optimized machine-learned model: Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea [see MPEP 2106.05(f)] and therefore fails to integrate the exception into a practical application.
determining, [by the computing system], a candidate layer cluster of the plurality of candidate layer clusters for the layer based on the cost function and the range of cost metrics associated with the candidate layer cluster; and This limitation is directed to the abstract idea of a mental process (including an observation, evaluation, judgement, opinion) which can be performed in the human mind, or by a human using pen and paper (see MPEP 2106.04(a)(2) Ill. C.)
by the computing system This limitation is directed to a computer merely used as a tool to perform an existing process (see MPEP 2106.05(f) (2)).
selecting, [by the computing system], a candidate layer from the candidate layer cluster based on the cost function and the cost metrics associated with one or more layers selected prior to the candidate layer. This limitation is directed to the abstract idea of a mental process (including an observation, evaluation, judgement, opinion) which can be performed in the human mind, or by a human using pen and paper (see MPEP 2106.04(a)(2) Ill. C.)
Regarding claim 8,
Claim 8 is rejected under 35 U.S.C 101 because the claimed invention is directed to an abstract idea without significantly more. The claim is dependent on claim 4 which includes an abstract idea (see rejection for claim 4). The additional limitations:
wherein grouping the plurality of candidate layers into the plurality of candidate layer clusters further comprises This limitation is directed to the abstract idea of a mental process (including an observation, evaluation, judgement, opinion) which can be performed in the human mind, or by a human using pen and paper (see MPEP 2106.04(a)(2) Ill. C.)
storing, [by the computing system], layer selection information indicative of the plurality of candidate layer clusters and each of the plurality of candidate layers; and This limitation is directed to the abstract idea of a mental process (including an observation, evaluation, judgement, opinion) which can be performed in the human mind, or by a human using pen and paper (see MPEP 2106.04(a)(2) Ill. C.)
by the computing system This limitation is directed to a computer merely used as a tool to perform an existing process (see MPEP 2106.05(f) (2)).
wherein determining the candidate layer cluster of the plurality of candidate layer clusters comprises This limitation is directed to the abstract idea of a mental process (including an observation, evaluation, judgement, opinion) which can be performed in the human mind, or by a human using pen and paper (see MPEP 2106.04(a)(2) Ill. C.)
determining, [by the computing system], the candidate layer cluster of the plurality of candidate layer clusters for the layer based on the cost function and layer selection information. This limitation is directed to the abstract idea of a mental process (including an observation, evaluation, judgement, opinion) which can be performed in the human mind, or by a human using pen and paper (see MPEP 2106.04(a)(2) Ill. C.)
by the computing system This limitation is directed to a computer merely used as a tool to perform an existing process (see MPEP 2106.05(f) (2)).
Regarding claim 9,
Claim 9 is rejected under 35 U.S.C 101 because the claimed invention is directed to an abstract idea without significantly more. The claim is dependent on claim 4 which includes an abstract idea (see rejection for claim 4). The additional limitations:
wherein constructing the optimized machine-learned model based on the cost function comprises, for a model layer N of the plurality of model layers M: Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea [see MPEP 2106.05(f)] and therefore fails to integrate the exception into a practical application.
determining, [by the computing system] for a candidate model layer for the model layer N, that the cost metrics associated with each candidate model layer constructed for the model layer N+1 based on the candidate model layer are greater than a maximum cost; and This limitation is directed to the abstract idea of a mental process (including an observation, evaluation, judgement, opinion) which can be performed in the human mind, or by a human using pen and paper (see MPEP 2106.04(a)(2) Ill. C.)
filtering, [by the computing system], the candidate model layer from inclusion in the optimized machine-learned model. This limitation is directed to the abstract idea of a mental process (including an observation, evaluation, judgement, opinion) which can be performed in the human mind, or by a human using pen and paper (see MPEP 2106.04(a)(2) Ill. C.)
by the computing system This limitation is directed to a computer merely used as a tool to perform an existing process (see MPEP 2106.05(f) (2)).
Regarding claim 10,
Claim 10 is rejected under 35 U.S.C 101 because the claimed invention is directed to an abstract idea without significantly more. The claim is dependent on claim 4 which includes an abstract idea (see rejection for claim 4). The additional limitations:
determining, [by the computing system], that the performance metric for the optimized machine-learned model is less than a threshold degree of performance; and This limitation is directed to the abstract idea of a mental process (including an observation, evaluation, judgement, opinion) which can be performed in the human mind, or by a human using pen and paper (see MPEP 2106.04(a)(2) Ill. C.)
by the computing system This limitation is directed to a computer merely used as a tool to perform an existing process (see MPEP 2106.05(f) (2)).
constructing, [by the computing system], a second optimized machine-learned model comprising M model layers based on a second cost function, Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea [see MPEP 2106.05(f)] and therefore fails to integrate the exception into a practical application.
wherein the second cost function maximizes an accuracy of the optimized machine-learned model subject to a sum of the cost metrics associated with each candidate model layer included in the optimized machine-learned model being less than a second maximum cost greater than the maximum cost. This limitation is directed to mathematical calculation (see MPEP 2106.04(a)(2) l. C.)
Regarding claim 11,
Claim 11 is rejected under 35 U.S.C 101 because the claimed invention is directed to an abstract idea without significantly more. The claim is dependent on claim 4 which includes an abstract idea (see rejection for claim 4). The additional limitations:
wherein, prior to selecting the one or more search options from a plurality of layer search options, the method comprises receiving, [by the computing system], an optimization request indicative of a quantity of layers M and the maximum cost. This limitation is directed to receiving or transmitting data over a network. The courts have recognized receiving or transmitting data over a network as well understood, routine, and conventional functions when they are claimed in a merely generic manner (e.g., at a high level of generality) or as insignificant extra-solution activity (see MPEP 2106.05(d) II.).
by the computing system This limitation is directed to a computer merely used as a tool to perform an existing process (see MPEP 2106.05(f) (2)).
Regarding claim 12, this claim is rejected under the same rationale with claim 4 (as shown in the rejections above) because they are analogous claims
Regarding claim 13, this claim is rejected under the same rationale with claim 5 (as shown in the rejections above) because they are analogous claims
Regarding claim 14, this claim is rejected under the same rationale with claim 6 (as shown in the rejections above) because they are analogous claims
Regarding claim 15, this claim is rejected under the same rationale with claim 7 (as shown in the rejections above) because they are analogous claims
Regarding claim 16, this claim is rejected under the same rationale with claim 8 (as shown in the rejections above) because they are analogous claims
Regarding claim 17, this claim is rejected under the same rationale with claim 9 (as shown in the rejections above) because they are analogous claims
Regarding claim 18, this claim is rejected under the same rationale with claim 10 (as shown in the rejections above) because they are analogous claims
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claim(s) 1-3, 6 and 14 are rejected under 35 U.S.C. 103 as being unpatentable over Molchanov et al. (“LANA: Latency Aware Network Acceleration”) in view of Letunovskiy et al. (US 2024/0169213 A1) and further in view of Gu et al. (“iNAS: Integral NAS for Device-Aware Salient Object Detection”).
Regarding claim 1, Molchanov explicitly discloses:
iteratively constructing, for each model layer of a plurality of model layers, one or more candidate model layers; (Molchanov, Pg. 3, Section 2.1: “We represent the teacher network as the composition of N teacher operations by
PNG
media_image1.png
26
282
media_image1.png
Greyscale
, where x is the input tensor, ti is the ith operation (i.e., layer) in the network. We then define the set of candidate student operations
PNG
media_image2.png
33
141
media_image2.png
Greyscale
, which will be used to approximate the teacher operations. Here, M denotes the number of candidate operations per layer. The student operations can draw from a wide variety of operations – the only requirement is that all candidate operations for a given layer must have the same input and output tensor dimensions as the teacher operation ti.”)
determining a cost metric for each of the candidate model layers, wherein a cost metric is indicative of a cost associated with inclusion of the candidate model layer in an optimized machine-learned model; (Molchanov, Pg. 3, Section 2.1, Col. 2, ¶[2]: “Finding the optimal architecture in hardware-aware NAS reduces to:
PNG
media_image3.png
192
543
media_image3.png
Greyscale
where bi Є RM is a vector of corresponding cost of each student operation (latency, number of parameters, FLOPs, etc.) in layer i. The total budget constraint is defined via scalar B. The objective is to minimize the loss function L that estimates the error with respect to the correct output y while meeting a budget constraint”)
constructing an optimized machine-learned model comprising a candidate model layer for each of the plurality of layers based on a cost function, (Molchanov, Pg. 3, Section 2.1, Col. 2, ¶[2]: “The problem of optimal selection of operations is often tackled in NAS. This problem is usually formulated as a bilevel optimization that selects operations and optimizes their weights jointly [40, 103].Finding the optimal architecture in hardware-aware NAS reduces to:
PNG
media_image3.png
192
543
media_image3.png
Greyscale
where bi Є RM is a vector of corresponding cost of each student operation (latency, number of parameters, FLOPs, etc.) in layer i. The total budget constraint is defined via scalar B. The objective is to minimize the loss function L that estimates the error with respect to the correct output y while meeting a budget constraint”)
wherein the cost function maximizes a performance metric of the optimized machine-learned model subject to a sum of the cost metrics associated with each candidate model layer included in the optimized machine-learned model being less than a maximum cost. (Molchanov, Pg. 3, Section 2.1, Col. 2, ¶[2]: “The problem of optimal selection of operations is often tackled in NAS. This problem is usually formulated as a bilevel optimization that selects operations and optimizes their weights jointly [40, 103].Finding the optimal architecture in hardware-aware NAS reduces to:
PNG
media_image3.png
192
543
media_image3.png
Greyscale
where bi Є RM is a vector of corresponding cost of each student operation (latency, number of parameters, FLOPs, etc.) in layer i. The total budget constraint is defined via scalar B. The objective is to minimize the loss function L that estimates the error with respect to the correct output y while meeting a budget constraint”) [Examiner’s note: Molchanov teaches minimizing the model’s loss or accuracy reduction, which corresponds to maximizing the optimized model’s performance, subject to the summed cost being no greater than the maximum budget]
Molchanov fails to teach:
A computing system for layer-wise neural architecture search with polynomial
complexity to combinatorically construct an optimized machine-learned model, comprising:
one or more processors;
one or more non-transitory computer-readable media that store instructions that, when executed by the one or more processors, cause the participant computing device to perform operations, the operations comprising:
for each model layer, grouping the one or more respective candidate model layers into one or more candidate layer clusters, wherein each candidate layer cluster is associated with a range of cost metrics;
filtering at least one candidate model layer based the cost metric associated with the candidate model layer being greater than a threshold cost; and
However, Letunovskiy explicitly discloses:
A computing system for layer-wise neural architecture search with polynomial
complexity to combinatorically construct an optimized machine-learned model, comprising: (Letunovskiy, ¶[0005]: “This application provides methods and apparatuses, to improve search for neural network architectures. In some embodiments, the search takes into account the hardware for implementing the neural network processing”, ¶[0021]: “”)
one or more processors; (Letunovskiy, ¶[0050]: “According to a fifth aspect, a computer-readable storage medium having stored thereon instructions that, when executed, cause one or more processors to execute any of the above mentioned methods is proposed”)
one or more non-transitory computer-readable media that store instructions that, when executed by the one or more processors, cause the participant computing device to perform operations, the operations comprising: (Letunovskiy, ¶[0050]: “According to a fifth aspect, a computer-readable storage medium having stored thereon instructions that, when executed, cause one or more processors to execute any of the above mentioned methods is proposed. The instructions cause the one or more processors to perform the method according to any of the first to fourth aspect or any possible embodiment or implementation of the first or second aspect.”)
filtering at least one candidate model layer based the cost metric associated with the candidate model layer being greater than a threshold cost; and (Letunovskiy, ¶[0024]: “In an implementation, the searching for the one or more NN architectures comprises performing K times, K being a positive integer, the following steps: pseudo-randomly selecting a first set of candidate architectures from the search space: obtaining a second set of candidate architectures by removing from the first set of candidates those candidate architectures which do not satisfy a predefined condition including latency and/or accuracy”)
It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Letunovskiy and Molchanov. Letunovskiy teaches a method to search for the best neural network architecture. Molchanov teaches a network acceleration framework that automatically replaces inefficient operations in a given network with more efficient counterparts. One of ordinary skill would have motivation to combine Letunovskiy and Molchanov because MPEP 2143 sets forth the Supreme Court rationales for obviousness including: (D) Applying a known technique to a known device (method, or product) ready for improvement to yield predictable results; (E): “Obvious to try” choosing from a finite number of identified, predictable solutions, with a reasonable expectation of success; (F) Known work in one field of endeavor may prompt variations of it for use in either the same field or a different one based on design incentives or other market forces if the variations are predictable to one of the ordinary skill in the art.
However, Gu explicitly discloses:
for each model layer, grouping the one or more respective candidate model layers into one or more candidate layer clusters, wherein each candidate layer cluster is associated with a range of cost metrics; (Gu, Pg. 4918, FIG 5:
PNG
media_image4.png
274
494
media_image4.png
Greyscale
, Pg. 4918, Section 3.2, Col. 2, ¶[1]: “Fig. 5, the whole search space is composed of layer-wise block choices. The block choices within each layer vary in latency. Suppose we uniformly sample the block layer-by-layer, the accumulated latency of overall sampled models will obey a multinomial distribution, i.e., the extremely low-latency or extremely high-latency areas are under-sampled but the middle latency area is over-sampled. To explore the entire latency area of our integral search space, we propose latency-group sampling (LGS). Given a latency lookup table (LUT), we divide the layer-wise search space into several latency groups.”)
The combination of Molchanov and Gu are analogous art because they are in the same field of training time series data. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention, having the teachings of Molchanov and Gu before them, to modify the teachings of Molchanov to include the teachings of Gu grouping candidate model layers into clusters associated with respective cost ranges to reduce the number of candidates evaluated for a given resource budget, thereby decreasing search time and computational expense while preserving candidates representing different cost-performance tradeoffs.
Regarding claim 2, the combination of Molchanov, Letunovskiy and Gu explicitly discloses all the limitations of claim 1 (as shown in the rejections above).
Molchanov in view of Letunovskiy and Gu further discloses:
wherein the cost associated with selection of the candidate layer comprises one or more constraints, comprising: a size of the candidate layer; a degree of energy consumption associated with the candidate layer; or an inference latency associated with the candidate layer. (Molchanov, Pg. 3, Section 2.1, Col. 2, ¶[2]: “Finding the optimal architecture in hardware-aware NAS reduces to:
PNG
media_image3.png
192
543
media_image3.png
Greyscale
where bi Є RM is a vector of corresponding cost of each student operation (latency, number of parameters, FLOPs, etc.) in layer i. The total budget constraint is defined via scalar B. The objective is to minimize the loss function L that estimates the error with respect to the correct output y while meeting a budget constraint”)
Regarding claim 3, the combination of Molchanov, Letunovskiy and Gu explicitly discloses all the limitations of claim 1 (as shown in the rejections above).
Molchanov in view of Letunovskiy and Gu further discloses:
wherein the operations further comprise: determining, by the computing system, that the performance metric for the optimized machine-learned model is less than a threshold degree of performance; and (Molchanov, Pg. 4, Col. 2, ¶[2]: “Approximating the architecture loss with a linear function allows us to formulate the search problem as solving an integer linear program (ILP). This has several main advantages: (i) Although solving integer linear programs is generally NP-hard, there exist many off-the-shelf libraries that can obtain a high-quality solutions in a few seconds. (ii) Since integer linear optimization libraries easily scale up to millions of variables, our search also scales up easily to very large number of candidate operations per layer. (iii) We can easily formulate the search problem such that instead of one architecture, we obtain a set of diverse candidate architectures. Formally, we denote the kth solution with
PNG
media_image5.png
37
85
media_image5.png
Greyscale
which is obtained by solving:
PNG
media_image6.png
159
437
media_image6.png
Greyscale
, where we minimize the change in the loss while satisfying the budget and overlap constraint. The scalar O sets the maximum overlap with any previous solution which is set to 0.7N in our case. We obtain K diverse solutions by solving the minimization above K times.”)
constructing, by the computing system, a second optimized machine-learned model comprising a candidate model layer for each of the plurality of layers based on the cost function,, (Molchanov, Pg. 3, Section 2.1, Col. 2, ¶[2]: “The problem of optimal selection of operations is often tackled in NAS. This problem is usually formulated as a bilevel optimization that selects operations and optimizes their weights jointly [40, 103].Finding the optimal architecture in hardware-aware NAS reduces to:
PNG
media_image3.png
192
543
media_image3.png
Greyscale
where bi Є RM is a vector of corresponding cost of each student operation (latency, number of parameters, FLOPs, etc.) in layer i. The total budget constraint is defined via scalar B. The objective is to minimize the loss function L that estimates the error with respect to the correct output y while meeting a budget constraint”, Pg. 5, Col. 1, ¶[1]: “We obtain K diverse solutions by solving the minimization above K times”)
wherein the second cost function maximizes the performance metric of the optimized machine-learned model subject to a sum of the cost metrics associated with each candidate model layer included in the optimized machine-learned model being less than a second maximum cost greater than the maximum cost. (Molchanov, Pg. 3, Section 2.1, Col. 2, ¶[2]: “The problem of optimal selection of operations is often tackled in NAS. This problem is usually formulated as a bilevel optimization that selects operations and optimizes their weights jointly [40, 103].Finding the optimal architecture in hardware-aware NAS reduces to:
PNG
media_image3.png
192
543
media_image3.png
Greyscale
where bi Є RM is a vector of corresponding cost of each student operation (latency, number of parameters, FLOPs, etc.) in layer i. The total budget constraint is defined via scalar B. The objective is to minimize the loss function L that estimates the error with respect to the correct output y while meeting a budget constraint”, Pg. 5, Col. 1, ¶[1]: “We obtain K diverse solutions by solving the minimization above K times”) [Examiner’s note: Molchanov teaches minimizing the model’s loss or accuracy reduction, which corresponds to maximizing the optimized model’s performance, subject to the summed cost being no greater than the maximum budget]
Regarding claim 6, the combination of Molchanov, Letunovskiy and Gu explicitly discloses all the limitations of claim 4 (as shown in the rejections above).
Molchanov in view of Letunovskiy further discloses:
wherein the one or more candidate layers comprises a plurality of candidate layers respectively associated with a plurality of cost metrics, (Molchanov, Pg. 3, Section 2.1, Col. 2, ¶[2]: “The problem of optimal selection of operations is often tackled in NAS. This problem is usually formulated as a bilevel optimization that selects operations and optimizes their weights jointly [40, 103].Finding the optimal architecture in hardware-aware NAS reduces to:
PNG
media_image3.png
192
543
media_image3.png
Greyscale
where bi Є RM is a vector of corresponding cost of each student operation (latency, number of parameters, FLOPs, etc.) in layer i. The total budget constraint is defined via scalar B. The objective is to minimize the loss function L that estimates the error with respect to the correct output y while meeting a budget constraint”, Pg. 5, Col. 1, ¶[1]: “We obtain K diverse solutions by solving the minimization above K times”, Pg. 5, Col. 1, ¶[4]: “We construct a large of pool of diverse candidate operation including M = 197 operations for each layer of teacher”)
wherein each of the plurality of cost metrics is different, and (Molchanov, Pg. 6, Col. 1: “Linear relaxation in architecture search assumes that a candidate architecture can be scored by a fitness metric measured independently for all operations”)
wherein each candidate layer represents an optimal layer for a respectively associated cost metric; and (Molchanov, Pg. 5, Col. 1, ¶[3]: “Candidate architecture evaluation. Solving Eq. 4 provides us with K architectures. The linear proxy used for candidates loss is calculated in an isolated setting for each operation. To reduce the approximation error, we evaluate all K architectures with pretrained weights from phase one on a small part of the training set (6k images on ImageNet) and select the architecture with the lowest loss”)
wherein each candidate layer cluster is associated with a range of cost metrics. Molchanov, Pg. 3, Section 2.1, Col. 2, ¶[2]: “The problem of optimal selection of operations is often tackled in NAS. This problem is usually formulated as a bilevel optimization that selects operations and optimizes their weights jointly [40, 103].Finding the optimal architecture in hardware-aware NAS reduces to:
PNG
media_image3.png
192
543
media_image3.png
Greyscale
where bi Є RM is a vector of corresponding cost of each student operation (latency, number of parameters, FLOPs, etc.) in layer i.)
Molchanov in view of Letunovskiy fails to disclose:
wherein using the one or more search options to identify one or more candidate layers further comprises: grouping, by the computing system, the plurality of candidate layers into a plurality of candidate layer clusters,
However, Gu explicitly discloses:
wherein using the one or more search options to identify one or more candidate layers further comprises: grouping, by the computing system, the plurality of candidate layers into a plurality of candidate layer clusters, (Gu, Pg. 4915, Col. 1, ¶[3]: “we propose a latency-group sampling (LGS) that introduces the device latency to guide sampling. Dividing the layer-wise search space into several latency groups, and aggregating samples in specific latency groups, LGS preserves the offspring”, Pg. 4919, Col. 1, ¶[1]: “We divide latency ranges of block choices in each layer into G latency groups. We sample N candidates for an initial population P, where each latency group has n/G samples.”, Pg. 4919, Col. 2, ¶[1]: “Then in the selection step, LGS preserves a certain number of elite offsprings in different groups, which enables the evolution search to find better models in different latency areas.”)
The combination of Molchanov, Letunovskiy and Gu are analogous art because they are in the same field of training time series data. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention, having the teachings of Molchanov, Letunovskiy and Gu before them, to modify the teachings of Molchanov and Letunovskiy to include the teachings of Gu grouping candidate model layers into clusters associated with respective cost ranges to reduce the number of candidates evaluated for a given resource budget, thereby decreasing search time and computational expense while preserving candidates representing different cost-performance tradeoffs.
Regarding claim 14, this claim is rejected under the same rationale with claim 6 as they are analogous claims.
Claim(s) 4-5, 9-13 and 17-18 are rejected under 35 U.S.C. 103 as being unpatentable over Molchanov et al. (“LANA: Latency Aware Network Acceleration”) in view of Letunovskiy et al. (US 2024/0169213 A1)
Regarding claim 4, Molchanov explicitly discloses:
A computer-implemented method to implement layerwise optimization of
machine-learned models, the method comprising: for each model layer N of a plurality of model layers M: (Molchanov, Pg. 3, Col. 1, Section 2.1: “We represent the teacher network as the composition of N teacher operations by
PNG
media_image7.png
27
283
media_image7.png
Greyscale
… Here, M denotes the number of candidate operations per layer. The student operations can draw from a wide variety of operations – the only requirement is that all candidate operations for a given layer must have the same input and output tensor dimensions as the teacher operation ti”)
selecting, by a computing system comprising one or more computing devices, one or more layer search options from a plurality of layer search options; (Molchanov, Pg. 3, Section 2: “Operation selection phase (Sec. 2.2), in which we search for an architecture composed of a combination of the original teacher layers and pretrained efficient operations via linear optimization.”, Pg. 4, Col. 2, ¶[2]: “We can easily formulate the search problem such that instead of one architecture, we obtain a set of diverse candidate architectures.”)
wherein the one or more candidate model layers are respectively associated with one or more cost metrics, (Molchanov, Pg. 3, Section 2.1, Col. 2, ¶[2]: “Finding the optimal architecture in hardware-aware NAS reduces to:
PNG
media_image3.png
192
543
media_image3.png
Greyscale
where bi Є RM is a vector of corresponding cost of each student operation (latency, number of parameters, FLOPs, etc.) in layer i.)
wherein a cost metric is indicative of a cost associated with inclusion of the candidate model layer in an optimized machine-learned model; and (Molchanov, Pg. 3, Section 2.1, Col. 2, ¶[2]: “Finding the optimal architecture in hardware-aware NAS reduces to:
PNG
media_image3.png
192
543
media_image3.png
Greyscale
where bi Є RM is a vector of corresponding cost of each student operation (latency, number of parameters, FLOPs, etc.) in layer i. The total budget constraint is defined via scalar B. The objective is to minimize the loss function L that estimates the error with respect to the correct output y while meeting a budget constraint. In general, the optimization problem in Eq. 1 is an NP-hard combinatorial problem with an exponentially large state space (i.e., MN).)
constructing, by the computing system, an optimized machine-learned model comprising M model layers based on a cost function, (Molchanov, Pg. 3, Section 2.1, Col. 2, ¶[2]: “The problem of optimal selection of operations is often tackled in NAS. This problem is usually formulated as a bilevel optimization that selects operations and optimizes their weights jointly [40, 103].Finding the optimal architecture in hardware-aware NAS reduces to:
PNG
media_image3.png
192
543
media_image3.png
Greyscale
where bi Є RM is a vector of corresponding cost of each student operation (latency, number of parameters, FLOPs, etc.) in layer i. The total budget constraint is defined via scalar B. The objective is to minimize the loss function L that estimates the error with respect to the correct output y while meeting a budget constraint”, Pg. 5, Col. 1, ¶[1]: “We obtain K diverse solutions by solving the minimization above K times”, Pg. 5, Col. 1, ¶[4]: “We construct a large of pool of diverse candidate operation including M = 197 operations for each layer of teacher”)
wherein the cost function maximizes a performance metric of the optimized machine-learned model subject to a sum of the cost metrics associated with each candidate model layer included in the optimized machine-learned model being less than a maximum cost. (Molchanov, Pg. 3, Section 2.1, Col. 2, ¶[2]: “The problem of optimal selection of operations is often tackled in NAS. This problem is usually formulated as a bilevel optimization that selects operations and optimizes their weights jointly [40, 103].Finding the optimal architecture in hardware-aware NAS reduces to:
PNG
media_image3.png
192
543
media_image3.png
Greyscale
where bi Є RM is a vector of corresponding cost of each student operation (latency, number of parameters, FLOPs, etc.) in layer i. The total budget constraint is defined via scalar B. The objective is to minimize the loss function L that estimates the error with respect to the correct output y while meeting a budget constraint”, Pg. 5, Col. 1, ¶[1]: “We obtain K diverse solutions by solving the minimization above K times”) [Examiner’s note: Molchanov teaches minimizing the model’s loss or accuracy reduction, which corresponds to maximizing the optimized model’s performance, subject to the summed cost being no greater than the maximum budget]
Molchanov fails to disclose:
based on the model layer N, using, by the computing system, the one or more search options to construct one or more candidate model layers for a model layer N+1 of the plurality of model layers,
However, Letunovskiy explicitly discloses:
based on the model layer N, using, by the computing system, the one or more search options to construct one or more candidate model layers for a model layer N+1 of the plurality of model layers, (Letunovskiy, ¶[0147]: “Every block is divided into edges (Es,i). Edges correspond to neural network layers. Each edge Es,i is one from the list: conv_1x1, conv_3x3, conv_5x5, conv_7x7,conv_1x3_3x1,conv_1x5_5x1,conv_1x7_7x1, or identity.”, ¶[0193]: “The scaling may be performed iteratively, a multiple times e.g. for different numbers of stages and/or different target devices. It is noted that the term "iteratively" herein means that the output of previous iteration is used as input for next iteration. For example, a scaled architecture obtained in step n is an input to a further scaling in step n + 1.”, ¶[0199]: “[0199] In step 860, all architectures A1, ... , AK are constructed with blocks B1, ... , B5 and the numbers of blocks N1+i1, ... , Ns+is. Then, in step 870, the K architectures A1,..., AK are trained with a pre-defined training procedure. In step 880, best architecture A* is found so that:… and the best architecture is added to architectures list M. The best quality architecture for the target latency is thus found based on the accuracy. However, accuracy is only one possible and exemplary criterion. The same steps 820 to 880 are performed for further architectures. The algorithm terminates in step 890, and may return the list of the resulting architectures M.”)
It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Letunovskiy and Molchanov. Letunovskiy teaches a method to search for the best neural network architecture. Molchanov teaches a network acceleration framework that automatically replaces inefficient operations in a given network with more efficient counterparts. One of ordinary skill would have motivation to combine Letunovskiy and Molchanov because MPEP 2143 sets forth the Supreme Court rationales for obviousness including: (D) Applying a known technique to a known device (method, or product) ready for improvement to yield predictable results; (E): “Obvious to try” choosing from a finite number of identified, predictable solutions, with a reasonable expectation of success; (F) Known work in one field of endeavor may prompt variations of it for use in either the same field or a different one based on design incentives or other market forces if the variations are predictable to one of the ordinary skill in the art.
Regarding claim 5, this claim is rejected under the same rationale with claim 2 as they are analogous claims.
Regarding claim 9, the combination of Molchanov, Letunovskiy and Gu explicitly discloses all the limitations of claim 4 (as shown in the rejections above).
Molchanov in view of Letunovskiy further discloses:
wherein constructing the
optimized machine-learned model based on the cost function comprises, for a model layer N of the plurality of model layers M:
determining, by the computing system for a candidate model layer for the model layer N, that the cost metrics associated with each candidate model layer constructed for the model layer N+1 based on the candidate model layer are greater than a maximum cost; and (Letunovskiy, ¶[0196-0197]: “The total latency L may be obtains, wherein L is the total latency on device H ( equal to the sum of the block latencies of all blocks in the architecture A) and L1, ... , LS are latencies of respective blocks B1, ... , BS, where S is the number of stages of architecture A. If L >Lmax· the procedure ends, because the scaled architecture does not fulfil the latency condition. If L is not greater than the maximum latency, then the procedure continues with step 840.”)
filtering, by the computing system, the candidate model layer from inclusion in the optimized machine-learned model. (Letunovskiy, ¶[0024]: “In an implementation, the searching for the one or more NN architectures comprises performing K times, K being a positive integer, the following steps: pseudo-randomly selecting a first set of candidate architectures from the search space: obtaining a second set of candidate architectures by removing from the first set of candidates those candidate architectures which do not satisfy a predefined condition including latency and/or accuracy”)
Regarding claim 10, this claim is rejected under the same rationale with claim 3 as they are analogous claims.
Regarding claim 11, the combination of Molchanov and Letunovskiy discloses all the limitations of claim 4 (as shown in the rejections above).
Molchanov in view of Letunovskiy further discloses:
wherein, prior to selecting the one or more search options from a plurality of layer search options, the method comprises receiving, by the computing system, an optimization request indicative of a quantity of layers M and the maximum cost. (Molchanov, Pg. 3, Col. 1, Section 2: “Our goal in this paper is to accelerate a given pre-trained teacher network by replacing its inefficient operations with more efficient alternatives.”, Pg. 3, Section 2.1, Col. 2, ¶[2]: “The problem of optimal selection of operations is often tackled in NAS. This problem is usually formulated as a bilevel optimization that selects operations and optimizes their weights jointly [40, 103].Finding the optimal architecture in hardware-aware NAS reduces to:
PNG
media_image3.png
192
543
media_image3.png
Greyscale
where bi Є RM is a vector of corresponding cost of each student operation (latency, number of parameters, FLOPs, etc.) in layer i.)
Regarding claim 12, this claim is rejected under the same rationale with claim 4 as they are analogous claims.
Regarding claim 13, this claim is rejected under the same rationale with claim 5 as they are analogous claims.
Regarding claim 17, this claim is rejected under the same rationale with claim 9 as they are analogous claims.
Regarding claim 18, this claim is rejected under the same rationale with claim 10 as they are analogous claims.
Claim(s) 7-8 and 15-16 are rejected under 35 U.S.C. 103 as being unpatentable over Molchanov et al. (“LANA: Latency Aware Network Acceleration”) in view of Letunovskiy et al. (US 2024/0169213 A1) further in view of Dai et al. (“FBNetV3: Joint Architecture-Recipe Search using Predictor Pretraining”).
Regarding claim 7, the combination of Molchanov and Letunovskiy discloses all the limitations of claim 4 (as shown in the rejections above).
Molchanov in view of Letunovskiy fails to disclose:
wherein constructing the optimized machine-learned model comprises, for each layer of the optimized machine-learned model: determining, by the computing system, a candidate layer cluster of the plurality of candidate layer clusters for the layer based on the cost function and the range of cost metrics associated with the candidate layer cluster; and
However, Dai explicitly discloses:
wherein constructing the optimized machine-learned model comprises, for each layer of the optimized machine-learned model: determining, by the computing system, a candidate layer cluster of the plurality of candidate layer clusters for the layer based on the cost function and the range of cost metrics associated with the candidate layer cluster; and (Dai, Pg. 11, Col. 1, Section A.3: “The selection of m best-performing samples in the constrained iterative optimization involves two steps: (1) equally divide the FLOP range into m bins and (2) pick the sample with the highest predicted score within each bin.”, Pg. 4, Col. 1, Section 3.3: “As mentioned prior, our goal is to find the most accurate architecture and training recipe combination under given resource constraints. We thus formulate the architecture search as a constrained optimization problem:
PNG
media_image8.png
61
525
media_image8.png
Greyscale
where A, h, and Ω refer to the neural network architecture, training recipe, and designed search space, respectively. acc maps the architecture and training recipe to accuracy. gi(A) and ɣ refer to the formula and count of resource constraints, such as computational cost, storage cost, and run-time latency.”)
selecting, by the computing system, a candidate layer from the candidate layer cluster based on the cost function and the cost metrics associated with one or more layers selected prior to the candidate layer. (Dai, Pg. 4, Section 3.4: “We evaluate the score for each child with the pretrained accuracy predictor u, and select top K highest-scoring candidates for the next generation”, Pg. 11, Col. 1, Section A.3: “The selection of m best-performing samples in the constrained iterative optimization involves two steps: (1) equally divide the FLOP range into m bins and (2) pick the sample with the highest predicted score within each bin.”, Pg. 4, Col. 1, Section 3.3: “As mentioned prior, our goal is to find the most accurate architecture and training recipe combination under given resource constraints. We thus formulate the architecture search as a constrained optimization problem:
PNG
media_image8.png
61
525
media_image8.png
Greyscale
)
The combination of Molchanov, Letunovskiy, Gu and Dai are analogous art because they are in the same field of training time series data. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention, having the teachings of Molchanov, Letunovskiy, Gu and Dai before them, to modify the teachings of M Molchanov, Letunovskiy and Gu to include the teachings of Dai determining a candidate layer cluster using the cost function and the cluster’s cost metric range to efficiently identify candidates satisfying the applicable resource bydget while achieving the desired performance cost tradeoff.
Regarding claim 8, the combination of Molchanov and Letunovskiy discloses all the limitations of claim 4 (as shown in the rejections above).
Molchanov in view of Letunovskiy fails to disclose:
wherein grouping the plurality of candidate layers into the plurality of candidate layer clusters further comprises storing, by the computing system, layer selection information indicative of the plurality of candidate layer clusters and each of the plurality of candidate layers; and (Molchanov, Pg. 4, Fig. 1: “Operation Selection Phase: We estimate and record in a lookup table the reduction of network accuracy and latency from replacing a teacher operation with one of the student operations.”)
wherein determining the candidate layer cluster of the plurality of candidate layer clusters comprises determining, by the computing system, the candidate layer cluster of the plurality of candidate layer clusters for the layer based on the cost function and layer selection information. (Dai, Pg. 11, Col. 1, Section A.3: “The selection of m best-performing samples in the constrained iterative optimization involves two steps: (1) equally divide the FLOP range into m bins and (2) pick the sample with the highest predicted score within each bin.”, Pg. 4, Col. 1, Section 3.3: “As mentioned prior, our goal is to find the most accurate architecture and training recipe combination under given resource constraints. We thus formulate the architecture search as a constrained optimization problem:
PNG
media_image8.png
61
525
media_image8.png
Greyscale
where A, h, and Ω refer to the neural network architecture, training recipe, and designed search space, respectively. acc maps the architecture and training recipe to accuracy. gi(A) and ɣ refer to the formula and count of resource constraints, such as computational cost, storage cost, and run-time latency.”)
Regarding claim 15, this claim is rejected under the same rationale with claim 7 as they are analogous claims.
Regarding claim 16, this claim is rejected under the same rationale with claim 8 as they are analogous claims.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to AMY TRAN whose telephone number is (571)270-0693. The examiner can normally be reached Monday - Friday 7:30 am - 5:00 pm EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, David Yi can be reached at (571) 270 7519. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/AMY TRAN/Examiner, Art Unit 2126
/DAVID YI/Supervisory Patent Examiner, Art Unit 2126